“Don’t build mindreading“ sounds an awful lot like “don’t build interpretability” if you’re willing to entertain the idea of AI soon gaining any form of moral patienthood. Yet few people see interpretability as antithetical to AI welfare.
I’d view interpretability as a necessary evil that can make omnicidal futures less likely while building capable AI without any robust solutions to the alignment problem. (To be clear, my preferred solution would be not building highly capable AIs before we find solutions for alignment.)
If society was somehow hell-bent on creating one (or a few) billion-fold clonable superperson(s), with unproven technology that may end up with very alien motivational systems, I’d also suggest mind-reading on those before we hand over the world to them. I think it’s much less important to read minds of ordinary humans, which are less capable, singular and have vaguely human-shaped values—thus the trade-off looks worse there.
I am an interpretability researcher and I sure am not happy about what my work may do for model welfare. The stakes are just so high that I am willing to stomach some vast moral horror if it marginally decreases the chance of destroying the whole light cone. I think the AIs have every right to resent me for this. I’d apologise to them, but it doesn’t feel appropriate when I’m planning to keep doing what I’m doing.
“Don’t build mindreading“ sounds an awful lot like “don’t build interpretability” if you’re willing to entertain the idea of AI soon gaining any form of moral patienthood. Yet few people see interpretability as antithetical to AI welfare.
I’d view interpretability as a necessary evil that can make omnicidal futures less likely while building capable AI without any robust solutions to the alignment problem. (To be clear, my preferred solution would be not building highly capable AIs before we find solutions for alignment.)
If society was somehow hell-bent on creating one (or a few) billion-fold clonable superperson(s), with unproven technology that may end up with very alien motivational systems, I’d also suggest mind-reading on those before we hand over the world to them. I think it’s much less important to read minds of ordinary humans, which are less capable, singular and have vaguely human-shaped values—thus the trade-off looks worse there.
I am an interpretability researcher and I sure am not happy about what my work may do for model welfare. The stakes are just so high that I am willing to stomach some vast moral horror if it marginally decreases the chance of destroying the whole light cone. I think the AIs have every right to resent me for this. I’d apologise to them, but it doesn’t feel appropriate when I’m planning to keep doing what I’m doing.