I’d view interpretability as a necessary evil that can make omnicidal futures less likely while building capable AI without any robust solutions to the alignment problem. (To be clear, my preferred solution would be not building highly capable AIs before we find solutions for alignment.)
If society was somehow hell-bent on creating one (or a few) billion-fold clonable superperson(s), with unproven technology that may end up with very alien motivational systems, I’d also suggest mind-reading on those before we hand over the world to them. I think it’s much less important to read minds of ordinary humans, which are less capable, singular and have vaguely human-shaped values—thus the trade-off looks worse there.
I’d view interpretability as a necessary evil that can make omnicidal futures less likely while building capable AI without any robust solutions to the alignment problem. (To be clear, my preferred solution would be not building highly capable AIs before we find solutions for alignment.)
If society was somehow hell-bent on creating one (or a few) billion-fold clonable superperson(s), with unproven technology that may end up with very alien motivational systems, I’d also suggest mind-reading on those before we hand over the world to them. I think it’s much less important to read minds of ordinary humans, which are less capable, singular and have vaguely human-shaped values—thus the trade-off looks worse there.