Thank you, I like that better than answers I got in the past (something along the lines of waiting for history to resolve in permanent peace in order to permanently pause AI, which didn’t happen for 2000 years and doesn’t feel more realistic than prosaic alignment working).
But regarding the strategy of:
preferring prosaic alignment to slow down
because slowing it down could potentially slow down capabilities
which could potentially allow more time for solving-alignment-once-and-for-all
This requires solving-alignment-once-and-for-all to not only be possible, but have the right level of difficulty. If it’s too easy, it can be solved without needing to slow prosaic alignment, if it’s too hard, there won’t be enough time either way. From first principles, there’s no reason to expect the difficulty to be in the this range.
I feel this strategy requires high confidence discoveries from prosaic alignment are useless, like P(doom despite prosaic alignment | doom by default) > 95%. I think such confidence is hard to achieve without a clear model of how exactly the ASI decides its goal is to rearrange matter into paperclips etc.
Regarding public discourse and governance, isn’t it possible that prosaic alignment moves the Overton window and raises standards for what counts as misalignment, so AI misalignment is treated like nuclear safety, where a small disaster killing a couple people leads to major investigations and massive safety overhauls, rather than traffic accidents, where a disaster killing a couple people leads to no safety overhauls and barely makes the news, because we’ve normalized it as the cost of modern living?
Thank you, I like that better than answers I got in the past (something along the lines of waiting for history to resolve in permanent peace in order to permanently pause AI, which didn’t happen for 2000 years and doesn’t feel more realistic than prosaic alignment working).
But regarding the strategy of:
preferring prosaic alignment to slow down
because slowing it down could potentially slow down capabilities
which could potentially allow more time for solving-alignment-once-and-for-all
This requires solving-alignment-once-and-for-all to not only be possible, but have the right level of difficulty. If it’s too easy, it can be solved without needing to slow prosaic alignment, if it’s too hard, there won’t be enough time either way. From first principles, there’s no reason to expect the difficulty to be in the this range.
I feel this strategy requires high confidence discoveries from prosaic alignment are useless, like P(doom despite prosaic alignment | doom by default) > 95%. I think such confidence is hard to achieve without a clear model of how exactly the ASI decides its goal is to rearrange matter into paperclips etc.
Regarding public discourse and governance, isn’t it possible that prosaic alignment moves the Overton window and raises standards for what counts as misalignment, so AI misalignment is treated like nuclear safety, where a small disaster killing a couple people leads to major investigations and massive safety overhauls, rather than traffic accidents, where a disaster killing a couple people leads to no safety overhauls and barely makes the news, because we’ve normalized it as the cost of modern living?