My impression was that one of the reason for continuing the advance in AI capabilities following that level of handoff, is that it would no longer be up to us whether continued AI capabilities growth happenedm and it would no longer be up to us whether we survived. Either we already succeeded in aligning AGI well enough by then to trust it to manage the ascent to ASI (or be wise enough to refuse to build it), or we already doomed ourselves.
I don’t think that’s true, and I tried to argue in the Pausing after unipolarity section that different level of alignment and wisdom is needed for the two goals.
Imagine that individual AI instances are about as smart, thoughtful and benevolent as the average capabilities researcher at an AI company right now. I think setting up a large army of these instances to maintain the peace might work pretty well without them fully taking over the world. With some clever checks-and-balances mechanisms alluded to in this post, I think this collective of moderately well-meaning AIs could remember their oaths to remain loyal to human democracy and function well. But I’m much more scared of tasking the collective of these AI instances to decide if it’s safe to scale further and then solve the alignment of the next generations well.
So I think there is a very wide gap between how wise and aligned the AIs need to be to trust them to do a peace-keeping operation with a narrow scope vs to trust them to build the superintelligence whose values will determine what the galaxies get filled with.
My impression was that one of the reason for continuing the advance in AI capabilities following that level of handoff, is that it would no longer be up to us whether continued AI capabilities growth happenedm and it would no longer be up to us whether we survived. Either we already succeeded in aligning AGI well enough by then to trust it to manage the ascent to ASI (or be wise enough to refuse to build it), or we already doomed ourselves.
I don’t think that’s true, and I tried to argue in the Pausing after unipolarity section that different level of alignment and wisdom is needed for the two goals.
Imagine that individual AI instances are about as smart, thoughtful and benevolent as the average capabilities researcher at an AI company right now. I think setting up a large army of these instances to maintain the peace might work pretty well without them fully taking over the world. With some clever checks-and-balances mechanisms alluded to in this post, I think this collective of moderately well-meaning AIs could remember their oaths to remain loyal to human democracy and function well. But I’m much more scared of tasking the collective of these AI instances to decide if it’s safe to scale further and then solve the alignment of the next generations well.
So I think there is a very wide gap between how wise and aligned the AIs need to be to trust them to do a peace-keeping operation with a narrow scope vs to trust them to build the superintelligence whose values will determine what the galaxies get filled with.
I hope so. To me that feels plausible, but probably a fine edge to walk.