I don’t think that’s true, and I tried to argue in the Pausing after unipolarity section that different level of alignment and wisdom is needed for the two goals.
Imagine that individual AI instances are about as smart, thoughtful and benevolent as the average capabilities researcher at an AI company right now. I think setting up a large army of these instances to maintain the peace might work pretty well without them fully taking over the world. With some clever checks-and-balances mechanisms alluded to in this post, I think this collective of moderately well-meaning AIs could remember their oaths to remain loyal to human democracy and function well. But I’m much more scared of tasking the collective of these AI instances to decide if it’s safe to scale further and then solve the alignment of the next generations well.
So I think there is a very wide gap between how wise and aligned the AIs need to be to trust them to do a peace-keeping operation with a narrow scope vs to trust them to build the superintelligence whose values will determine what the galaxies get filled with.
I don’t think that’s true, and I tried to argue in the Pausing after unipolarity section that different level of alignment and wisdom is needed for the two goals.
Imagine that individual AI instances are about as smart, thoughtful and benevolent as the average capabilities researcher at an AI company right now. I think setting up a large army of these instances to maintain the peace might work pretty well without them fully taking over the world. With some clever checks-and-balances mechanisms alluded to in this post, I think this collective of moderately well-meaning AIs could remember their oaths to remain loyal to human democracy and function well. But I’m much more scared of tasking the collective of these AI instances to decide if it’s safe to scale further and then solve the alignment of the next generations well.
So I think there is a very wide gap between how wise and aligned the AIs need to be to trust them to do a peace-keeping operation with a narrow scope vs to trust them to build the superintelligence whose values will determine what the galaxies get filled with.
I hope so. To me that feels plausible, but probably a fine edge to walk.