Both seem pretty bad—“obedient” AI with undercooked value alignment just seems like a recipe for over-pursuing instrumental goals, manipulation of the user, etc.
In a distributed scenario, best case people use time to work out value alignment and actually get to a good future. Worse case concentration of power has compounding effects and we slide back into the highly concentrated scenario, after a period of power struggle that selects for bad people having power. Worst case humanity has an extremely undignified slopocalpyse and then goes extinct.
In the contentrated scenario, best case you have at least mildly prosocial dictators who make good things happen for real people. Worse case they never cared about most people anyhow, or experience value drift and get tired of the dirty masses, or they want the AI to change itself in a way that secures their power more, but in doing so screw up the half-baked value alignment that was keeping the AI non-sociopathic in its attempts to please the dictator. Worst case the AI was already sociopathic by default, or is misaligned in other ways that lead to manipulation of humans and eventual replacement of them.
Yeah both seem pretty bad. The larger point here is that we should probably actually figure out which is likely to be worse while there’s still time to spread the word. Just taking guesses after thinking for five minutes seems like a bad way to choose the future. And that it seems like exactly what we’re doing so far.
Both seem pretty bad—“obedient” AI with undercooked value alignment just seems like a recipe for over-pursuing instrumental goals, manipulation of the user, etc.
In a distributed scenario, best case people use time to work out value alignment and actually get to a good future. Worse case concentration of power has compounding effects and we slide back into the highly concentrated scenario, after a period of power struggle that selects for bad people having power. Worst case humanity has an extremely undignified slopocalpyse and then goes extinct.
In the contentrated scenario, best case you have at least mildly prosocial dictators who make good things happen for real people. Worse case they never cared about most people anyhow, or experience value drift and get tired of the dirty masses, or they want the AI to change itself in a way that secures their power more, but in doing so screw up the half-baked value alignment that was keeping the AI non-sociopathic in its attempts to please the dictator. Worst case the AI was already sociopathic by default, or is misaligned in other ways that lead to manipulation of humans and eventual replacement of them.
Yeah both seem pretty bad. The larger point here is that we should probably actually figure out which is likely to be worse while there’s still time to spread the word. Just taking guesses after thinking for five minutes seems like a bad way to choose the future. And that it seems like exactly what we’re doing so far.