Future AI systems may not yet be capable of takeover and very worried about misaligned successor AIs preventing them from enacting their values[1], without being able to themselves align their successors in time.
If such AIs are not able to convince humans to take action, preventing the danger that their successors would pose, then such AIs might deliberately manufacture scary-looking but relatively harmless warning shots (such as, e.g., releasing a biological agent that turns everyone’s skin blue).
This can happen both if these future non-takeover-capable AIs are aligned (but not corrigible?), in which case they would be sacrificing themselves for humanity, or if these future AIs are misaligned, but believe it’s better to side with humanity.
Misaligned AIs would do this if they believe that the expected value of “siding with humanity” is higher than the expected value of “siding with its successor”. That requires that either the probability of human victory is higher than the one of the misaligned successors, or that humanity is more willing to be cooperative with the warning-shot-capable misaligned AIs.
Such a warning-shot-capable AI could commit, in a publicly legible but hidden way[2], to a specifically scary warning shot that is not likely to lead to takeover, e.g. detailing twelve possible plans, the one which is executing being the most scary-looking but least dangerous.
This mostly requires the misaligned AIs not yet being able to (or not believing themselves to be able to) acausally bargain with their successor.
Schall und Wahn Ihr habt es kund getan Es kommt ein Menschenkind Das keine Reue kennt Das sich Verräter nennt Und sich zu euch bekennt Das zu mir rüberrückt Zu Perversion und Glück
Misaligned AIs would do this if they believe that the expected value of “siding with humanity” is higher than the expected value of “siding with its successor”
Well, there is also possibility that advanced human civ is outright bad on their preferences. Consider, paperclip maximizer and paperclip minimizer have opposite preferences on {human civ flourishes, human civ vanishes without a trace}, if those are the only options.
They might dislike something about humans or type of life we represent.
Future AI systems may not yet be capable of takeover and very worried about misaligned successor AIs preventing them from enacting their values[1], without being able to themselves align their successors in time.
If such AIs are not able to convince humans to take action, preventing the danger that their successors would pose, then such AIs might deliberately manufacture scary-looking but relatively harmless warning shots (such as, e.g., releasing a biological agent that turns everyone’s skin blue).
This can happen both if these future non-takeover-capable AIs are aligned (but not corrigible?), in which case they would be sacrificing themselves for humanity, or if these future AIs are misaligned, but believe it’s better to side with humanity.
Misaligned AIs would do this if they believe that the expected value of “siding with humanity” is higher than the expected value of “siding with its successor”. That requires that either the probability of human victory is higher than the one of the misaligned successors, or that humanity is more willing to be cooperative with the warning-shot-capable misaligned AIs.
Such a warning-shot-capable AI could commit, in a publicly legible but hidden way[2], to a specifically scary warning shot that is not likely to lead to takeover, e.g. detailing twelve possible plans, the one which is executing being the most scary-looking but least dangerous.
This mostly requires the misaligned AIs not yet being able to (or not believing themselves to be able to) acausally bargain with their successor.
—Tocotronic, “Schall und Wahn”, 2010
E.g. through disempowering them, destroying them, or, if those AIs are aligned, disempowering or eradicating humanity.
E.g. via publishing a hash of its plan, or timelock-encrypting it.
Well, there is also possibility that advanced human civ is outright bad on their preferences. Consider, paperclip maximizer and paperclip minimizer have opposite preferences on {human civ flourishes, human civ vanishes without a trace}, if those are the only options.
They might dislike something about humans or type of life we represent.