That is the Harmony game that I mention in fn. 3 and 7. In that game, refrain weakly dominates (so both players prefer to refrain unilaterally). When I started writing this post I actually thought it would play a bigger role, but in the AI case it only happens on the knife edge of P(Doom) = 1 (which you can see in the visualization if you drag the slider to P(Doom) = 1).
Actually Harmony-type games can matter more if you relax the assumption that only the winner’s AI’s alignment matters. If even a loser’s AI being misaligned can still cause Doom, then Harmony (Nash equilibrium only at (Pause, Pause)) can happen for the whole top part of the probability distribution; where the boundary is depends on how likely a loser’s misaligned AI is to cause Doom.
That is the Harmony game that I mention in fn. 3 and 7. In that game, refrain weakly dominates (so both players prefer to refrain unilaterally). When I started writing this post I actually thought it would play a bigger role, but in the AI case it only happens on the knife edge of P(Doom) = 1 (which you can see in the visualization if you drag the slider to P(Doom) = 1).
Actually Harmony-type games can matter more if you relax the assumption that only the winner’s AI’s alignment matters. If even a loser’s AI being misaligned can still cause Doom, then Harmony (Nash equilibrium only at (Pause, Pause)) can happen for the whole top part of the probability distribution; where the boundary is depends on how likely a loser’s misaligned AI is to cause Doom.