Great post! Big upvote. As usual I’d like to see the math presaged and summarized in text in the abstract or intro, but that’s a nitpick. I’ll summarize when I reference it.
One small issue is that Win payoff could actually be as big or bigger than Doom loss. Doom just kills everyone, while Win means using the whole resources of the lightcone for your values. Depending on your assumptions about likely futures if no one builds it, this could make W a much larger absolute value than D.
But actual decision-makers don’t make decisions with utilitarian math, so I think that’s irrelevant.
Intuitively, even a modest chance of killing everyone is worth avoiding. The payoff of winning for you, subjectively, isn’t nearly as good as that is bad...
Hm, as I write that I realize it’s wrong in an important way. For an individual decision-maker, building ASI rapidly can easily be the difference between suffering and dying from aging and living approximately forever. And here the point that real people in power are rarely utilitarians cuts the other way.
The other point is that even if the other guy wins, they’re unlikely to wipe you out completely, and most of your values will probably control the future if the other team shares many of your values.
Anyway, excellent post. This is one discussion we should be having much more. Politicians and the public haven’t started thinking in these terms, but they will, and we should be ready to make that debate as sensible as possible before it happens.
Thanks! To your point about most decision-makers not being utilitarians, most decision-makers also aren’t taking life-extension into account in their decisions on this (well, government decision-makers anyway; lab CEOs very well might be).
The other point is that even if the other guy wins, they’re unlikely to wipe you out completely, and most of your values will probably control the future if the other team shares many of your values.
I agree with this. This impacts the Lose payoff, but I can’t figure out if it impacts the Win:Lose ratio, or only the Lose:Doom ratio. If it impacts the Win:Lose ratio, there’s an interesting point I didn’t highlight in the piece which is that, if Win > |Lose|, there’s a zone in the very low P(Doom) region that becomes Deadlock, where both parties prefer mutual racing to mutual pausing (you can see it if you play with the Win:Lose ratio in the visualizer in the Appendix). But this doesn’t impact q*, so I didn’t highlight it.
I think decision-makers will take life extension into account just about exactly when they take existential risk into account. They are both fairly obvious implications of taking RSI and ASI seriously.
I think the second point is crucial and equally important to the main point. I think actually most humans in charge of ASI would produce very good futures in the long term. Unfortunately their political enemies are probably the absolute worst off in the short term so perhaps that doesn’t matter for the game theory from the leaders perspective. But it matters a lot for the group will, which might matter for electing or otherwise choosing a leader.
I am hoping we elect a next US president who both takes alignment x-risk seriously and who takes the possibility of cooperating with China seriously. That’s one reason I think it’s timely to hammer out this logic now.
Great post! Big upvote. As usual I’d like to see the math presaged and summarized in text in the abstract or intro, but that’s a nitpick. I’ll summarize when I reference it.
One small issue is that Win payoff could actually be as big or bigger than Doom loss. Doom just kills everyone, while Win means using the whole resources of the lightcone for your values. Depending on your assumptions about likely futures if no one builds it, this could make W a much larger absolute value than D.
But actual decision-makers don’t make decisions with utilitarian math, so I think that’s irrelevant.
Intuitively, even a modest chance of killing everyone is worth avoiding. The payoff of winning for you, subjectively, isn’t nearly as good as that is bad...
Hm, as I write that I realize it’s wrong in an important way. For an individual decision-maker, building ASI rapidly can easily be the difference between suffering and dying from aging and living approximately forever. And here the point that real people in power are rarely utilitarians cuts the other way.
The other point is that even if the other guy wins, they’re unlikely to wipe you out completely, and most of your values will probably control the future if the other team shares many of your values.
Anyway, excellent post. This is one discussion we should be having much more. Politicians and the public haven’t started thinking in these terms, but they will, and we should be ready to make that debate as sensible as possible before it happens.
Thanks! To your point about most decision-makers not being utilitarians, most decision-makers also aren’t taking life-extension into account in their decisions on this (well, government decision-makers anyway; lab CEOs very well might be).
I agree with this. This impacts the Lose payoff, but I can’t figure out if it impacts the Win:Lose ratio, or only the Lose:Doom ratio. If it impacts the Win:Lose ratio, there’s an interesting point I didn’t highlight in the piece which is that, if Win > |Lose|, there’s a zone in the very low P(Doom) region that becomes Deadlock, where both parties prefer mutual racing to mutual pausing (you can see it if you play with the Win:Lose ratio in the visualizer in the Appendix). But this doesn’t impact q*, so I didn’t highlight it.
I think decision-makers will take life extension into account just about exactly when they take existential risk into account. They are both fairly obvious implications of taking RSI and ASI seriously.
I think the second point is crucial and equally important to the main point. I think actually most humans in charge of ASI would produce very good futures in the long term. Unfortunately their political enemies are probably the absolute worst off in the short term so perhaps that doesn’t matter for the game theory from the leaders perspective. But it matters a lot for the group will, which might matter for electing or otherwise choosing a leader.
I am hoping we elect a next US president who both takes alignment x-risk seriously and who takes the possibility of cooperating with China seriously. That’s one reason I think it’s timely to hammer out this logic now.