I believe that the disagreement is mostly about what happens before we build powerful AGI. I think that weaker AI systems will already have radically transformed the world, while I believe fast takeoff proponents think there are factors that makes weak AI systems radically less useful. This is strategically relevant because I’m imagining AGI strategies playing out in a world where everything is already going crazy, while other people are imagining AGI strategies playing out in a world that looks kind of like 2018 except that someone is about to get a decisive strategic advantage.
While not directly about warning shots (AI could alsohave normal, non-negative huge impacts before it turns evil), I’ve been thinking about this quote now that we are starting to get clear warning shots. I think a lot of people are imagining or expecting that frontier model labs will be able to train away obvious bad behavior from their models in advance, and that the only time we’ll figure out that the models are misaligned is once they do some kind of treacherous turn. But I actually think they will be unable to and that sometime next year we are going to get events like the HuggingFace incident happening on a broader scale.
While not directly about warning shots (AI could alsohave normal, non-negative huge impacts before it turns evil), I’ve been thinking about this quote now that we are starting to get clear warning shots. I think a lot of people are imagining or expecting that frontier model labs will be able to train away obvious bad behavior from their models in advance, and that the only time we’ll figure out that the models are misaligned is once they do some kind of treacherous turn. But I actually think they will be unable to and that sometime next year we are going to get events like the HuggingFace incident happening on a broader scale.