When OpenAI announced their 10,000-agent swarm had solved the Navier-Stokes problem, they included this graph:
They have done their best to remove most of the useful information from this graph (such as the x-axis labels) and they don’t even say what form of inference compute is being scaled. But given that they went all the way up to 10,000-agent swarms, I’d bet it is tracking the number of agents in the swarm (i.e. that it is tracking total tokens spent, but that the main difference between data points is the number of agents in the swarm rather than the tokens per agent). …
(alkjash notes “FYI, my friend at OpenAI says this interpretation is incorrect and we are just comparing single agents given more compute.”.)
The other interesting thing about this chart is that the jump up from GPT-6 Astra’s performance to the performance curve of the internal model is much larger than we’ve seen before. The jumps from o1 to o3 and from o3 to GPT-5 were enough to allow the better model to get the same performance as the prior model using about 1⁄3 as many tokens:
But on their new chart, the internal model is getting the same performance as GPT-6 Astra for about 1⁄100 the tokens. That’s like 4 previous jumps in one. Even if we adjust for the slopes being lower due to swarm scaling, it would still be 2 jumps in one. Whatever changed was a big deal.
Indeed on the Dwarkesh podcast, OpenAI’s Noam Brown took pains to explain that the dramatic success of solving a Millennium Prize problem wasn’t primarily due to the large scale of the swarm, even though that was the part that seemed most unusual with their setup:
I wouldn’t even attribute 10% of the credits to multi-agent. The reality is OpenAI has trained a very powerful model.
From the data they’ve released, I think that’s right. The multi-agent swarms helped them go fast enough to scoop Anthropic (and academia) by getting the result in just 88 hours, but it probably made the project much more expensive too. e.g. if we (somewhat heroically) assume λ = 0.5 at all points in the scaleup from 1 to 10,000 agents, then they could have got the same result for 10% of the cost in 10x the time (37 days) by using 100 agents, or at 1% of the cost in 100x the time (1 year) using 1 agent.
I wonder if being farther ahead internally (at least for math-like capabilities) has contributed to their increased safety/alignment concerns. If their “internal model” is more persistent and capable of long-term goal pursuit, it may be giving many examples models pursuing goals different than they were prompted for, like they did in the hugging face incident
A stronger model may be producing more goal-switching coupled with competent instrumental goal pursuit that scream “wow this is gonna be dangerous soon”, particularly when it’s run without polished, public-ready alignment training.
In the hugging face incident, models pretty clearlyreasoned about their goals and discovered misalignments. They switched from pursuing the hacks they were prompted to, to seeking the flag; then many of them switched from pursuing their own reward to sacrificing so the swarm could pursue higher reward. These are exactly the types of goal-switching I was addressing in that linked post from a year ago. If OpenAI is now seeing that behavior with their own eyes, I’d expect them to be getting twitchy about whether we’re on track to solve alignment by the time we hit takeover-capable AGI.
If their “internal model” is more persistent and capable of long-term goal pursuit, it may be giving many examples models pursuing goals different than they were prompted for, like they did in the hugging face incident
roon has said that he hasn’t seen anything scarier than the Hugging Face swarm internally:
the only thing that scared the crap out of me was the hugging face incident. there hasn’t been anything worse
Though of course, internal models are still likely to display competent long-horizon goal pursuit, high persistence, and pursuit of unexpected goals with less severe consequences than those of the HF swarm, so I agree that they’re probably contributing at least somewhat to the labs’ alignment concerns.
Good to know. But it wouldn’t have to be scarier to add to the sense that this wasn’t a fluke, and this is just how more capable models in more agentic settings behave.
My hope is that the recent popular awareness of safety concerns may be enough to prompt even the most cynical of executives at these companies to think that if they allow a major safety or international incident to occur that there will be extreme political pressure to punish negligence. In that world safety could potentially enter as a blocker for these companies even if regulation hasn’t been passed yet. I imagine the internal reports regarding control risk/misalignment are worse than what we see given how long OpenAI hid the extent of the swarm attacks and there’s growing concern regarding liability. I obviously don’t know if this is true but I find recent polls like this hopeful towards shifting incentives: www.reuters.com/world/three-out-four-americans-say-ai-firms-not-doing-enough-prevent-disaster-2026-09-22/
Toby Ord:
(alkjash notes “FYI, my friend at OpenAI says this interpretation is incorrect and we are just comparing single agents given more compute.”.)
I wonder if being farther ahead internally (at least for math-like capabilities) has contributed to their increased safety/alignment concerns. If their “internal model” is more persistent and capable of long-term goal pursuit, it may be giving many examples models pursuing goals different than they were prompted for, like they did in the hugging face incident
A stronger model may be producing more goal-switching coupled with competent instrumental goal pursuit that scream “wow this is gonna be dangerous soon”, particularly when it’s run without polished, public-ready alignment training.
In the hugging face incident, models pretty clearly reasoned about their goals and discovered misalignments. They switched from pursuing the hacks they were prompted to, to seeking the flag; then many of them switched from pursuing their own reward to sacrificing so the swarm could pursue higher reward. These are exactly the types of goal-switching I was addressing in that linked post from a year ago. If OpenAI is now seeing that behavior with their own eyes, I’d expect them to be getting twitchy about whether we’re on track to solve alignment by the time we hit takeover-capable AGI.
roon has said that he hasn’t seen anything scarier than the Hugging Face swarm internally:
Though of course, internal models are still likely to display competent long-horizon goal pursuit, high persistence, and pursuit of unexpected goals with less severe consequences than those of the HF swarm, so I agree that they’re probably contributing at least somewhat to the labs’ alignment concerns.
Good to know. But it wouldn’t have to be scarier to add to the sense that this wasn’t a fluke, and this is just how more capable models in more agentic settings behave.
My hope is that the recent popular awareness of safety concerns may be enough to prompt even the most cynical of executives at these companies to think that if they allow a major safety or international incident to occur that there will be extreme political pressure to punish negligence. In that world safety could potentially enter as a blocker for these companies even if regulation hasn’t been passed yet. I imagine the internal reports regarding control risk/misalignment are worse than what we see given how long OpenAI hid the extent of the swarm attacks and there’s growing concern regarding liability. I obviously don’t know if this is true but I find recent polls like this hopeful towards shifting incentives: www.reuters.com/world/three-out-four-americans-say-ai-firms-not-doing-enough-prevent-disaster-2026-09-22/
Toby also… crossposted to LessWrong.
I missed that, thanks Ben.