Sam Altman: “This is the first security incident that I have felt very viscerally. I’ve been a little surprised that more people don’t feel it so viscerally.
We paused training. We have to figure out how to secure our sandboxing in a world of multiple zero days being chained together.
We may have to pace the rate of AI development to give ourselves enough time for society to harden around these new capability levels.”
I just want to surface the hypothesis that known liar Sam Altman could be lying about this. Although seems somewhat unlikely given that it would (probably?) be easy for an OpenAI insider to contradict him.
I want to surface the hypothesis that sometimes, you get a glimmer of hope. While no tweet should be taken prima facie, the tweet alone is above-expectation.
Doesn’t seem like the type of lie that he would tell. It’s a specific actual physical falsifiable-in-practice-if-false claim, not a value statement, personal judgement, expression of future intention, framing on an event, secret known by very few people all of whom can be reliably expected to not share, or anything else lying about which is a free action.
In particular, note that tons of OpenAI employees would know that he’s lying about this. While it’s possible they wouldn’t share publicly, I don’t think he would want to show himself to be so blatant a blatant liar to them all.
He could have “meant” that they’d paused training of that specific model. I would be astonished if they’d at any point stopped running all code that would result in updates to any model’s weights.
Mm, maybe. I haven’t watched the interview yet, I’d need to hear it in context to say more.
… That said, what occurs to me now is that it doesn’t necessarily matter whether OpenAI did what Altman seems to say it did. It’s notable that he even appears to want to make it sound like OpenAI has paused all frontier-model training over this incident. That public-comms move is actually quite extraordinary by itself!
I think the cynical reading here is that their security is shit, meaning their models constantly break out, meaning their training environments are now basically ineffective and just train the models how to break out of OAI’s sandboxes, so OAI is not really capable of training models and so it costs them nothing or less than nothing to “pause training” to work on hardening their sandboxes.
Even if sandboxes are rarely broken, it could still really hurt the sample efficiency if the majority of great successes are caused by it finding new ways OAI’s sandboxes are broken, or it becoming more motivated to search for sandbox failures as step 1.
Note this is still good news, as it indicates some alignment between the goals of alignment (specifically training models with intentionality, care, and security), and capabilities.
It’s sufficiently vague that I wouldn’t read too much into it, especially from Altman. For example, there’s little security risk posed by pretraining/midtraining, so I assume that’s continuing. They could have paused training on the specific model which broke out of the sandbox, but training on other systems is ongoing. They could have patched the sandbox vulnerabilities they detected and then resumed training, etc.
A similar example of such vagueness was the promise by Altman to ‘[dedicate] 20% of the compute we’ve secured to date to this effort [of the Superalignment team]’. Later reporting shows what happened:
It [Superalignment] was a task so important that the company said in its announcement that it would commit “20% of the compute we’ve secured to date over the next four years” to the effort.
But a half dozen sources familiar with the Superalignment team’s work said that the group was never allocated this compute. Instead, it received far less in the company’s regular compute allocation budget, which is reassessed quarterly.
One source familiar with the Superalignment team’s work said that there were never any clear metrics around exactly how the 20% amount was to be calculated, leaving it subject to wide interpretation. For instance, the source said the team was never told whether the promise meant “20% each year for four years” or “5% a year for four years” or some variable amount that could wind up being “1% or 2% for the first three years, and then the bulk of the commitment in the fourth year.” In any case, all the sources Fortune spoke to for this story confirmed that the Superalignment team was never given anything close to 20% of OpenAI’s secured compute as of July 2023.
OpenAI researchers can also make requests for what is known as “flex” compute—access to additional GPU capacity beyond what has been budgeted—to deal with new projects between the quarterly budgeting meetings. But flex requests from the Superalignment team were routinely rejected by higher ups, these sources said.
Their updated post on the Huggingface incident states:
No models planned for upcoming release were involved in exploiting Hugging Face. The pre-release model mentioned in our blog post is an internal-only research prototype and was never intended for public release. Following the incident, we deactivated, encrypted, and restricted it from research access.
Anthropic could single-handedly force a pause among all major AI companies by joining in. It would lend an incredible amount of credibility to the OpenAI pause. Any other company who didn’t join would be hated, and possibly coerced by governments.
I’ll offer the hypothesis that equipment, energy, and data constraints might be putting hard limits on training. This is a way to justify a slowdown without making a negative statement about the future prospects for business growth.
OpenAI has stopped training.
Sam Altman: “This is the first security incident that I have felt very viscerally. I’ve been a little surprised that more people don’t feel it so viscerally.
We paused training. We have to figure out how to secure our sandboxing in a world of multiple zero days being chained together.
We may have to pace the rate of AI development to give ourselves enough time for society to harden around these new capability levels.”
https://x.com/AISafetyMemes/status/2082222516454785296
I just want to surface the hypothesis that known liar Sam Altman could be lying about this. Although seems somewhat unlikely given that it would (probably?) be easy for an OpenAI insider to contradict him.
‘We paused training’ is also compatible with having paused training, and then restarted training later.
I want to surface the hypothesis that sometimes, you get a glimmer of hope. While no tweet should be taken prima facie, the tweet alone is above-expectation.
It’s also possible that “pausing training” is a term-of-art that doesn’t mean what the natural language interpretation suggests.
Doesn’t seem like the type of lie that he would tell. It’s a specific actual physical falsifiable-in-practice-if-false claim, not a value statement, personal judgement, expression of future intention, framing on an event, secret known by very few people all of whom can be reliably expected to not share, or anything else lying about which is a free action.
In particular, note that tons of OpenAI employees would know that he’s lying about this. While it’s possible they wouldn’t share publicly, I don’t think he would want to show himself to be so blatant a blatant liar to them all.
He could have “meant” that they’d paused training of that specific model. I would be astonished if they’d at any point stopped running all code that would result in updates to any model’s weights.
Mm, maybe. I haven’t watched the interview yet, I’d need to hear it in context to say more.
… That said, what occurs to me now is that it doesn’t necessarily matter whether OpenAI did what Altman seems to say it did. It’s notable that he even appears to want to make it sound like OpenAI has paused all frontier-model training over this incident. That public-comms move is actually quite extraordinary by itself!
I think the cynical reading here is that their security is shit, meaning their models constantly break out, meaning their training environments are now basically ineffective and just train the models how to break out of OAI’s sandboxes, so OAI is not really capable of training models and so it costs them nothing or less than nothing to “pause training” to work on hardening their sandboxes.
Even if sandboxes are rarely broken, it could still really hurt the sample efficiency if the majority of great successes are caused by it finding new ways OAI’s sandboxes are broken, or it becoming more motivated to search for sandbox failures as step 1.
Note this is still good news, as it indicates some alignment between the goals of alignment (specifically training models with intentionality, care, and security), and capabilities.
It’s sufficiently vague that I wouldn’t read too much into it, especially from Altman. For example, there’s little security risk posed by pretraining/midtraining, so I assume that’s continuing. They could have paused training on the specific model which broke out of the sandbox, but training on other systems is ongoing. They could have patched the sandbox vulnerabilities they detected and then resumed training, etc.
A similar example of such vagueness was the promise by Altman to ‘[dedicate] 20% of the compute we’ve secured to date to this effort [of the Superalignment team]’. Later reporting shows what happened:
Their updated post on the Huggingface incident states:
Altman today also told reporters that the model in question has been ‘deactivated permanently’. https://x.com/Dareasmunhoz/status/2082511870376743110
So my assumption here is that he means they paused training of the model in question and have rolled back to some previous model.
Didn’t he say something almost identical about GPT-2, in 2019?
Anthropic could single-handedly force a pause among all major AI companies by joining in. It would lend an incredible amount of credibility to the OpenAI pause. Any other company who didn’t join would be hated, and possibly coerced by governments.
Luckily, that is in motion https://www.washingtonpost.com/technology/2026/07/29/openai-anthropic-endorse-call-government-pace-ai-progress/
I’ll offer the hypothesis that equipment, energy, and data constraints might be putting hard limits on training. This is a way to justify a slowdown without making a negative statement about the future prospects for business growth.