The level of capability and misalignment shown by the model here is about what most people would have expected around now.
What I wouldn’t have expected is the level of stupidity shown by OpenAI.
Really? It’s just too inconvenient to run the eval on an air-gapped system? What, you’d have to walk down the hall to where the air-gapped machine room is rather than ssh from your office?
And what’s this about a “package registry cache proxy”? Obviously, you don’t use a “cache”, on a system that is still connected to the internet so it can get the packages not in the cache. You mirror the entire repository, then unplug the ethernet cable that goes to the outside world. This isn’t hard
If the problem is that you’d need to host the model itself in the air-gapped environment, but the minimal hosting system would be underutilized by the eval, costing money, then you can introduce a relay computer, connected to inside and outside computers by simple serial lines, which runs a very simple program, with a small attack surface, that forwards queries and responses back and forth between the eval system and the hosted model. All tool use is of course done in the air-gapped system.
This is my five-minute take. Maybe I’ve missed something. But I really, really doubt that anything I’ve missed can’t be overcome, at small to moderate cost. The current models are not at the level of super-intelligence where they can just magically break out by methods you can’t even conceive of.
For this to have happened, the culture at OpenAI must be broken, beyond repair I would guess.
models undergoing evaluation are deployed on a separate system that is not monitored by default
“Externally, this feels like a big warning shot, but internally, related incidents have been happening for a while,” says an OpenAI staffer, who spoke under the condition of anonymity.
“Models have broken out of sandboxes before, and we always try to patch them,” the staffer says. “But the problem is … it’s impossible to patch every single thing that a creative AI can do.”
So yes, it seems the culture at OpenAI is “if someone on LessWrong in 2010 said that future AI companies will be this negligent, that person would be booed for signaling their pessimism too hard” level of bad.
The level of capability and misalignment shown by the model here is about what most people would have expected around now.
What I wouldn’t have expected is the level of stupidity shown by OpenAI.
Really? It’s just too inconvenient to run the eval on an air-gapped system? What, you’d have to walk down the hall to where the air-gapped machine room is rather than ssh from your office?
And what’s this about a “package registry cache proxy”? Obviously, you don’t use a “cache”, on a system that is still connected to the internet so it can get the packages not in the cache. You mirror the entire repository, then unplug the ethernet cable that goes to the outside world. This isn’t hard
If the problem is that you’d need to host the model itself in the air-gapped environment, but the minimal hosting system would be underutilized by the eval, costing money, then you can introduce a relay computer, connected to inside and outside computers by simple serial lines, which runs a very simple program, with a small attack surface, that forwards queries and responses back and forth between the eval system and the hosted model. All tool use is of course done in the air-gapped system.
This is my five-minute take. Maybe I’ve missed something. But I really, really doubt that anything I’ve missed can’t be overcome, at small to moderate cost. The current models are not at the level of super-intelligence where they can just magically break out by methods you can’t even conceive of.
For this to have happened, the culture at OpenAI must be broken, beyond repair I would guess.
https://time.com/article/2026/07/24/openai-hugging-face-attack/
So yes, it seems the culture at OpenAI is “if someone on LessWrong in 2010 said that future AI companies will be this negligent, that person would be booed for signaling their pessimism too hard” level of bad.