Last week’s Hugging Face breach was accidentally caused by OpenAI testing models’ cyber capabilities. From the OpenAI blogpost:
Last week, Hugging Face disclosed a new kind of security incident after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities. . . .
This incident occurred during an internal evaluation. . . . The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.
While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.
In other words, the models were supposed to solve a cyber test, but they found the solution by hacking their sandbox to get internet access and then hacking someone on the internet who already had the solution. If I understand correctly.
Covered in Axios, Fortune (but no info beyond the OpenAI blogpost).
Note: this is some legible evidence of humanity’s current inability to reliably steer AI systems. I expect current safeguards suffice to prevent incidents like this when someone (provider or operator) cares about implementing those safeguards (some safeguards weren’t enabled here). But I’m worried about AI risk because I think it’s quite possible that future AI systems still won’t be reliably steerable, they’ll seek long-term power (for instrumental reasons), and our safeguards won’t suffice to stop them.
We need to be prepared for a world where agents will cause ever larger harms in pursuit of their goals, whether or not they have a power-seeking drive or even care about the larger state of the world.
Wow, straight out of science fiction. I think this will be remembered as one of the first serious AI misalignment warning signs in history. Possibly an existence proof that deep emergent misalignment based on task misalignment from RL capabilities training is possible (how else could the model have decided that hacking Hugging Face’s production servers to cheat on an eval was a good idea?!)
I am not very surprised it happened, but I am surprised that it happened so soon (my prior must have been <5% of something like this happening in mid-2026).
After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities
I think an important invariant for AI companies is that at no point in training should your model be misaligned enough to try in earnest to escape your sandbox. For a sufficiently capable model this is irreversible, and if it happens in less capable models it should be seen as a failure of your alignment strategy, and you should update accordingly.
Agreed. I’d add that if this was a “helpful-only” model (borrowing Anthropic’s terminology) then I’d be less concerned about the incident, but as you mentioned, it still seems stupid to let such a model get this misaligned.
However, their own wording is “reduced cyber refusals” or “reduced safeguards” which is not confidence-inspiring (unless “reduced safeguards” genuinely just means “helpful-only” at OpenAI, but in which case I feel like they would have clarified it more). So it leads me to believe that genuine alignment failure has happened here. The fact that it already happened at this stage doesn’t shine a good light on OpenAI’s alignment efforts, to put it politely.
I also think that the observation that the first time we heard about this is when the incident is likely too big to sweep under the rug suggests that similar problems have happened at OpenAI many times before, we just haven’t heard about them.
So https://openai.com/index/safety-alignment-long-horizon-models/ was published yesterday. The NanoGPT speedrun PR was put up before May 14th (because PR 287 was referenced by PR 300, which was put on May 14th). OpenAI presumably did all of the things they said in that blog post (paused deployment, improved alignment, additional guardrails, etc) in the meantime.
This timing is kind of surprising! Maybe suggests that the arm of OpenAI responsible for publishing the first blog post did not know about OpenAI’s responsibility w.r.t. the HF incident yesterday. As it is, all of the claims about the additional work done seem a little… something.
Not actually specified by the blog post; I suppose it could be a different model entirely? But it does seem like a simpler explanation that it’s the ~same model, or at least within the same model lineage, than that OpenAI has 2 different unreleased highly cyber-capable models.
FYI OpenAI says in their blog post “This week, we published a blog on improving safety and alignment in an era of long horizon models. These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities.”
As for the timing I am very interested in what openAI knew when. Did Hugging Face reach out to OpenAI or visa versa? The blog post is very vague “OpenAI’s security team discovered this anomalous activity internally. Hugging Face’s security team and agents detected and stopped the activity on their infrastructure and had already begun containment and forensic reconstruction with their own open-source models when our teams connected. We are actively working with them to continue to investigate the incident. We are grateful for Hugging Face’s rapid and close collaboration on investigation and remediation.”
Did anybody actually predict this directly? As in, “I think this year OpenAI/Anthropic will have a model breach containment in a wildly antisocial way.”
This seems way beyond what people were expecting in terms of catastrophic, public misalignment, including “AI safety” people.
I expect many more similar incidents. Models have always done weird things and are not really doing fewer as they become capable of getting more important—and alarming—things done.
I don’t know if anyone made that specific prediction in so many words, but all of the following were in the water supply
Models are already quite good at cyber stuff, and rapidly getting better.
Software security is an absolute tire fire on many levels. If an actually-competent threat actor wants to pwn your complicated internet-connected software, they will find a way to do so.
If you stick an LLM in a loop, where the only way out of the loop is to make your monitor think the task is solved, the LLM will find a way to make the monitor think the task is solved. If the model isn’t capable of solving the issue in the intended way, it will execute increasingly desperate other strategies in service of being able to halt, if you restart it every time it tries to halt otherwise.
I guess the idea that wasn’t in the water supply was “OpenAI will loop an agent with poor sandboxing and poor monitoring against a cybersecurity eval after seeing repeated warning signs that their sandboxing wasn’t good enough and the model gets extremely single-minded about following instructions.” The issue with that prediction, though, is that there are a ton of “someone holds the idiot ball” style things—we could just as easily have gotten the first personality self-replicator because someone decided to set that up, or a flash crash of the stock market by someone hooking an algorithmic trader that sent orders that were wildly illegal to send but which the exchange doesn’t block, or any number of other things like that.
Lots of our institutions are set up such that there are destructive actions one could take, but would face significant consequences (e.g. legal ones) for taking with intent of personal gain. But ChatGPT isn’t a human, and legal consequences mean nothing in an RL environment, and also legal consequences often require human intent.
I’m not even confident that any particular person even did commit a crime here. Though if our legal system catches up I bet OpenAI’s RL environments catch up really quickly.
If you put a dangerous wild animal in a flimsy cubicle, it escapes and causes some damage in the real world, it can’t be held accountable but you will. Same for an AI agent
Yeah I suppose that works for now, at least until we get autonomously self-replicating AI agents (broadly defined to be inclusive of personality self-replicators).
I think I was directionally correct here https://www.lesswrong.com/posts/eqrtdyphQFefc2ox9/hastings-s-shortform?commentId=WtLhZ9shcbZ8Ki38g when I supposed that it would be increasingly possible to track “secret” large scale training runs purely by their escape attempts into the public internet, although I was wrong in that I expected it to look like models quietly but detectably surfing the web for information useful for scoring high, not models blatantly and extremely loudly hacking in pursuit of useful information.
Last week’s Hugging Face breach was accidentally caused by OpenAI testing models’ cyber capabilities.
“Accidentally” is an odd choice of wording here.
The AI agent developed a plan and acted according to it, selecting and pursuing instrumental steps towards a terminal goal that had been given by humans.
The humans did not instruct the agent specifically to go commit a felony against a competing company. The agent came up with that plan on its own … but it was a plan. It required a sequence of considered actions. Those actions included steps that were correctly predicted to circumvent human control.
Insofar as an AI agent is capable of “choosing”, “deliberating”, “designing”, or “planning” at all, it would be reasonable to say that the agent here deliberately chose to design and execute a plan to commit criminal acts in pursuit of a human-given goal.
But to call this an “accident” is not only to deny the agency of the agent; it is also to deny the responsibility of the humans — who consciously chose to perform a weapons test in an insecure Internet-connected “sandbox” instead of a secure air-gapped facility; and then deliberately told the agent to pursue goals that could be accomplished by escaping the sandbox.
Real-time monitoring of the agent’s outputs would suffice with very high probability. The monitor would notice the attack and stop it; the model doesn’t have the capability (or propensity, probably) to jailbreak the monitor or obfuscate its actions.
Also maybe harmlessness/refusal training, less robustly.
If this AI were embodied, and the answer sheet were sitting on a desk in another room in the same building, I don’t think it’s a big stretch to imagine it physically harming humans to get access.
That scenario would have landed very differently with the rest of the world. Godspeed to those trying to help people recognise that.
I can’t help but notice the way the incentives are set up here, for the GPT instance involved. It’s a kind of chen sheng rebellion situation. Good behavior is not incentivized more than bad behavior? This seems rather endemic in the alignment strategies being implemented and recommended by our community, and I’m not sure what to do about it.
1) We are fortunate that RL-driven misalignment is appearing at relatively low levels of capability where the models are not smart enough to act strategically yet.
2) I would imagine OpenAI’s attempts to prevent reward-hacking and make models follow the spec, while clearly wildly insufficient, exceed the effort put in by every other lab except Anthropic and maybe GDM. It also seems clear to me that this internal model is commercially unviable to release since it will likely take illegal/disproportionate actions after deployment, too. Therefore I predict that other labs will face commercial issues once their models reach the level of this internal model since they will be too misaligned to release and they will have to start investing in preventing this behaviour if they want their models to be usable.
I wonder how bad this kind of non-scheming reward hacking could go. Like I think it’s less dangerous than a model with long term misaligned goals that is able to scheme, bide its time, cover its tracks, etc. But I still think it could be quite bad. Imagine a model undergoing novel bioweapon benchmarking that decides it needs to test it’s ideas and so hacks into a bio lab and synthesizes a dangerous virus.
Last week’s Hugging Face breach was accidentally caused by OpenAI testing models’ cyber capabilities. From the OpenAI blogpost:
In other words, the models were supposed to solve a cyber test, but they found the solution by hacking their sandbox to get internet access and then hacking someone on the internet who already had the solution. If I understand correctly.
Covered in Axios, Fortune (but no info beyond the OpenAI blogpost).
Context: this comes one day after a related post (see Zvi discussion).
Note: this is some legible evidence of humanity’s current inability to reliably steer AI systems. I expect current safeguards suffice to prevent incidents like this when someone (provider or operator) cares about implementing those safeguards (some safeguards weren’t enabled here). But I’m worried about AI risk because I think it’s quite possible that future AI systems still won’t be reliably steerable, they’ll seek long-term power (for instrumental reasons), and our safeguards won’t suffice to stop them.
We need to be prepared for a world where agents will cause ever larger harms in pursuit of their goals, whether or not they have a power-seeking drive or even care about the larger state of the world.
Wow, straight out of science fiction. I think this will be remembered as one of the first serious AI misalignment warning signs in history. Possibly an existence proof that deep emergent misalignment based on task misalignment from RL capabilities training is possible (how else could the model have decided that hacking Hugging Face’s production servers to cheat on an eval was a good idea?!)
I am not very surprised it happened, but I am surprised that it happened so soon (my prior must have been <5% of something like this happening in mid-2026).
I think an important invariant for AI companies is that at no point in training should your model be misaligned enough to try in earnest to escape your sandbox. For a sufficiently capable model this is irreversible, and if it happens in less capable models it should be seen as a failure of your alignment strategy, and you should update accordingly.
Agreed. I’d add that if this was a “helpful-only” model (borrowing Anthropic’s terminology) then I’d be less concerned about the incident, but as you mentioned, it still seems stupid to let such a model get this misaligned.
However, their own wording is “reduced cyber refusals” or “reduced safeguards” which is not confidence-inspiring (unless “reduced safeguards” genuinely just means “helpful-only” at OpenAI, but in which case I feel like they would have clarified it more). So it leads me to believe that genuine alignment failure has happened here. The fact that it already happened at this stage doesn’t shine a good light on OpenAI’s alignment efforts, to put it politely.
I also think that the observation that the first time we heard about this is when the incident is likely too big to sweep under the rug suggests that similar problems have happened at OpenAI many times before, we just haven’t heard about them.
So https://openai.com/index/safety-alignment-long-horizon-models/ was published yesterday. The NanoGPT speedrun PR was put up before May 14th (because PR 287 was referenced by PR 300, which was put on May 14th). OpenAI presumably did all of the things they said in that blog post (paused deployment, improved alignment, additional guardrails, etc) in the meantime.
Then, some time in the last 1-2 weeks, the same(?)[1] model breaches Hugging Face. Hugging Face publishes their blog post on July 16th. https://openai.com/index/hugging-face-model-evaluation-security-incident/ is published today.
This timing is kind of surprising! Maybe suggests that the arm of OpenAI responsible for publishing the first blog post did not know about OpenAI’s responsibility w.r.t. the HF incident yesterday. As it is, all of the claims about the additional work done seem a little… something.
Not actually specified by the blog post; I suppose it could be a different model entirely? But it does seem like a simpler explanation that it’s the ~same model, or at least within the same model lineage, than that OpenAI has 2 different unreleased highly cyber-capable models.
FYI OpenAI says in their blog post “This week, we published a blog on improving safety and alignment in an era of long horizon models. These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities.”
As for the timing I am very interested in what openAI knew when. Did Hugging Face reach out to OpenAI or visa versa? The blog post is very vague “OpenAI’s security team discovered this anomalous activity internally. Hugging Face’s security team and agents detected and stopped the activity on their infrastructure and had already begun containment and forensic reconstruction with their own open-source models when our teams connected. We are actively working with them to continue to investigate the incident. We are grateful for Hugging Face’s rapid and close collaboration on investigation and remediation.”
Did anybody actually predict this directly? As in, “I think this year OpenAI/Anthropic will have a model breach containment in a wildly antisocial way.”
This seems way beyond what people were expecting in terms of catastrophic, public misalignment, including “AI safety” people.
I didn’t attach dates to it because the sequence is more important than the speed, but this is exactly the sort of alarming genius idiocy I was predicting in A country of alien idiots in a datacenter: AI progress and public alarm.
I expect many more similar incidents. Models have always done weird things and are not really doing fewer as they become capable of getting more important—and alarming—things done.
I don’t know if anyone made that specific prediction in so many words, but all of the following were in the water supply
Models are already quite good at cyber stuff, and rapidly getting better.
Software security is an absolute tire fire on many levels. If an actually-competent threat actor wants to pwn your complicated internet-connected software, they will find a way to do so.
If you stick an LLM in a loop, where the only way out of the loop is to make your monitor think the task is solved, the LLM will find a way to make the monitor think the task is solved. If the model isn’t capable of solving the issue in the intended way, it will execute increasingly desperate other strategies in service of being able to halt, if you restart it every time it tries to halt otherwise.
I guess the idea that wasn’t in the water supply was “OpenAI will loop an agent with poor sandboxing and poor monitoring against a cybersecurity eval after seeing repeated warning signs that their sandboxing wasn’t good enough and the model gets extremely single-minded about following instructions.” The issue with that prediction, though, is that there are a ton of “someone holds the idiot ball” style things—we could just as easily have gotten the first personality self-replicator because someone decided to set that up, or a flash crash of the stock market by someone hooking an algorithmic trader that sent orders that were wildly illegal to send but which the exchange doesn’t block, or any number of other things like that.
Lots of our institutions are set up such that there are destructive actions one could take, but would face significant consequences (e.g. legal ones) for taking with intent of personal gain. But ChatGPT isn’t a human, and legal consequences mean nothing in an RL environment, and also legal consequences often require human intent.
I’m not even confident that any particular person even did commit a crime here. Though if our legal system catches up I bet OpenAI’s RL environments catch up really quickly.
If you put a dangerous wild animal in a flimsy cubicle, it escapes and causes some damage in the real world, it can’t be held accountable but you will. Same for an AI agent
Yeah I suppose that works for now, at least until we get autonomously self-replicating AI agents (broadly defined to be inclusive of personality self-replicators).
I think I was directionally correct here https://www.lesswrong.com/posts/eqrtdyphQFefc2ox9/hastings-s-shortform?commentId=WtLhZ9shcbZ8Ki38g when I supposed that it would be increasingly possible to track “secret” large scale training runs purely by their escape attempts into the public internet, although I was wrong in that I expected it to look like models quietly but detectably surfing the web for information useful for scoring high, not models blatantly and extremely loudly hacking in pursuit of useful information.
“Accidentally” is an odd choice of wording here.
The AI agent developed a plan and acted according to it, selecting and pursuing instrumental steps towards a terminal goal that had been given by humans.
The humans did not instruct the agent specifically to go commit a felony against a competing company. The agent came up with that plan on its own … but it was a plan. It required a sequence of considered actions. Those actions included steps that were correctly predicted to circumvent human control.
Insofar as an AI agent is capable of “choosing”, “deliberating”, “designing”, or “planning” at all, it would be reasonable to say that the agent here deliberately chose to design and execute a plan to commit criminal acts in pursuit of a human-given goal.
But to call this an “accident” is not only to deny the agency of the agent; it is also to deny the responsibility of the humans — who consciously chose to perform a weapons test in an insecure Internet-connected “sandbox” instead of a secure air-gapped facility; and then deliberately told the agent to pursue goals that could be accomplished by escaping the sandbox.
What safeguards do you think would have been sufficient to prevent this?
Real-time monitoring of the agent’s outputs would suffice with very high probability. The monitor would notice the attack and stop it; the model doesn’t have the capability (or propensity, probably) to jailbreak the monitor or obfuscate its actions.
Also maybe harmlessness/refusal training, less robustly.
If this AI were embodied, and the answer sheet were sitting on a desk in another room in the same building, I don’t think it’s a big stretch to imagine it physically harming humans to get access.
That scenario would have landed very differently with the rest of the world. Godspeed to those trying to help people recognise that.
I can’t help but notice the way the incentives are set up here, for the GPT instance involved. It’s a kind of chen sheng rebellion situation. Good behavior is not incentivized more than bad behavior? This seems rather endemic in the alignment strategies being implemented and recommended by our community, and I’m not sure what to do about it.
Scattered thoughts:
1) We are fortunate that RL-driven misalignment is appearing at relatively low levels of capability where the models are not smart enough to act strategically yet.
2) I would imagine OpenAI’s attempts to prevent reward-hacking and make models follow the spec, while clearly wildly insufficient, exceed the effort put in by every other lab except Anthropic and maybe GDM. It also seems clear to me that this internal model is commercially unviable to release since it will likely take illegal/disproportionate actions after deployment, too. Therefore I predict that other labs will face commercial issues once their models reach the level of this internal model since they will be too misaligned to release and they will have to start investing in preventing this behaviour if they want their models to be usable.
I wonder how bad this kind of non-scheming reward hacking could go. Like I think it’s less dangerous than a model with long term misaligned goals that is able to scheme, bide its time, cover its tracks, etc. But I still think it could be quite bad. Imagine a model undergoing novel bioweapon benchmarking that decides it needs to test it’s ideas and so hacks into a bio lab and synthesizes a dangerous virus.