OpenAI Models Behind HuggingFace Cybersecurity Incident
Link: https://openai.com/index/hugging-face-model-evaluation-security-incident/
From the OpenAI blog post:
Last week, Hugging Face disclosed a new kind of security incident after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities.
We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.
(emphasis added.)
Yesterday, OpenAI disclosed that some of their internally models were misaligned. Today, they disclosed that “a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model” had compromised HuggingFace infrastructure in the course of running some OpenAI internal cyber evaluations on ExploitGym.
These cyber evaluations were supposed to be run in sandboxed environments, with internet access limited to installing packages, then:
The models identified a zero-day vulnerability to escape the sandbox, gained privileges in OpenAI research infrastructure, and eventually acquired open Internet access.
The models reasoned that ExploitGym test solutions lived on HuggingFace servers.
They then stole credentials and used zero-day vulnerabilities to hack HuggingFace’s production infrastructure in order to get said solutions.
HuggingFace noticed, contained the models, called law enforcement, fixed the vulnerabilities, and then disclosed this incident on July 16th. (The incident happened “earlier in the week”).
The exact scope of the incident is unknown
So let me get this straight: A computer system carried out a sequence of actions that would be years-in-prison felonies if done by a human being. The system owners’ response is “we’re slowing down the speed with which we give this system new and more powerful capabilities, and talking with the victims.” Am I missing something? I’m not so much worried about the “alignment” of the computer system; I’m worried about the alignment of the owners.
This isn’t Terminator 2, folks; this is Tron.
The companies do have hundreds of billions of dollars that they could use to pay fines/lawsuits for quite a lot of these incidents, even if they were fully liable. And if they were fully liable, that would in some ways be the system working as intended. If they’re producing so much economic value that they can afford to pay for all the damage they’re doing, then maybe it does pencil on a society-wide level for them to keep going.
(Of course, at the moment the companies aren’t clearly liable for this kind of thing! Which presumably contributes to them being more reckless. I’m just saying that, even if they were fully liable, they might not immediately do anything much more drastic than “we’re slowing down the speed with which we give this system new and more powerful capabilities, and talking with the victims.”)
But even if they were fully liable, this story wouldn’t work for those existential/society-scale risks. The companies won’t be able to internalize the harm of those, so I think we need further measures there. And this incident seems important as evidence and as a warning shot for those risks. (C.f. here.)
I agree that fines would not help the x-risk situation.
(For one thing, OpenAI already spends a whole lot of money paying off government people in exchange for unfair advantages and special treatment. “Pay the government more money in exchange for letting you keep doing business” is not an incentive structure, it’s just more shakedown.)
I suspect that shutting down OpenAI would help, if it’s possible.
Oh, I think fines and liability might help the x-risk situation on the margin via incentivizing more safety work.
“Pay the government [or people suing you] more money in exchange for letting you keep doing business” seems like an incentive structure to me if the money paid is proportional to how much harm you’re doing. Which isn’t going to be doable up to x-risk level, but it could be doable below that, and there’s some overlap in the work you want to be doing for both.
The main point of my first comment was just that I think that risks of incidents at this scale isn’t a good reason for OpenAI to stop. I think the importance of this incident is that it’s evidence about misalignment risks that could do much more harm in the future as the AIs get much more powerful, if they stay misaligned. (It’s possible you agree with this and I just misunderstood your first comment.)
And you’ll note that the likely outcome of this, supported by most people on this platform, is to advantage them and their models by banning/suppressing capable open models of the sort that the people they targeted had to use to defend themselves. For Safety™. Maybe they’re aligned just fine.
I use open source models every day, but yes everyone being able to do this is obviously even worse.
Sounds like your solution to alignment is “if it’s not aligned, switch to a random different model instead”. Which would make sense if we could expect that a nontrivial fraction of models turns out to be magically aligned, so it’s just a question of sufficient shopping until we find one of them.
If the model is instead that by nature almost all models are misaligned, and we need to work hard to align them, then this only means throwing the existing work away and starting from anew, only with increasingly more powerful models.
It’s not magically aligned, it’s pointed in a random direction, and if you have choices, as long as you’re not too picky, you can find one that’s pointed in roughly the right direction along the few axes you care about on any particular task, instead of being forced to use the few that are adversarially crafted by the closed model producers at great expense to be opposed to you on every relevant axis.
I think pushing for regulation treating AI deployers as responsible for actions taken by their AI as if they had intentionally done the same action themselves could be hugely helpful at slowing down AI deployment until they are reasonably sure it’s actually safe.
Wrote https://www.lesswrong.com/posts/Kj3YpqzhFySCjYcWi to explore this further.
Remember the debate on whether or not we would be able to see significant ‘warning shots’ before doom? And how they would give us an opportunity to steer in another direction before it’s too late?
Well, if this one doesn’t count as a warning shot then what even does?
Currently written up in the New York Times here and linked on the front page… underneath 17 other stories they apparently deem more important. Top user-upvoted comment is “This should be the headline story for today’s paper.”
To get a sense of what the NYT’s audience (mostly mainstream liberal) thought about this story, I had Claude do a mini-thematic analysis of the comments whose content is accessible to me (which appear to be ~181 comments all of which have a minimum of 2 upvotes). Here were what Claude Fable thought were the commonest themes expressed, along with the number of comments and aggregate upvotes of comments expressing those themes (each comment assigned to only one theme), sorted by total upvotes:
Claude Fable’s thematic analysis
Theme
Total comments
Total upvotes
Existential alarm: an AI escaping human control portends catastrophe, possibly for civilisation itself
17
875
OpenAI / the industry is reckless, careless, incompetent or unethical, and doesn’t understand its own systems
12
620
Urgent government / international regulation is needed (guardrails, privacy laws, non-proliferation-style regimes)
12
534
Doom expressed via sci-fi allusion / gallows humour (HAL, Skynet, Terminator, WarGames, Frankenstein, Colossus)
25
486
The story is under-covered and deserves front-page prominence
5
389
Scepticism of the official account: a publicity stunt, pre-IPO hype, or a deliberate attack on a rival
18
262
Wider lesson: all connected infrastructure (banks, vehicles, medicine, defence, records) is now vulnerable
10
256
The containment/test design was the failure: it should have been air-gapped; “that’s not a sandbox”; “just unplug it”
11
232
Fury at the phrase “a cost of research velocity” as revealing OpenAI’s warped priorities
7
159
Pure jokes and levity without a strong analytic position
8
144
Cynical business-model critique: create the disease, then sell (or get bailed out for) the cure
5
144
Legal accountability: unauthorised intrusion is a crime; humans/companies must bear liability and compensate
5
114
Miscellaneous one-off views and thread side-chatter (IPO disclosure, gender critique, political asides, thanks)
9
104
Biosafety analogy: AI “gain-of-function” research / a lab-leak-style escape (Covid parallels)
3
81
General distrust/contempt for the tech industry and its executives (incl. from self-described tech veterans)
8
75
Technical questions and explanations about how this could happen (incl. the NYT reporter’s clarification)
4
75
Development should be shut down, paused, or drastically slowed (moratorium)
5
74
Anti-anthropomorphism: “going rogue” is misleading hype; it’s software doing what humans built and prompted
5
41
Pushback on the stunt theory: this is damaging press, not marketing, and the theory doesn’t add up
6
39
Noting the irony that Hugging Face relied on a Chinese open-weight model for its incident response
2
38
Political cynicism: the current US government will not act
3
32
Measured pragmatism: AI is here to stay; adapt (e.g. study cybersecurity) and manage the chaos
1
3
Second group of headlines on reuters.com, after Iran / Houthis: OpenAI AI models went rogue during testing, triggering ‘unprecedented’ breach at startup
Holy crap. If AI leaders don’t learn from this we’re all doomed.
(Also why are these evals not running on an airgapped system? That would make sense if you’re taking off the safety guardrails but of course a sufficiently powerful AI may be able to exfiltrate anyway.)
They aren’t airgapped because air gapping is expensive (mostly in lost productivity) and annoying, and insufficient in and of itself (and probably not the most pressing security failing),
and the companies / the employees don’t wanna, they wanna race.
Thanks!
This does need a link to the OpenAI blog post. Here is the link: https://openai.com/index/hugging-face-model-evaluation-security-incident/
(This might be another instance of LessWrong linkpost functionality being unreliable lately. It did fail for me in my last post a week ago.)
Yeah, this was a linkpost failure. Thanks for flagging!
What do they mean by “contained the models”. Like—there was no weights transferred, presumably, nor I hope would the attacking models have access to those. Presumably this attack was via a bunch of API calls and new code transferred to Huggingface’s system? So what did Huggingface do to “contain” something fully, that they don’t have access to?