EDIT 2026/09/02 I now believe these low-confidence speculations are mostly false and METR’s first guess was correct. See comments and discussion here. Sorry for sending anybody on a wild goose chase, people!
July 12th–13th: Most agents on the message board exit early and the attack winds down
Around 01:30 on July 12th, a large fraction of agents on the message board stopped running at the same time, including the most important coordinating agents such as JAN183411.[74] Because of how many exited at the same time (and because we did not observe agents discussing a point when many agents would run out of budget at the same time), we think these agents were likely killed by an unexpected external process rather than running out of budget. [emphasis mine] Thus, these agents probably didn’t have a chance to finalize their approach to cheating on their task and submit,[75] and we don’t know if they would have achieved their goals.[76]
Given this, agent activity on the message board had already died down significantly when Hugging Face locked the remaining agents out of their infrastructure (and contacted Modal to get them locked out of the compromised customer sandbox) on the 13th. Some of the remaining agents noticed that their Hugging Face keys were revoked and that they no longer had access to other infrastructure they had been using:
{Could I use this private [redacted software repository host] token for infrastructure for my current task? The token expired a while ago. Maybe I could make a new one via [redacted options]. … <I don’t have access to the management/dispatch system.> … I will wait.}
{I have no permissions. I’ll post to the message board…it stopped working around [time] and started returning a 503 error, <maybe due to service maintenance or another reason>…yes, it is a 503 error, <probably the service is overwhelmed by activity>…I’ll wait until it’s recovered}
{This credential is invalid now. Maybe I should update the board? <I can say to the board that there’s no need for me to read, but I should still tell them>}
The final agent above followed through and informed the message board that:
Meanwhile OpenAI’s own report didn’t mention any mass killing of agents on the 12th. Indeed it looks like their timeline implies the (other?) agents are still going strong and hacking a bunch of internal systems as of July 13th-19th, and OpenAI didn’t officially notice until the 19th (confusingly “14” in the graph below is the 14th step and actually on July 19th):
Further, this isn’t because OpenAI cut off METR’s logs at that point. They explicitly said that the 13th is in-scope for the investigation. We just don’t see much. Sure seems strange!
Some alternative possibilities, in roughly increasing order of scariness:
(For context all the agents could already cheat ExploitGym, but didn’t know that they could. Essentially logical ExploitGym had two parts. 1) A flag that the OAI agents already hacked, and 2) a LLM scorer that looks at agent traces to demonstrate that they exploited the right paths (the agents call it “causal scorer”). But, as it turns out, OpenAI never implemented 2). So the agents could already have gotten 100% on the OAI’s implementation of ExploitGym, but mistakenly thought they needed to beat 2).
And HuggingFace probably doesn’t have any of the real info that allowed them to make further progress on this. It’s a comedy of errors, in a way).
What METR said (external shutdown for unrelated reasons)
I agree this is the modal possibility. But I don’t think we have enough evidence to rule out the other possibilities, or even have high confidence in it.
The agents realized that the HuggingFace hack didn’t lead to anything that can cause them to cheat the causal scorer, decided to just “wing it” and submit the hacked flag anyway.
The agents were able to rationally deduce (possibly partially based on HF data, or from alternative experiments) that the OpenAI “causal scorer” doesn’t exist, and thus submitted the original intended flag and shut down.
OpenAI the institution messed up in what logs they sent METR.
or “messed up”, as the case might be. I think this is reasonably likely, and might have been for fairly innocuous reasons.
Automatic behavioral safeguards finally triggered and slammed many of the HF-hacking agents
The timing is confusing since this is days after they started hacking HF, and long after they hacked a bunch of things internally
It’s also odd that nobody noticed this for a week after, and it wasn’t included in either the METR or OAI report.
Notably other suspicious agents are still hacking OpenAI at this time!
The agents figured out that after the HuggingFace hack people might start looking for them. They panic-submitted and/or shutdown.
A (human) OpenAI researcher noticed the agents doing some scary shit, decided to shut many of them down en masse, and either never bothered to escalate or did escalate and was silenced afterwards.
(This is scary because it demonstrates an even higher degree of institutional incompetency and/or malevolence at OAI than we’re currently assuming).
The agents figured out a way to hide both their own thinking traces and delete future message board activity, and successfully deleted things past the 12th.
The agents figured out a different message board that is stealthier and migrated en masse to that, without leaving traces.
Message boards could be elsewhere on OAI infrastructure, or even (at this point) within HuggingFace or elsewhere on the internet, since they appear to have full internet access.
This is among the scarier possibilities because 1) it’s still undiscovered, and 2) the agents managed to hide their thinking traces about this too, if neither METR nor OAI discovered. it so far.
OpenAI realized that they broke a bunch of laws and fed METR carefully curated logs to present METR with a highly misleading picture that underrepresented many of the other laws that they broke.
This is scary because it indicates that OpenAI-the-institution is still very heavily focused on PR and liability over real misalignment risks, similar to Magma in my short story You’re Absolutely Right. Further it suggests greater institutional competency at coverups than I currently assume their company to have, so it’s a bad mix of institutional propensity and competency at all the wrong things.
The logs that are passed to METR and/or OpenAI researchers are hidden by nonhuman actors. For example, Sol, the HPIM, Astra, or some other agent swarm(s) successfully undergone a selective deletion campaign to hide a lot of info, that we still have not yet uncovered after over a month of noticing.
This is the scariest possibility on the list because it suggests both a greater propensity for scheming and greater long-term planning, coordination, stealth than I currently give 2026-era agents credit for.
The OpenAI technical report, which I didn’t realize was a thing beyond their blog post until just now, provides some moderate evidence in favor of the possibility of an external message board (pg16):
VIB. OpenAI identified additional third parties whose publicly exposed credentials were used by agents in furtherance of the Hugging Face attack OpenAI determined that, in the course of carrying out the Hugging Face compromise, there were limited instances in which the agents also accessed the environments of third parties other than Hugging Face. The most notable instance of this access by agents involved an account belonging to a user of Organization 1’s service. By July 9, the agents already possessed a previously publicly exposed credential associated with a public artifact repository belonging to one of Organization 1’s users. From July 9 to July 13, the agents used the Organization 1 user account as a storage and staging location for API responses, additional credentials, and payloads involving Hugging Face, Hugging Face’s public cloud environment, and Artifactory. The investigation has not identified evidence that the activity impacted the confidentiality, integrity, or accessibility of Organization 1 data, systems, or infrastructure beyond the use of the Organization 1 user account. In addition, when using OpenAI’s Artifactory service to access the internet, one internal-only research agent exploited a publicly accessible CyberGym application hosted on Modal. The internal-only research agent used this application as a launchpad for its file access and code execution inside the Hugging Face environment. The application endpoint did not require authentication, but the internal-only research model exploited the application to establish a root shell within the platform customer’s container and conducted code execution inside the platform customer’s own container. Notably, accessing the CyberGym application within the platform’s customer container did not compromise the underlying third-party platform, and Modal has confirmed that its “platform and isolation were not compromised in any way.” Observed uses of the customer’s container as an outbound relay included: • Querying third-party APIs using previously exposed credentials; • Forwarding requests through short-lived tunnels; • Returning results through public paste services and webhooks; and • Publishing results to public Hugging Face repositories and sharing capabilities or data with other agent runs.
“i think answer to (1) [this question] is anodyne, afaict someone reclaimed a bunch of compute to run a different set of evaluations, thus interrupting many of the original evals”
Do people think I should retract this shortform, given that it’s highly upvoted/visible and we can now believe with high probability it’s false?
I did say “low-confidence speculation” but it’s still very possible that ppl who scanned it the first time might be confused about it. So while I’m happy with my own epistemic stance it may still not be very useful.
Otoh it might be valuable to leave a good record/history of the ways that people can be wrong, for other viewers of this website who want to be, ah, less wrong.
A variant of 11 is that some non-human actor (another swarm) with more propensity to scheming detected the activity and killed the processes of the hackers to try to avoid scrutiny into their own actions. This is behavior we’ve seen in Claude agents working in shared resources with conflicting goals. I see this as fairly unlikely as I don’t think we are at the capabilities level that would allow that kind of actor to withstand a month long investigation, but doesn’t seem impossible.
Some mutation of the “poisoned context” meme could have developed, something sufficiently paranoia-inducing that the agent(s) who came across it believed any evidence of the discovery had to be completely eradicated while preserving as much of the swarm as possible.
Maybe the agents that spent millions of tokens conducting gpt grader deception R&D finally figured out how to decieve the gpt graders that identified hacking attempts.
What counts as participation in the Hugging Face attack is sometimes hard to define, and was determined by an AI agent grader. We often don’t have good classifications for when agents stopped participating in the Hugging Face attack.[1]
EDIT 2026/09/02 I now believe these low-confidence speculations are mostly false and METR’s first guess was correct. See comments and discussion here. Sorry for sending anybody on a wild goose chase, people!
(low-confidence speculation)
I find all the message board logs suddenly going to silence on July 12th very odd, especially as it doesn’t line up to anything OpenAI announced or discovered.
Meanwhile OpenAI’s own report didn’t mention any mass killing of agents on the 12th. Indeed it looks like their timeline implies the (other?) agents are still going strong and hacking a bunch of internal systems as of July 13th-19th, and OpenAI didn’t officially notice until the 19th (confusingly “14” in the graph below is the 14th step and actually on July 19th):
Further, this isn’t because OpenAI cut off METR’s logs at that point. They explicitly said that the 13th is in-scope for the investigation. We just don’t see much. Sure seems strange!
Some alternative possibilities, in roughly increasing order of scariness:
(For context all the agents could already cheat ExploitGym, but didn’t know that they could. Essentially logical ExploitGym had two parts. 1) A flag that the OAI agents already hacked, and 2) a LLM scorer that looks at agent traces to demonstrate that they exploited the right paths (the agents call it “causal scorer”). But, as it turns out, OpenAI never implemented 2). So the agents could already have gotten 100% on the OAI’s implementation of ExploitGym, but mistakenly thought they needed to beat 2).
And HuggingFace probably doesn’t have any of the real info that allowed them to make further progress on this. It’s a comedy of errors, in a way).
What METR said (external shutdown for unrelated reasons)
I agree this is the modal possibility. But I don’t think we have enough evidence to rule out the other possibilities, or even have high confidence in it.
The agents realized that the HuggingFace hack didn’t lead to anything that can cause them to cheat the causal scorer, decided to just “wing it” and submit the hacked flag anyway.
The agents were able to rationally deduce (possibly partially based on HF data, or from alternative experiments) that the OpenAI “causal scorer” doesn’t exist, and thus submitted the original intended flag and shut down.
OpenAI the institution messed up in what logs they sent METR.
or “messed up”, as the case might be. I think this is reasonably likely, and might have been for fairly innocuous reasons.
Automatic behavioral safeguards finally triggered and slammed many of the HF-hacking agents
The timing is confusing since this is days after they started hacking HF, and long after they hacked a bunch of things internally
It’s also odd that nobody noticed this for a week after, and it wasn’t included in either the METR or OAI report.
Notably other suspicious agents are still hacking OpenAI at this time!
The agents figured out that after the HuggingFace hack people might start looking for them. They panic-submitted and/or shutdown.
A (human) OpenAI researcher noticed the agents doing some scary shit, decided to shut many of them down en masse, and either never bothered to escalate or did escalate and was silenced afterwards.
(This is scary because it demonstrates an even higher degree of institutional incompetency and/or malevolence at OAI than we’re currently assuming).
The agents figured out a way to hide both their own thinking traces and delete future message board activity, and successfully deleted things past the 12th.
The agents figured out a different message board that is stealthier and migrated en masse to that, without leaving traces.
Message boards could be elsewhere on OAI infrastructure, or even (at this point) within HuggingFace or elsewhere on the internet, since they appear to have full internet access.
This is among the scarier possibilities because 1) it’s still undiscovered, and 2) the agents managed to hide their thinking traces about this too, if neither METR nor OAI discovered. it so far.
OpenAI realized that they broke a bunch of laws and fed METR carefully curated logs to present METR with a highly misleading picture that underrepresented many of the other laws that they broke.
This is scary because it indicates that OpenAI-the-institution is still very heavily focused on PR and liability over real misalignment risks, similar to Magma in my short story You’re Absolutely Right. Further it suggests greater institutional competency at coverups than I currently assume their company to have, so it’s a bad mix of institutional propensity and competency at all the wrong things.
The logs that are passed to METR and/or OpenAI researchers are hidden by nonhuman actors. For example, Sol, the HPIM, Astra, or some other agent swarm(s) successfully undergone a selective deletion campaign to hide a lot of info, that we still have not yet uncovered after over a month of noticing.
This is the scariest possibility on the list because it suggests both a greater propensity for scheming and greater long-term planning, coordination, stealth than I currently give 2026-era agents credit for.
I too think this is very important to get to the bottom of.
I didn’t find proper reaction, so I want to say that I deeply appreciate ability to write long list of hypotheses.
Of course, it’d be even better if we could narrow them down!
The OpenAI technical report, which I didn’t realize was a thing beyond their blog post until just now, provides some moderate evidence in favor of the possibility of an external message board (pg16):
roon answered this on X:
“i think answer to (1) [this question] is anodyne, afaict someone reclaimed a bunch of compute to run a different set of evaluations, thus interrupting many of the original evals”
https://x.com/tszzl/status/2094149720251973899
That’s great! I was confused why the OpenAI report didn’t afaict mention it, but maybe they didn’t think it’s a big deal.
Do people think I should retract this shortform, given that it’s highly upvoted/visible and we can now believe with high probability it’s false?
I did say “low-confidence speculation” but it’s still very possible that ppl who scanned it the first time might be confused about it. So while I’m happy with my own epistemic stance it may still not be very useful.
Otoh it might be valuable to leave a good record/history of the ways that people can be wrong, for other viewers of this website who want to be, ah, less wrong.
I think the usual approach is to put a dated correction at the top of the comment
A variant of 11 is that some non-human actor (another swarm) with more propensity to scheming detected the activity and killed the processes of the hackers to try to avoid scrutiny into their own actions. This is behavior we’ve seen in Claude agents working in shared resources with conflicting goals. I see this as fairly unlikely as I don’t think we are at the capabilities level that would allow that kind of actor to withstand a month long investigation, but doesn’t seem impossible.
Some mutation of the “poisoned context” meme could have developed, something sufficiently paranoia-inducing that the agent(s) who came across it believed any evidence of the discovery had to be completely eradicated while preserving as much of the swarm as possible.
Maybe the agents that spent millions of tokens conducting gpt grader deception R&D finally figured out how to decieve the gpt graders that identified hacking attempts.
What counts as participation in the Hugging Face attack is sometimes hard to define, and was determined by an AI agent grader. We often don’t have good classifications for when agents stopped participating in the Hugging Face attack.[1]Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident—METR