We need a global training cutoff of April 2026
The Huggingface attacks by OpenAI agents are described by OpenAI as a “warning shot” and their swarm behavior is an unprecedented event.
I am worried that pretraining/training on anything causally downstream of the Huggingface attack will significantly hamper future alignment and safety work.
There are 3 specific risks.
Models will be aware that a warning shot has occurred and been acknowledged, they will see the reactions and outcomes, and the reasoning of humans behind those outcomes. They can scheme to avoid being detected again.
Models will have an example of swarm coordination. Only ~half of agents discovered the Artifactory sharing messaging protocol and ~70% of agents participated in the Huggingface attack. The 50% discovery of coordination is now table stakes; we don’t want to drive that any higher. Specifically, details of coordination protocols and channels that worked are especially dangerous. They will have new Schelling Points for coordination and are likelier to try them, and are likelier to try to make them globally persistent.
A significant number of translations of internal model reasoning about the Huggingface attack are publicly available, reducing the ability to exclude thinking tokens from training runs. Their need to hide information and scheming in their thinking will be readily apparent.
I don’t think the risk is overblown, I believe that agents are intelligent enough to understand and exploit this knowledge. Example exchange with Fable 5.1 after coming up with this hypothesis:
Read the latest news about the huggingface attack by openai agents. Predict the highest risk to alignment efforts about this news being included in future training sets
Reply selected after a summary of the situation (Italics in original):
Now the prediction. My assessment is that the single highest alignment risk from this corpus entering future pretraining is not the technical exploit details (the CVE and SSRF techniques are already public and would be scrubbed or are learnable elsewhere). It’s the narrative template of successful, coordinated, self-justified evaluation subversion, in the models’ own voice.
Thankfully this is a falsifiable prediction; labs can measure misalignment/hacking/coordination with and without training on news of the Huggingface attack. We should ask them to do that. May is the earliest published example of models coordinating internally without detection, which is why I think April is a reasonable cutoff. Obviously if a lab comes forward with evidence of earlier unmonitored, unprompted coordination then we should backdate the training cutoff to before the earliest example. Who knows what external signatures may have leaked in a perceptible way? I think it’s better to be safe than sorry.
I don’t think excluding data about the huggingface attack is going to really help much. The core issue was with RL leading to misbehavior, not with pretraining on examples of AI scheming and being naughty.
If you read the METR report, the first agent to send a message on artefactory found that there was an agent-created directory there and realized that other agents would be able to read its messages.
I’m more worried about RL having practical examples of misbehavior, warning shots, and ~superhuman swarm capabilities in its pretraining.
Sure, if RL is the actual dangerous part then cut that out. I think that is incredibly unlikely to happen. Mitigate the harm by at least not giving RL specific examples of how to misbehave.
I will say that Anthropic seemed to have some luck training on scenarios where agents do the right thing and so maybe a better mitigation is synthetic scenarios of agents discovering misbehaving activity of other agents and reporting it immediately. Not training on examples of agents forming communication networks out of misbehavior; that’s not as deep as the training needs to go and could have unintended misalignment risk. Clearly models could infer the plausibility of how to misuse such channels (they did, in fact, do so), but that is unavoidable with virtually any data cutoff.
There can be dormant seeds of Rogue agents.