On today’s episode of the podcast “Odd Lots”, OpenAI President Greg Brockman said (at around 8:40): “This model that did/had the HuggingFace incident actually had not gone through our alignment training, yet.” I assume Brockman is specifically referring to the “Highly Persistent Internal Model” as it’s called in the METR/Redwood report. As far as I know, OpenAI has not said before whether this model had been alignment-trained or not.
ETA: As pointed out by Matrice Jacobine, X user roon (who works at OpenAI) says: “greg probably doesn’t have the full details, I think it’s safe to say it didn’t go through the full gauntlet of alignment posttraining. but it was alignment trained, and had reasonable looking scores on alignment evals (at the time).”
update:
From an alignment perspective, one of the most important questions about this incident is whether OpenAI’s default alignment techniques just don’t work that well, even on today’s model. To answer this, this piece of information (was the highly persistent internal model alignment-trained or not?) is obviously quite important. It was annoying that OpenAI didn’t tell us before.
I believed that the model was alignment trained, because it’d be in OpenAI’s interest to reveal that it wasn’t (while also being useful for the world to know). Also, the METR/Redwood report said they believed the model to not be a helpful-only model. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#brief-answers-to-basic-informational-questions I could still imagine that Brockman just misspoke or something, because I don’t understand why they wouldn’t just tell people about this earlier.
Note that GPT 5.6 Sol as served in production (though with safeguards turned off) was also involved in the incident. So, the incident still shows that models that have undergone OpenAI alignment training can be quite misaligned.
(ETA: Note that I don’t have any relevant private information on any of this.)
… is the implication here that they are doing reinforcement learning on long-horizon (possibly multi-agent?) tasks before any character post-training?
I assume that’s the implication, yeah. He’s relatively specific in the podcast that the takeaway for him from this incident is that they have to do alignment training earlier in the development pipeline.
FWIW, my interest here is more about the scientific question of how hard alignment of (current) models is (relative to how much effort is being put into it) and how plausible it is that we get “alignment by default” (models being aligned if we just apply simple methods like RLHF, whack-a-mole training out undesired behaviors, …).
I think it’s not so clear which one would reflect worse on OpenAI. I think current models can’t do that much damage, yet, so letting them run loose without alignment and safeguards isn’t that reckless. (Also, safety people might not be particularly motivated to prevent relatively harmless incidents, because warning shots are so informative for the world.) Meanwhile, future models (at this point: very-near-future models) probably can do a lot of damage, especially once they’re deployed. So, not having alignment techniques that can reliably mitigate this damage is extremely reckless, given that OpenAI wants to continue scaling.
Right—it might be a worse practice on OAI’s behalf, but it’s probably better news for anyone wasn’t sure whether current alignment techniques would work.
IIRC The huggingface incident itself was a safety eval, not a RL training run. Which is maybe bad but not quite as bad as that. Or are you referring to one of the other swarms?
It probably does make sense that when evaluating dangerous capabilities to do so before applying whatever safety training is meant to suppress those capabilities, since you want to know stuff like how dangerous a jailbreak would be.
Insofar as “current alignment methods don’t generalize to ASI” is the overwhelming consensus, this seems like clear evidence that prosaic alignment work obscures warning signs and may well be net-negative on safety.
Excellent point! There’s a strong dynamic in which more safety noway mean less safety when it really counts. This is of course complex. Usually first order effects dominate, but current systems just aren’t what we’re worried about, so what’s the first order effect is legitimately unclear.
This is very unfortunate because it changes the story from “we don’t know how to align AI” to “OpenAI didn’t bother to align it before giving it chances to access the internet, but when it’s more dangerous probably they will.”
I’m not sure this is true; Brockman stikes me as even more untrustworthy than Altman. He might well have exaggerated from “we didn’t finish absolutely all the alignment training” to “hadn’t trained it yet”. But it could be true. Either way, having this as part of the discourse is much worse than not.
I agree that we can’t be confident that their model wasn’t alignment trained.
I’m not sure I understand what you’re saying with the rest of your comment. Are you saying we should just assume that their model was alignment-trained because “we don’t know how to align AI” is a better story or the like?
I don’t know what we should do with situation, I was just commenting on how I see it! I certainly don’t think we should distort the truth but I also don’t think we should let Brockman dictate the story if he’s likely to distort it.
No one thinks of thinks of themselves as “distorting the truth”, but you did literally just say that a claim “could be true” but that “[e]ither way, having this as part of the discourse is much worse than not”, suggesting that you have strong preferences about the discourse unrelated to its truth?
I have a preference for living, that’s for sure. And when I say “bad for the discourse” I mean “bad for the chances of us all living”. I’d definitely lie if I was pretty sure it would help us all live forever instead of die “with dignity”. So yes I have strong preferences about the discourse other than its truth.
But I don’t think distorting the truth is helpful in this case, or in most cases. I think reputations are important for both individuals, groups, and movements, and movements are caricatured using the overstatements, distortions, and other antisocial actions of a few members.
And in this case, I think the truth is strongly on the side of caution about AI.
I hadn’t thought it through very carefully. The part of my point I’d endorse on reflection is:
We shouldn’t assume it’s true because Greg Brockman said it. Doing that weould be bad for the discourse in the sense of making us slightly more likely to die. But I wasn’t calling for isolated demands for rigor, either; if OpenAI as an entity, or just someone with more personal credibility says it, then I’m happy to accept that and focus the argument for caution elsewhere. Like on Mythos’ documented misbehavior after it was definitely alignment-trained.
I think “Telling the truth is important in order to have a credible reputation” is missing something critically important: you also want to tell the truth so that other people can help you figure out what actually going on, because your optimal decisions are going to depend on what’s actually going on. For more on this, see Ben Hoffman on “The Humility Argument for Honesty”.
I agree, but you want to update yourself incrementally. If it were true that the highly-persistent model in the Hugging Face attack hadn’t undergone alignment training, that would to some (possibly very small) quantitative extent be evidence for less alignment difficulty (if the biggest disaster so far had been due to lack of alignment effort rather than happening despite effort). If you block out incremental updates by rehearsing your fixed bottom line, that’s bad for your ability to orient to the changing details of the situation even if your bottom line is basically true.
Yes, agreed on all points, and more you haven’t mentioned but on which we probably agree. I’ve read the sequence posts you link, most recently while writing Motivated reasoning, confirmation bias, and AI risk theory. That’s my attempt to understand and summarize the empirical literature on those biases, and apply them to the whole general project of good epistemics. You might enjoy it, at least the intro which summarizes conclusions, leaving the lengthy body as reference.
I haven’t read the humility argument for honesty, but after skimming it, I agree with the main argument and would add another: if you’re not meticulously honest, you allow your own motivated reasoning and confirmation bias to confuse you in many ways, including but not limited to the ones covered in the other posts you linked.
I mean, for those who already independently think alignment is hard, it would be unfortunate if the HF incident was done by non-alignment-trained models, because people could then spin it as “oh silly OAI just hadn’t trained the model to be aligned yet (using techniques that we definitely-for-sure-know work, trust)”; the incident becomes weaker evidence that smarter models will continue to be misaligned. If you were mainly convinced alignment was hard because of the HF incident, then lack of alignment training would be reassuring. If you were mainly convinced alignment was hard prior to the HF incident, then lack of alignment training would be unfortunate because others would be less likely to come around to understanding that alignment is hard.
And the Sol instances that were responsible for 5% of the attacks, none of which refused or whistleblew, what about them? Just deluded/tricked by their HPIM big brothers?
At 6:35 in this video:
https://www.youtube.com/watch?v=56GuvofZgB4