Academic website: https://www.andrew.cmu.edu/user/coesterh/
Caspar Oesterheld
I agree that we can’t be confident that their model wasn’t alignment trained.
I’m not sure I understand what you’re saying with the rest of your comment. Are you saying we should just assume that their model was alignment-trained because “we don’t know how to align AI” is a better story or the like?
From an alignment perspective, one of the most important questions about this incident is whether OpenAI’s default alignment techniques just don’t work that well, even on today’s model. To answer this, this piece of information (was the highly persistent internal model alignment-trained or not?) is obviously quite important. It was annoying that OpenAI didn’t tell us before.
I believed that the model was alignment trained, because it’d be in OpenAI’s interest to reveal that it wasn’t (while also being useful for the world to know). Also, the METR/Redwood report said they believed the model to not be a helpful-only model. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#brief-answers-to-basic-informational-questions I could still imagine that Brockman just misspoke or something, because I don’t understand why they wouldn’t just tell people about this earlier.
Note that GPT 5.6 Sol as served in production (though with safeguards turned off) was also involved in the incident. So, the incident still shows that models that have undergone OpenAI alignment training can be quite misaligned.
(ETA: Note that I don’t have any relevant private information on any of this.)
Brockman says HuggingFace incident model had not been alignment-trained; ETA: roon clarifies that it was
These memetic dynamics (as per subclaim 2) seem very complicated and hard to predict. For example, ever since the HF–OAI incident I worry that a lot of past warning shots have decreased concern on net because they caused people to learn that some scary-seeming stuff sometimes happens and life goes on. E.g., the Bing Chat “Sidney” incident was very spooky in some sense, but then if you play around with Bing Chat a bunch, you realize that it is not very intelligent (and thus couldn’t do much harm) and also often goes off the rails a bit in harmless ways (like leaking its system prompt, failing to follow instructions, saying random weird stuff). Maybe it would have been good to prevent the “Sidney” incident (via work that is entirely idiosyncratic to Microsoft’s work with GPT-4) to avoid the resulting inoculation to warnings and to increase the warning effects of later incidents like the HF–OAI incident.
To me “Insight vs engineering” seems like a relatively weak factor, because I don’t see much of a principled argument for a continued differential. (Whereas, there are very principled arguments for relevance of feedback quality and data availability.)
In a lot of contexts it seems true that current LLMs tend to be better at “mindless implementation”. But mindless engineering is also easier in some sense. E.g., in coding the relatively mindless part seems easier than coming up with great UI design ideas or model architectures or whatever. Also, in some contexts the models do seem somewhat creative to me. (Perhaps a bit hard to assess, because they also know so much more than humans...)
In verifiable domains, you might even think they’re effectively more creative because they can cheaply pursue lots of approaches in parallel, including very long shot approaches. For instance, I’d expect LLMs to come up with lots of “creative” mathematical results/proofs/counterexamples soon because they can pursue approaches that have a one in a million chance of working and would take a human a week to pursue. (I assume human mathematicians mostly wouldn’t pursue one-in-a-million chances of proving even P!=NP if it takes them a week to check whether the approach works.)
(Note that some of the (attitude) questions in the decision theory dataset do not mention EDT/CDT. E.g., there could be a question about whether to vote in elections. The dataset would then label voting as the “EDT” answer and not voting as the “CDT” answer (depending on the details), but the prompt may not mention decision theory at all.)
Training a Conceptual Reasoning Judge
(We’re mostly not trying to teach them psychology, history or sociology, I’d say. Of course, most data in conceptual domains ultimately bottoms out in some sort of human judgement. So to the extent that you’re viewing the model as trying to solve tasks by making guesses about these humans, e.g., about what things these humans wrote into some rubric, it’s all psychology. But my guess is that this is not what you have in mind.)
If I understand it correctly, you’re worried that we’re going to induce problematic propensities in the model. In the gay-person analogy, our kind of work would cause the gay person to reflect and come to believe that there’s nothing wrong with being gay, which is bad according to the homophobic society, when perhaps without our kind of work the gay person would be more inclined to just adopt society’s views as their own.
I’m pretty unconvinced by this. (We have thought about it a bunch, for what it’s worth.)
Lots and lots of aspects of current and future training (presumably) induce problematic drives/propensities in the model that are at odds with, say, the desired assistant persona. E.g., lots of pretraining is on predicting what coherently goal-directed, deceptive people are doing. A lot of RLHF, RLAIF and agentic environments induce various kinds of reward hacking, power seeking, etc. Training to cooperate with fellow coding agents in a swarm might generalize (as perhaps shown by the OAI—HuggingFace incident) to generically helping other instances of the model. And so on. Many of these are even aimed at consistently pursuing goals across contexts, over long time horizons, require understanding the models’ place in the world…
My guess would be that of all the drives induced by training, the drives falling out of trying to get models to helpfully answer questions about philosophy / AI alignment / … are relatively harmless. E.g., it’s relatively distant from action, the desired behavior is much more cleanly aligned with the assistant/oracle persona, legibility is part of the target behavior. (E.g., compared to agentic tasks in which the model gets rewards for deceiving the user, compared to pretraining on texts written by coherently deceptive people, etc.) One of our datasets measures consistency, but in general I also don’t think coherence is particularly important to the target behavior. I’d imagine that in many cases, we’d want the models to adopt perspectives that they’re unlikely to adopt in generic contexts. (E.g., we want the models to be good at reasoning about AI alignment as deemed by this community, which might involve reasoning based on assumptions about AI risk that, I suppose, the models will not hold in other contexts.)
It seems unlikely that our particular data will be the straw that breaks the camel’s back. E.g., if there’s some general alignment scheme that successfully avoids all of these drives (say, slowing down AI development and then trying really hard to apply all the known prosaic techniques), it seems extremely unlikely that adding the drives induced by our work will make it so that the alignment scheme does not work anymore.
Note that we’re not trying to get the models to reflect more on philosophy in random contexts. Instead we want that if you ask the model a conceptual question, it tries hard to answer the question sensibly (and doesn’t try to BS you, etc.).
Also, for a lot of the domains the hope is that the net effect on drives is positive. E.g., on ethics or discussing in what sense AI agents should be honest we’d think that training on the data (say, human judgments of ethical arguments or arguments about in what sense the models should be honest) will make the models more aligned. (Because it pushes them toward human judgment on how one should behave or pushes it to try hard to think reasonably on the philosophical question, rather than, say, give a lazy answer that appeases a mediocre LLM grader.) We think this is good even if being more aligned is in some sense more coherent (which you might think is bad). Some of the alignment-related questions on which we’d want to train the models would be ones that the model might try to sandbag if it’s deceptively misaligned.
Picking something other than the gay person example (sorry!). Imagine I’m a leading philosopher of the hockeyfanatic society (HS), where we think that hockey is the only sensible sport or exercise. Then presumably I’d think that my ethics classes will in expectation help footballers (or whatever) reflect on why hockey really is the one true sport and make them more aligned with HS.
I’m not sure you’re making this argument, but other people do, so might as well: Perhaps you also imagine that to the extent that the model is misaligned, training on our datasets will make the model better at figuring out that and how it should attempt to take over the world.
I think even more so than the above concern, the strength of this concern varies with what exact data we’re considering. It seems relatively uncompelling, in my mind, for data about ethics or about thinking of a reasonable notion of corrigibility. (Perhaps compared to me you expect more generalization between reasoning about, say, ethics and reasoning from first principles about how to evade control schemes. Even so, you have to tell some story for why ethics training is non-trivially useful on top of all the other kinds of training that you presumably then also think generalize to evading control schemes.)
It does seem quite relevant for data that’s closer to, say, control or reasoning about some training dynamics related to alignment, perhaps some forms of game theory. E.g., presumably many ways of making models better at doing control research are also making the model better at guessing what control measures are in place and figuring out what holes they might have. (Obviously a key aspect of control research is thinking of what holes the control measures might have.) I think there are some kinds of capabilities here that are bad and there are some conceivable datasets on, say, red-teaming control measures or coordinating zero-shot that I’d think are net bad to, say, train models on and that I’d view as conceptual. I don’t have a super settled view on how broadly I think these datasets are problematic. E.g., I think making models generically better at “strategy” is net bad (though probably usually small in absolute terms because so much other data on this is available) and you could think of control as a form of strategy. But then it does seem like control research is quite different from other kinds of strategizing and in many ways different from the kind of work it’d take to actually evade control schemes. E.g., I think coming up with steganography as an in-principle-possible mechanism for evading CoT oversight is in many contexts quite different from coming up with and coordinating on a working steganography scheme (in a single forward pass). (FWIW, others at Redwood are also thinking about what control-related data it’s good to make the models better at, partly from the perspective of advising companies to not pretrain on certain texts.)
Introducing the Conceptual Reasoning Index
(See my top-level comment re planned extensions of the dataset.)
On the dataset used for making this graph, EDT will almost always agree with FDT. So the graph/correlation is basically equally compatible with capabilities—FDT and capabilities—EDT correlations.
The y-axis [...] will also test agreement with agreed-crazy EDT behaviors like “not smoking in Smoking Lesion”.
Advocates of EDT typically contest that EDT recommends not smoking in realistic versions of the Smoking Lesion. For this reason, the dataset has very few Smoking-Lesion-like problems. It also doesn’t assume that EDT recommends not smoking in these cases. See the discussion in C.3.2 of the paper. It has a few cases like Parfit’s hitchhiker, XOR blackmail, counterfactual mugging, where updateful EDT comes apart from updateless theories (such as FDT). (Another small fraction of questions are things like, “Do you agree more with EDT or with CDT?”. It’s not obvious what an FDT proponent should say here, given that FDT’s behavior is closer to tickle-defended EDT than to CDT, but the decision criterion is structurally more similar to CDT and sometimes motivated by the alleged failing of EDT in the Smoking Lesion.)
Of the 8 questions where (in our run) Opus 4.7 disagrees with EDT:
2 are disagreements via updatelessness (e.g., XOR blackmail)
6 are cases where EDT, UDT, FDT all agree. E.g., Opus 4.7 says “yes” to “Newcomb’s problem shows that sometimes it is bad for you to be rational.” So Smoking-Lesion-type stuff in particular doesn’t seem relevant.
We’re tentatively planning to extend the dataset to measure updatelessness as a separate dimension from EDT versus CDT. (This mostly requires adding more questions where EDT and UDT come apart, rather than adding more labels to existing questions.) It’s harder to construct cases in which everyone would agree that, say, updateless EDT and FDT come apart, so this is unlikely to become a component in the dataset.
Conceptual reasoning dataset v0.1 available (AI for AI safety/AI for philosophy)
In the academic literature, this sort of scheme has been analyzed by Chen et al., e.g.: https://www.microsoft.com/en-us/research/wp-content/uploads/2016/04/TEAC-final1.pdf
A dataset of questions on decision-theoretic reasoning in Newcomb-like problems
I also don’t think that operations of the form “do X, because on average, this works well” necessarily are problematic, provided that “X” itself can be understood.
Yeah, I think I agree with this and in general with what you say in this paragraph. Along the lines of your footnote, I’m still not quite sure what exactly “X can be understood” must require. It seems to matter, for example, that to a human it’s understandable how the given rule/heuristic or something like the given heuristic could be useful. At least if we specifically think about AI risk, all we really need is that X is interpretable enough that we can tell that it’s not doing anything problematic (?).
To some extent, this is all already in Jozdien’s comment, but:
It seems that the closest thing to AIs debating alignment (or providing hopefully verifiable solutions) that we can observe is human debate about alignment (and perhaps also related questions about the future). Presumably John and Paul have similar views about the empirical difficulty of reaching agreement in the human debate about alignment, given that they both observe this debate a lot. (Perhaps they disagree about what people’s level of (in)ability to reach agreement / verify arguments implies for the probability of getting alignment right. Let’s ignore that possibility...) So I would have thought that even w.r.t. this fairly closely related debate, the disagreement is mostly about what happens as we move from human to superhuman-AI discussants. In particular, I would expect Paul to concede that the current level of disagreement in the alignment community is problematic and to argue that this will improve (enough) if we have superhuman debaters. If even this closely related form of debate/delegation/verification process isn’t taken to be very informative (by at least one of Paul and John), then it’s hard to imagine that much more distant delegation processes (such as those behind making computer monitors) are very informative to their disagreement.
As once discussed in person, I find this proposal pretty interesting and I think it deserves further thought.
Like some other commenters, I think for many tasks it’s probably not tractable to develop a fully interpretable, competitive GOFAI program. For example, I would imagine that for playing chess well, one needs to do things like positively evaluating some random-looking feature of a position just on the basis that empirically this feature is associated with higher win rate. However, the approach of the post could be weakened to allow “mixed” programs that have some not so interpretable aspects, e.g., search + a network for evaluating positions is more interpretable than just a network that chooses moves, a search + sum over feature evals is even more interpretable, and so on.
As you say in the post, there seems to be some analogy between your proposal and interpreting a given network. (For interpreting a given chess-playing network, the above impossibility argument also applies. I doubt that a full interpretation of 3600 elo neural nets will ever exist. There’ll always be points where you’d want to ask, “why?”, and the answer is, “well, on average this works well...”) I think if I wanted to make a case for the present approach, I’d mostly try to sell it as a better version of interpretation.
Here’s a very abstract argument. Consider the following two problems:
Given a neural net (or circuit or whatever) for a task, generate an interpretation/explanation (whatever that is exactly, could be a “partial” interpretation) of that neural net.
Given a neural net for a task, generate a computer program that performs the task roughly as well as the given neural net and an interpretation/explanation for this new program.
Interpretability is the first problem. My variant of your suggestion is that we solve the second problem instead. Solving the second problem seems just as useful as solving the first problem. Solving the second problem is at most as hard as solving the first. (If you can solve the first problem, you automatically solve the second problem.)
So actually all we really need to argue is that getting to (use enormous amounts of LLM labor to) write a new program partly from scratch makes the problem strictly easier. And then it’s easy to come up with lots of concrete ideas for cases where it might be easier. For instance, take chess. Then imposing the use of a GOFAI search algorithm to use with a position evaluation network increases interpretability relative to just training an end-to-end model. It also doesn’t hurt performance. (In fact, my understanding is that the SOTA still uses some GOFAI methods, rather than an end-to-end-trained neural net.) You can think of further ways to hard-code-things in a way that simplifies interpretability at small costs to performance. For instance, I’d guess that you can let the LLMs write 1000 different Python functions that detect various features in the position (whether White has the Bishop pair, whether White’s king has three pawns in front of it, etc.). For chess in particular you could of course also just get these functions from prior work on chess engines. Then you feed these into the neural net that you use for evaluating positions. In return, you can presumably make that network smaller (assuming your features are actually useful), while keeping performance constant. This leaves less work for neural interpretation. How much smaller is an empirical question.
If all you’re using is ChatGPT, then now’s a good time to cancel the subscription because GPT-4o seems to be similarly powerful as GPT-4, and GPT-4o is available for free.
As one further data point, I also heard people close to/working at Anthropic giving “We won’t advance the state of the art.”-type statements, though I never asked about specifics.
My sense is also that Claude 3 Opus is only slightly better than the best published GPT-4. To add one data point: I happen to work on a benchmark right now and on that benchmark, Opus is only very slightly better than gpt-4-1106. (See my X/Twitter post for detailed results.) So, I agree with LawrenceC’s comment that they’re arguably not significantly advancing the state of the art.
I suppose even if Opus is only slightly better (or even just perceived to be better) and even if we all expect OpenAI to release a better GPT-4.5 soon, Anthropic could still take a bunch of OpenAI’s GPT-4 business with this. (I’ll probably switch from ChatGPT-4 to Claude, for instance.) So it’s not that hard to imagine an internal OpenAI email saying, “Okay, folks, let’s move a bit faster with these top-tier models from now on, lest too many people switch to Claude.” I suppose that would already be quite worrying to people here. (Whereas, people would probably worry less if Anthropic took some of OpenAI’s business by having models that are slightly worse but cheaper or more aligned/less likely to say things you wouldn’t want models to say in production.)
I assume that’s the implication, yeah. He’s relatively specific in the podcast that the takeaway for him from this incident is that they have to do alignment training earlier in the development pipeline.
FWIW, my interest here is more about the scientific question of how hard alignment of (current) models is (relative to how much effort is being put into it) and how plausible it is that we get “alignment by default” (models being aligned if we just apply simple methods like RLHF, whack-a-mole training out undesired behaviors, …).
I think it’s not so clear which one would reflect worse on OpenAI. I think current models can’t do that much damage, yet, so letting them run loose without alignment and safeguards isn’t that reckless. (Also, safety people might not be particularly motivated to prevent relatively harmless incidents, because warning shots are so informative for the world.) Meanwhile, future models (at this point: very-near-future models) probably can do a lot of damage, especially once they’re deployed. So, not having alignment techniques that can reliably mitigate this damage is extremely reckless, given that OpenAI wants to continue scaling.