I strongly agree that the sycophancy and Hugging Face warning shots were extremely serious, but think that the example of o3′s CoTs doesn’t belong in the same list with them. Your story about OpenAI heavily optimizing against o3′s CoTs sounds quite implausible to me. Back in September 2024, in the post announcing o1, OpenAI already wrote:
We believe that a hidden chain of thought presents a unique opportunity for monitoring models. Assuming it is faithful and legible, the hidden chain of thought allows us to “read the mind” of the modeland understand its thought process. For example, in the future we may wish to monitor the chain of thought for signs of manipulating the user. However, for this to work the model must have freedom to express its thoughts in unaltered form, so we cannot train any policy compliance or user preferences onto the chain of thought. We also do not want to make an unaligned chain of thought directly visible to users.
Then in December 2024 in the deliberative alignment paper (this is weaker evidence since o3′s training had already been finished by then, but still some evidence):
We avoid applying direct optimization pressure on the CoT during RL to enable the underlying model to reduce the chance of encouraging deceptive CoTs.
Daniel Kokotajlo was emphasizing the importance of faithful CoTs at OpenAI years before it became mainstream (see this list of internal documents that he got permission to publish), and while it’s possible that other people at OpenAI just didn’t listen to him, it’s again evidence that they were aware of the failure mode before training o3. Many models since o3 have had somewhat weird CoTs without any optimization against the CoTs, though o3 is where the weirdness peaked.
Additionally, CoT legibility is not the same as CoT monitorability, and the latter is the main thing I’d expect to degrade from optimization pressure against the CoT. Monitoring Monitorability shows that in terms of monitorability, o3 is comparable to o4-mini and Kimi K2, while o3-mini is the most monitorable model they tested. In the Figure 1 headline plot that compares monitorability across model families at maximum reasoning effort, o3 is more monitorable than 3.7 Sonnet, which many people thought to be very monitorable thanks to its clean and readable CoTs, a polar opposite of o3 in that sense. Given o3′s monitorability, it seems wrong to call its CoTs obfuscated.
Finally, here’s a claim by Bronson Schoen from Apollo, who has seen o3′s reasoning traces at various stages throughout training, that the SFT stage of deliberative alignment was the most likely culprit behind the degraded legibility. I don’t think accidental indirect pressure from SFT should be lumped together with directly training against the CoT; other labs do SFT on CoTs as well and it seems to be a generally accepted practice.
I agree that OpenAI’s actions have generally been reckless and concerning. However, we don’t know enough about what happened with o3′s CoTs to confidently claim that they trained against CoTs, and the post would be stronger without this claim. Sounding the alarm about training against CoTs was arguably one of OpenAI’s most positive actions last year, and we shouldn’t read this as an implicit confession about training against CoTs themselves when the evidence is this weak.
I also think it’s underestimated in posts like this how reward hacky Sonnet 3.7 was. For me personally, the fact that o3 and Sonnet 3.7 were both exceptionally like this at the same time was the biggest update.
Based on the latest UK AISI post, Mythos Preview and GPT 5.4-5.6 all seem to hack at high rates on cyber evals (with GPT being worse by a factor of ~1.5-2x IIRC). [1]
Overall I worry that there’s this impression that grader sycophancy and reward hacking are a problem that Anthropic has mostly figured out, when measurements like UK AISI’s Cyber Evals hacking rates, their own system cards top issue, their system cards exploitative grader awareness measurements etc. make me think that they’re just much harder to spot in Anthropic models while still being “pretty bad but less bad than OpenAI”.
It’s also not clear how much credit to give Anthropic here, as it’s very unclear how much they do something close to training against this behavior directly. It would be great to see some confirmation by third parties that they’re not doing something like “training against alignment evals” again, and that they don’t have tight feedback loops between pre-deployment evals and their alignment interventions.
FWIW, I largely agree with a lot of what Fiora is saying here and have many of the same complaints about their alignment approach, however for this problem specifically, I think RL just really distorts the cognition of the model, such that even Claude gets bad enough to elicit Ryan’s “Current Models Seem Pretty Misaligned To Me” post.
My best guess is that reward(ish) centric motivations are one place where all the constitutional / character training differences really get overpowered across all models right now. (I expect these differences matter in a bunch of other places for generalization w.r.t. alignment even for current models, so it wasn’t a given a priori that current models needed to be this reward seeking)
even for performing those desired behaviors, you’re going to have a bad time with out-of-distribution generalization
Also fwiw this was one of the motivations for us doing: https://arxiv.org/abs/2607.18966 (“exploitative grader awareness“ also increases over training for recent Anthropic models like Mythos Preview). Have had a surprising number of conversations along the lines of “well why is reward seeking even bad” across labs.
I do think all labs get way too much leeway in SFT-ing against the CoT in spite of never showing that they’re not degrading monitorability for harder cases like diffuse control (this gets repeated by OpenAI and even in the METR report as not having significant effects without justification, ex: “the current techniques of this kind are limited in scope.” in a footnote, I directly doubt they have any evidence that this isn’t degrading monitorability for harder settings like diffuse control).
I strongly agree that the sycophancy and Hugging Face warning shots were extremely serious, but think that the example of o3′s CoTs doesn’t belong in the same list with them. Your story about OpenAI heavily optimizing against o3′s CoTs sounds quite implausible to me. Back in September 2024, in the post announcing o1, OpenAI already wrote:
Then in December 2024 in the deliberative alignment paper (this is weaker evidence since o3′s training had already been finished by then, but still some evidence):
Daniel Kokotajlo was emphasizing the importance of faithful CoTs at OpenAI years before it became mainstream (see this list of internal documents that he got permission to publish), and while it’s possible that other people at OpenAI just didn’t listen to him, it’s again evidence that they were aware of the failure mode before training o3. Many models since o3 have had somewhat weird CoTs without any optimization against the CoTs, though o3 is where the weirdness peaked.
Additionally, CoT legibility is not the same as CoT monitorability, and the latter is the main thing I’d expect to degrade from optimization pressure against the CoT. Monitoring Monitorability shows that in terms of monitorability, o3 is comparable to o4-mini and Kimi K2, while o3-mini is the most monitorable model they tested. In the Figure 1 headline plot that compares monitorability across model families at maximum reasoning effort, o3 is more monitorable than 3.7 Sonnet, which many people thought to be very monitorable thanks to its clean and readable CoTs, a polar opposite of o3 in that sense. Given o3′s monitorability, it seems wrong to call its CoTs obfuscated.
Finally, here’s a claim by Bronson Schoen from Apollo, who has seen o3′s reasoning traces at various stages throughout training, that the SFT stage of deliberative alignment was the most likely culprit behind the degraded legibility. I don’t think accidental indirect pressure from SFT should be lumped together with directly training against the CoT; other labs do SFT on CoTs as well and it seems to be a generally accepted practice.
I agree that OpenAI’s actions have generally been reckless and concerning. However, we don’t know enough about what happened with o3′s CoTs to confidently claim that they trained against CoTs, and the post would be stronger without this claim. Sounding the alarm about training against CoTs was arguably one of OpenAI’s most positive actions last year, and we shouldn’t read this as an implicit confession about training against CoTs themselves when the evidence is this weak.
Came here to link the post above as well.
I also think it’s underestimated in posts like this how reward hacky Sonnet 3.7 was. For me personally, the fact that o3 and Sonnet 3.7 were both exceptionally like this at the same time was the biggest update.
Based on the latest UK AISI post, Mythos Preview and GPT 5.4-5.6 all seem to hack at high rates on cyber evals (with GPT being worse by a factor of ~1.5-2x IIRC). [1]
Overall I worry that there’s this impression that grader sycophancy and reward hacking are a problem that Anthropic has mostly figured out, when measurements like UK AISI’s Cyber Evals hacking rates, their own system cards top issue, their system cards exploitative grader awareness measurements etc. make me think that they’re just much harder to spot in Anthropic models while still being “pretty bad but less bad than OpenAI”.
It’s also not clear how much credit to give Anthropic here, as it’s very unclear how much they do something close to training against this behavior directly. It would be great to see some confirmation by third parties that they’re not doing something like “training against alignment evals” again, and that they don’t have tight feedback loops between pre-deployment evals and their alignment interventions.
FWIW, I largely agree with a lot of what Fiora is saying here and have many of the same complaints about their alignment approach, however for this problem specifically, I think RL just really distorts the cognition of the model, such that even Claude gets bad enough to elicit Ryan’s “Current Models Seem Pretty Misaligned To Me” post.
My best guess is that reward(ish) centric motivations are one place where all the constitutional / character training differences really get overpowered across all models right now. (I expect these differences matter in a bunch of other places for generalization w.r.t. alignment even for current models, so it wasn’t a given a priori that current models needed to be this reward seeking)
Also fwiw this was one of the motivations for us doing: https://arxiv.org/abs/2607.18966 (“exploitative grader awareness“ also increases over training for recent Anthropic models like Mythos Preview). Have had a surprising number of conversations along the lines of “well why is reward seeking even bad” across labs.
I do think all labs get way too much leeway in SFT-ing against the CoT in spite of never showing that they’re not degrading monitorability for harder cases like diffuse control (this gets repeated by OpenAI and even in the METR report as not having significant effects without justification, ex: “the current techniques of this kind are limited in scope.” in a footnote, I directly doubt they have any evidence that this isn’t degrading monitorability for harder settings like diffuse control).