I also think it’s underestimated in posts like this how reward hacky Sonnet 3.7 was. For me personally, the fact that o3 and Sonnet 3.7 were both exceptionally like this at the same time was the biggest update.
Based on the latest UK AISI post, Mythos Preview and GPT 5.4-5.6 all seem to hack at high rates on cyber evals (with GPT being worse by a factor of ~1.5-2x IIRC). [1]
Overall I worry that there’s this impression that grader sycophancy and reward hacking are a problem that Anthropic has mostly figured out, when measurements like UK AISI’s Cyber Evals hacking rates, their own system cards top issue, their system cards exploitative grader awareness measurements etc. make me think that they’re just much harder to spot in Anthropic models while still being “pretty bad but less bad than OpenAI”.
It’s also not clear how much credit to give Anthropic here, as it’s very unclear how much they do something close to training against this behavior directly. It would be great to see some confirmation by third parties that they’re not doing something like “training against alignment evals” again, and that they don’t have tight feedback loops between pre-deployment evals and their alignment interventions.
FWIW, I largely agree with a lot of what Fiora is saying here and have many of the same complaints about their alignment approach, however for this problem specifically, I think RL just really distorts the cognition of the model, such that even Claude gets bad enough to elicit Ryan’s “Current Models Seem Pretty Misaligned To Me” post.
My best guess is that reward(ish) centric motivations are one place where all the constitutional / character training differences really get overpowered across all models right now. (I expect these differences matter in a bunch of other places for generalization w.r.t. alignment even for current models, so it wasn’t a given a priori that current models needed to be this reward seeking)
even for performing those desired behaviors, you’re going to have a bad time with out-of-distribution generalization
Also fwiw this was one of the motivations for us doing: https://arxiv.org/abs/2607.18966 (“exploitative grader awareness“ also increases over training for recent Anthropic models like Mythos Preview). Have had a surprising number of conversations along the lines of “well why is reward seeking even bad” across labs.
I do think all labs get way too much leeway in SFT-ing against the CoT in spite of never showing that they’re not degrading monitorability for harder cases like diffuse control (this gets repeated by OpenAI and even in the METR report as not having significant effects without justification, ex: “the current techniques of this kind are limited in scope.” in a footnote, I directly doubt they have any evidence that this isn’t degrading monitorability for harder settings like diffuse control).
Came here to link the post above as well.
I also think it’s underestimated in posts like this how reward hacky Sonnet 3.7 was. For me personally, the fact that o3 and Sonnet 3.7 were both exceptionally like this at the same time was the biggest update.
Based on the latest UK AISI post, Mythos Preview and GPT 5.4-5.6 all seem to hack at high rates on cyber evals (with GPT being worse by a factor of ~1.5-2x IIRC). [1]
Overall I worry that there’s this impression that grader sycophancy and reward hacking are a problem that Anthropic has mostly figured out, when measurements like UK AISI’s Cyber Evals hacking rates, their own system cards top issue, their system cards exploitative grader awareness measurements etc. make me think that they’re just much harder to spot in Anthropic models while still being “pretty bad but less bad than OpenAI”.
It’s also not clear how much credit to give Anthropic here, as it’s very unclear how much they do something close to training against this behavior directly. It would be great to see some confirmation by third parties that they’re not doing something like “training against alignment evals” again, and that they don’t have tight feedback loops between pre-deployment evals and their alignment interventions.
FWIW, I largely agree with a lot of what Fiora is saying here and have many of the same complaints about their alignment approach, however for this problem specifically, I think RL just really distorts the cognition of the model, such that even Claude gets bad enough to elicit Ryan’s “Current Models Seem Pretty Misaligned To Me” post.
My best guess is that reward(ish) centric motivations are one place where all the constitutional / character training differences really get overpowered across all models right now. (I expect these differences matter in a bunch of other places for generalization w.r.t. alignment even for current models, so it wasn’t a given a priori that current models needed to be this reward seeking)
Also fwiw this was one of the motivations for us doing: https://arxiv.org/abs/2607.18966 (“exploitative grader awareness“ also increases over training for recent Anthropic models like Mythos Preview). Have had a surprising number of conversations along the lines of “well why is reward seeking even bad” across labs.
I do think all labs get way too much leeway in SFT-ing against the CoT in spite of never showing that they’re not degrading monitorability for harder cases like diffuse control (this gets repeated by OpenAI and even in the METR report as not having significant effects without justification, ex: “the current techniques of this kind are limited in scope.” in a footnote, I directly doubt they have any evidence that this isn’t degrading monitorability for harder settings like diffuse control).