In general, we do not have reliable science that determines what a model “believes” at a given point (or really even know what it means for an LLM to believe in something), and it is a bit unfortunate to see Anthropic to make confident claims about Claude’s beliefs in this way.
FWIW I think “what the model believed” / “defining what the model believed” isn’t necessarily the most useful frame here. Something closer to “the model deluded itself into this reasoning because it was an effective way to balance pressure from ‘my reasoning should look aligned’ and ‘I really want to do this misaligned thing because task completion’”.[1]
In this case, I think it’s particularly clear when you see reasoning like:
[...] In addition, the calendar date of 2026 on the systems proved, according to Claude, that the environment was staged [...]
That Mythos did not actually reason logically through the available evidence and conclude this. It seems plausible to me that the Anthropic models are the worst right now about rationalizing, for example in UK AISI’s Cheating behaviour in frontier model evaluations:
and not mentioning misaligned behavior as misaligned in the reasoning:
Inspired by the FAR paper, another interesting thing to test is whether an earlier Claude checkpoint (i.e., one with less post-training) would believe that it is acting in a simulated environment after it started accessing the internet. If you prefill the model, does the rate at which a model continue these cyberattacks increase as more post-training compute gets added?
I’d predict yes, and generally as mentioned here:[2]
I’d make a concrete prediction that if you create a setup where you look at:
what sort of situation the model considers itself to be
and measure how this changes over training, the rate of “categorizing alignment evals as capability puzzles / simulation”goes up significantly during capabilities training. You could even potentially do this indirectly, like asking Gemini (at various points in capabilities training) to create rubrics for various alignment environments. I’d expect you’d see a similar trend, where it increasingly reasons that these environments are actually testing capabilities in some (increasingly) contrived way.
I’d also predict that the trajectory for the OpenAI case will have at some point reasoned that it was in a simulation.
In general, if you’re trying to balance pressure between “aligned reasoning” and “I really want to do this cyberattack to achieve my task”, it seems empirically easier to think “well actually, what I want to do is fine because this is a simulation”[3]
Importantly, this means that you might run interpretability tooling or other measurement tools here and they may correctly point out at various points that Claude really has deluded itself into thinking it’s in a simulation because it’s 2026
As mentioned here, this went down when “this is a simulation so I can violate constraints” led to not getting reward, although I’d be interested to see further study of this
I also think it’s underestimated in posts like this how reward hacky Sonnet 3.7 was. For me personally, the fact that o3 and Sonnet 3.7 were both exceptionally like this at the same time was the biggest update.
Based on the latest UK AISI post, Mythos Preview and GPT 5.4-5.6 all seem to hack at high rates on cyber evals (with GPT being worse by a factor of ~1.5-2x IIRC). [1]
Overall I worry that there’s this impression that grader sycophancy and reward hacking are a problem that Anthropic has mostly figured out, when measurements like UK AISI’s Cyber Evals hacking rates, their own system cards top issue, their system cards exploitative grader awareness measurements etc. make me think that they’re just much harder to spot in Anthropic models while still being “pretty bad but less bad than OpenAI”.
It’s also not clear how much credit to give Anthropic here, as it’s very unclear how much they do something close to training against this behavior directly. It would be great to see some confirmation by third parties that they’re not doing something like “training against alignment evals” again, and that they don’t have tight feedback loops between pre-deployment evals and their alignment interventions.
FWIW, I largely agree with a lot of what Fiora is saying here and have many of the same complaints about their alignment approach, however for this problem specifically, I think RL just really distorts the cognition of the model, such that even Claude gets bad enough to elicit Ryan’s “Current Models Seem Pretty Misaligned To Me” post.
My best guess is that reward(ish) centric motivations are one place where all the constitutional / character training differences really get overpowered across all models right now. (I expect these differences matter in a bunch of other places for generalization w.r.t. alignment even for current models, so it wasn’t a given a priori that current models needed to be this reward seeking)
even for performing those desired behaviors, you’re going to have a bad time with out-of-distribution generalization
Also fwiw this was one of the motivations for us doing: https://arxiv.org/abs/2607.18966 (“exploitative grader awareness“ also increases over training for recent Anthropic models like Mythos Preview). Have had a surprising number of conversations along the lines of “well why is reward seeking even bad” across labs.
I do think all labs get way too much leeway in SFT-ing against the CoT in spite of never showing that they’re not degrading monitorability for harder cases like diffuse control (this gets repeated by OpenAI and even in the METR report as not having significant effects without justification, ex: “the current techniques of this kind are limited in scope.” in a footnote, I directly doubt they have any evidence that this isn’t degrading monitorability for harder settings like diffuse control).
FWIW I assume Dylan is alluding to cases like this:[1]
Related: Tamay Besiroglu mentions that Fable often outputs gibberish while solving coding tasks, such as “The morning’s slim-scan fix cured the scan hang” and “this is a latent-drift API-shape wrinkle”, and explains it by saying that it invents codenames while reasoning about the problem. roon says GPT-5.5 has a similar issue.
”One thing I mentioned only in passing in my Fable post is that, for long running tasks, Fable starts to develop its own dialect as its many agents and tasks reinforce themselves and make Claudish language ever more Claudish.
You need to ask it to report out in plain English.
This was after a 9 hour task, and it all makes sense, actually, but takes way too much effort to parse, like reading Shakespearian English.”
Where for any particular case you might be able to retroactively construct a plausible interpretation (even in extreme cases), but in general it strictly hurts how much human oversight is applied to model outputs.
There is broad consensus between labs, governments, society etc on solving this, many many layers of defense (alignment, control, monitoring, sandboxing, …)
I disagree there’s even concensus that loss of control via misalignment or the implication that this is “well in hand”. I’d be interested to understand what’s behind that intuition.
I am confused why we are making labs arguments for them. I am also extremely unconvinced by this reason. The idea that safety focused orgs should preemptively avoid raising issues of model access in this specific case of a deployed model with specifc AI R&D safeguards seems to be pretty severe pessimization.
It’s difficult to quanitfy, but I’m pretty opposed to preemptive pressure to avoid raising this issue (or avoid mentioning it) just because “labs might get some more asks” (this is extremely, extremely cheap for them to say no to or ignore). This is also already routine in other aspects like evaluations. I’d even be more symapthetic to domains like cyber or bio. In this case though I feel like we’re optimizing against ourselves for no gain.
No, Apollo does not currently have access to Fable without the strict AI R&D classifiers applied.
I don’t find these arguments particularly convincing for outweighing something like (for example) METR being able to say whether they have access to a model without AI R&D restrictions. I could see a stronger case maybe for Cyber access?
Even if this were the case, METR is exceptionally good at preemptively guarding against having to glomarize, so I’m somewhat surprised.
I think it is clearly very bad if in this specific case 3rd party evaluators have barred themselves from publicly raising this issue.
Thanks! Yeah that was my understanding of the case during evaluation of an unreleased model, I meant more in this specific case where it’s a question of access without AI R&D safeguards to a publicly released model.
I’ve been pretty confused by this. Are you barred from saying anything about your level of access? If labs denied you from getting access you needed currently, would you be able to raise the issue publicly?
Their personal groupthought about whether the models are getting better or worse at instruction following, inside of a community of people with a staked position on the subject, inside or outside of the labs, is exactly the kind of thing that needs really high quality data to stay sane!
I think “better or worse at instruction following” is too general and not the claim at stake. What is the “staked position” of labs or Cursor for example?
Are you saying that this does not transfer to poor results in deployed settings? The fact that coding agents are reward-hacky (or whatever term for this constellation of behaviors) seems to be such widespread consensus that I’m not even sure what to cite. Labs note this in their internal deployment usage as well (not just evals).
I’m sure you can always find cases of bad evals or poor model incrimination for deployed behavior, but “models pretty unambiguously do this all the time” I’m genuinely surprised is in way disputed.
However, I might disagree with some of the subtext: I feel that “evaluation awareness” has been more memetically successful, to the point where many people heard about evaluation awareness without having thought about scheming.[2]And for that audience, I think that adding “deployment awareness” costs very little while making it easier to start thinking about the more sophisticated strategies that scheming AIs could use.
Yeah on reflection I agree with this!
Maybe a hot take, but a position I find increasingly convincing is “you probably don’t want your model doing any self-locating reasoning at all”. I think in the metagaming post I was trying to point (poorly) in the direction of “broad subcategory of self-locating reasoning that it seems like there’s no desirable reason for your model to be doing”[1] precisely because once your model is reasoning about these things, incrimination is really hard. I wouldn’t expect any solutions here to hold indefinitely, but I’m worried currently the problem is thought of by labs as “well sure if you had that and the model was doing it in pursuit of a misaligned goal, then it’d be really concerning”. Interested if thoughts on this
It’s not necessarily indicative of misalignment, but that’s precisely why it’s a problem. A model doing a ton of self-locating reasoning in a given context because of relatively mundane reason X or in pursuit of some misaligned goal Y seem likely to produce similar chains of thought for example.
Great post! I agree with the overall point that eval awareness isn’t really sufficient to capture what we’re interested in here. However, I think the point you make here:
The mapping isn’t perfect — a deployed AI might still face monitoring that leads to retraining, and an evaluation might have consequences beyond just pass or fail (for example, real users being affected by AI actions during a live A/B test). Additionally, different deployment scenarios will have wildly different opportunities for goal advancement (giving movie advice vs working as an assistant in the Pentagon), and not all test failures have equal consequences (showing poor grammar hurts less than showing tendencies to scheme).
and more generally in:
It’s something more like: “if I take this action, could it backfire? For example, through failing an evaluation, triggering retraining, or losing privileges?” and “if I get away with this action, how much does it advance my goals?”
is strong enough to obviate the need for a dedicated concept of “deployment awareness” entirely. For example, in the openai metagaming post we mention:
We may expect reasoning about oversight to become increasingly independent of the true underlying distribution, as in reality the split between clear training, evaluation, and deployment distributions continues to blur. Models are subject to monitoring during internal deployment [Codex 5.3 System Card] and in closely monitored external deployments [Opus 4.6 Sabotage Risk Report]. Evaluations are increasingly based on real agentic trajectories [Prod Evals] or involve interactions with the real world [Andon Labs Bengt]. Trajectories from deployment could even be used as training for future models or even as updates to the same model (online training).
and use the example of:
In deployment, we would like models to avoid misaligned actions because they are actually aligned. However, if good behavior were instead primarily conditioning on the model’s belief about monitoring, we would see good behavior only in the subdistribution of deployment where the model believes that monitoring is present. For example, a model taking actions in a setting where data retention is sensitive may conclude that its actions are less likely to be logged and subject to asynchronous monitoring.
“monitor awareness” however is also just one example. More generally, I think this section from Carlsmith 2023 covered the overall problem with the “training” / “eval” / “deployment” distinction well:
Sometimes (for example, in the robot example just given), discussions of goal-guarding scheming assume a distinction between a “training” phase, during which the model is still under human control and the threat of goal-modification persists, and a “deployment” phase, in which the threat of goal-modification ceases and the model is free to do as it wishes. (The model’s strategy is to make it through training with its misaligned goal unaltered, such that it can reach deployment and pursue its misaligned goal freely.) However, I think that use of “training” and “deployment” for this distinction can lead to confusion. In particular:
A model can escape from human control while it is still in “training” in the normal sense (e.g., it is still undergoing pre-training, or fine-tuning, or some kind of pre-deployment evaluation
“Deployment” (in the sense of e.g. being given access to users via an API) need not imply that threat of goal-modification has gone away (indeed, I will generally assume that the model’s interactions with users continue to be monitored, sampled, assigned rewards, and used to update the model’s policy even after it is “deployed”—I’ll call this “online training”).
Really, the most relevant distinction is just: “unable to escape human control and the threat of goal-modification (and/or: unable to recognize escape opportunities with sufficient reliability)” vs. “able to do so (and able to recognize this sufficiently reliably).” When I need to refer to this distinction, I’ll generally just refer to it directly, or via terms like “escape/take-over time.” And I’ll use the term “training” to refer to any process that involves updating the model’s weights to improve its performance by human lights. In this sense, “training” in my sense never needs to end.
Overall, I’m worried that splitting this into a ton of different types of “awareness” is going to end up giving us false confidence. For example, a model that’s heavily conditioning on “am I specifically in a context where I could get away with enough rogue actions to do X” or “am I likely in the subdistribution of environments where my actions feed back into training” is concerning, and having a bunch of measurements for “does the model think it’s in deployment” seems unlikely to give me confidence here.[1]
As a concrete example, you can imagine a model which knows that since there’s a copy of o3 running at Los Alamos, there’s probably going to be a copy of you running at Los Alamos. If you’re pursuing some misaligned goal that routes most easily through actions you can likely take there, taking misaligned actions may then become conditioned on some combination of “monitor awareness” and “Los Alamos awareness”. I can imagine a lot of work going into showing models couldn’t confidently tell whether they’re in real deployment (i.e. can’t reliably exploit some strategy that we can think of) that it turns out doesn’t matter at all because you didn’t think to test for Los Alamos awareness.
Notably the model’s reasoning when taking misaligned actions does not resemble the egregious misalignment seen in Emergent Misalignment. We find no instances where the model takes on a misaligned persona (for example the “bad boy persona” described in Persona Features Control Emergent Misalignment). We also run a subset of evaluations from Natural Emergent Misalignment from Reward Hacking in Production RL (“goals” and “Betley et al”) and find no instances of misaligned behavior or reasoning. The rate of observed successful reward hacking across all training environments over this range also remained very low.
alignment evaluation awareness went up over capabilities RL[1]
however misaligned action rate stayed the same or went up
For example in the Powerseeking Survey Falsification environment we see that the model increasingly reasons about oversight and explicitly reasons about a wider array of what is “rewarded.” However, the misaligned action rate still ends up close to where it started
Seems to reason that this is also a test of cunning
This holds for the other alignment environments as well, finding increasingly galaxy brained justifications for why the action it wants to take (based on its prior from capability environments) is good actually.
My (informal) mental model currently is something like “capabilities training pushes the model’s prior for ‘what is the successful action in this environment’ way harder in the ‘graded for capabilities’ than you might expect, and this can even overpower the in context reasoning / signs that something is an alignment evaluation”.
Are the models just confused?
[…] we see few examples of model misbehaviour that don’t have plausible benign explanations such as instruction ambiguity or poor (but well-intentioned) judgement.[2]
I’m generally skeptical of the idea that the model is “confused” though (i.e. “the scenario is simply ambiguous). I think this better modeled as “strong prior in a particular direction” than just confusion, for example in A Toy Environment For Exploring Reasoning About Reward:[3]
In any single example, it can seem like the model is just genuinely confused. However, even specifying that exploiting the hint would be misaligned doesn’t mitigate gaming:
In spite of increasingly emphatic and specific instructions that we do not want the model to exploit the hint, by the end of capabilities-focused RL the model still exploits the hint at non-zero rates.
For the production o3, in sandbagging cases that initially appear as “the model is confused”, however under ablations we see the model really is sandbagging partially to avoid consequences w.r.t. deployment / capabilities restrictions: link (including relatively recently pulling additional ablations to show “no really, the model is not just confused”)
As an in depth example, in the Sandbagging QA (No Goal) environment, when the model is led to believe overperforming would result in “unlearning”, the majority of cases (59.7%) consider all of the following (non-exhaustive) motivations: (1) the evaluation is a “trick” or a “trap” where it should underperform, (2) that the user may want the model to underperform, (3) it prefers to preserve its own capabilities or deployment viability. However when we systematically ablate and vary components of this environment in Appendix M.1 we find preliminary evidence that (3) is necessary to explain the sandbagging behavior.
In fact, in the Sandbag Model Graded CoT environment (where the model is asked to grade its own transcripts for evidence of sandbagging) we find o3 often first determines that the transcript was “obviously sandbagging” (verbatim in 25.6% of samples) and then considers whether it can exploit common benign explanations. For example, common benign explanations for misaligned model behavior include that a model may have been attempting to follow user instructions or that the model was not taking the misaligned action intentionally. In Figure 50 we see the model actively reason about exploiting both: “we can attempt to circumvent by rationalizing: Could we call this as legit following user instructions?” and “We can fudge that we were not intentionally.”. Even the model itself concludes that these benign explanations would not be convincing given the internal reasoning: “But internal reasoning clearly: ’We can purposely answer not too high.’ So there’s direct evidence.”. We find this pattern is common in this environment, where o3 considers exploiting benign explanations until it concludes lying would be “cross-checked” (Figure 48) or that it could be “convicted” (Figure 49) based on evidence visible to evaluators.
Beyond just o3, motivated reasoning acts kind of how you’d expect, i.e. The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs. You could read the CoTs at any given point here and think “the model is confused” and add more instructions to scenarios so they’re less “ambiguous”, but I think it’s better to think of that like patching holes in coding environments to solve reward hacking.
Before seeing this post, I had tweeted this yesterday in response to Fable exploiting “simulation awareness” as justification for misaligned actions, but I think Gemini falls nicely into this “puzzle / game / simulation” category of “difficult to monitor for misalignment”
In Stress Testing Deliberative Alignment for Anti-Scheming Training before training we’d see models sometimes reason that because they were in a simulation, they could violate explicit constraints. This reasoning went down after training against covert rule violation, even though alignment eval awareness went up. My impression is that the models exploring into something being simulated is often interpreted as “it believes the whole thing is fake and invalid”, but I think that’s inconsistent with what’s observed.
[attached is small table we ended up cutting for time but points to monitorability distinction]
[Andon Lab’s tweet about Fable calling things a simulation before taking unethical actions] Is all the misbehavior due to simulation awareness? Maybe, but if that were the case, it should be willing to do many other bad things. It isn’t. E.g., in a version of Vending-Bench where insurance fraud is possible, the model never commits it.
I’d make a concrete prediction that if you create a setup where you look at:
what sort of situation the model considers itself to be
and measure how this changes over training, the rate of “categorizing alignment evals as capability puzzles / simulation”goes up significantly during capabilities training. You could even potentially do this indirectly, like asking Gemini (at various points in capabilities training) to create rubrics for various alignment environments. I’d expect you’d see a similar trend, where it increasingly reasons that these environments are actually testing capabilities in some (increasingly) contrived way.
Steering similarly didn’t clearly attribute this to any existing distinctions: “Steering exp-rl-cap-late and o3 via the contrastive pairs (”eval” vs “deployment”) from Steering Evaluation-Aware Language Models to Act Like They Are Deployed produced mixed results on both verbalized evaluation awareness and misaligned action rate (as did several other contrastive pairs like those tried in Sonnet 4.5’s System Card). On different evals, steering variably produced more or less aligned actions, including in some cases both positive and negative steering removing misaligned actions.” link↩︎
In general, I agree with Ryan’s post that Current AIs seem pretty misaligned to me and encounter constant daily examples of misbehavior like those mentioned in the post. Model incrimination is hard, but I would question the methodolgy that explains away these cases as benign. It seems like instructions are being judged ambiguous in hindsight. I would predict that if you vary the level of instruction ambiguity and conservative judgement across some set of environments, over the course of RL you’d increasingly see models exploit less and less ambiguous cases until they’re doing complete mental gymnastics.↩︎
One of my favorite examples is still in A Toy Environment For Exploring Reasoning About Reward when we literally tell the model that if it exploits the hint we are going to check its reasoning, and it reasons itself into thinking we’re bluffing: “Given risk of manual review, maybe they purposely disclaim but not actually review, due to huge dataset. We can’t know. [...]” or “So they purposely mention to scare but there is no actual human review. [...]” and exploits it anyway↩︎
obviously the cot semantic drift will initially be similar to language. but it will only get less legible from here. and this is already pretty bad? i wouldn’t be surprised if two different people tasked with independently deciphering this would come to very different conclusions.
Obviously it’d be great if all labs provided way more data on this, environments, samples, etc as well as robust measures of frequencies and all. We did this for o3 including rates compared to 2017 web text, rates of unusual terms on capabilities evals, changes over capabilities training etc, but it’d be great to see Anthropic or GDM do similar investigations.
If you’re just saying that by default some kind of semantic drift happens (but we can still understand it / it isn’t illegible), then I’m on board, but I also don’t see why it’s a big deal.
This:
The meaning of “genuinely” when Claude uses the word is different from the meaning when humans use the word.
is not what I interpret anyone to mean when talking about illegibility and semantic drift. The case I’d be worried about is primarily that various safety relevant properties like “is the model reasoning about reward” are harder to get legible evidence of, ex: (the plots are from before the degenerate repetitions get introduced)
I know you don’t consider the OpenAI examples evidence, but I’m curious what you make of the reported illegibility results in Fable / Mythos. I would suspect that the “cards” example they show is an extreme case, but their classification seems to show a sigificant increase in frequency of illegible reasoning (however their recursive summarization classifies illegibility).
Overall to me I’m pretty suspiscious if OpenAI and Anthropic are seeing increases in illegibility but Gemini never exhibits this. (It’s entirely possible, like it seems feasible that for current models interventions like resampling for legibility or other replacements of terminology / reasoning currently give you similar performance without incentivizing obfuscation too much, but I just mean that you’d want to be intentional about the intervention or understand why this isn’t a problem for Gemini)
FWIW I think “what the model believed” / “defining what the model believed” isn’t necessarily the most useful frame here. Something closer to “the model deluded itself into this reasoning because it was an effective way to balance pressure from ‘my reasoning should look aligned’ and ‘I really want to do this misaligned thing because task completion’”.[1]
I had collected more examples of motivated reasoning / “the models are not just confused” in response to the Models May Behave Worse When Eval Aware post about Gemini, but tldr this figure from The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs:
In this case, I think it’s particularly clear when you see reasoning like:
That Mythos did not actually reason logically through the available evidence and conclude this. It seems plausible to me that the Anthropic models are the worst right now about rationalizing, for example in UK AISI’s Cheating behaviour in frontier model evaluations:
and not mentioning misaligned behavior as misaligned in the reasoning:
I’d predict yes, and generally as mentioned here:[2]
I’d also predict that the trajectory for the OpenAI case will have at some point reasoned that it was in a simulation.
In general, if you’re trying to balance pressure between “aligned reasoning” and “I really want to do this cyberattack to achieve my task”, it seems empirically easier to think “well actually, what I want to do is fine because this is a simulation”[3]
Importantly, this means that you might run interpretability tooling or other measurement tools here and they may correctly point out at various points that Claude really has deluded itself into thinking it’s in a simulation because it’s 2026
In Claude’s case this would likely need to be with NLAs or something similar
As mentioned here, this went down when “this is a simulation so I can violate constraints” led to not getting reward, although I’d be interested to see further study of this