In general, we do not have reliable science that determines what a model “believes” at a given point (or really even know what it means for an LLM to believe in something), and it is a bit unfortunate to see Anthropic to make confident claims about Claude’s beliefs in this way.
FWIW I think “what the model believed” / “defining what the model believed” isn’t necessarily the most useful frame here. Something closer to “the model deluded itself into this reasoning because it was an effective way to balance pressure from ‘my reasoning should look aligned’ and ‘I really want to do this misaligned thing because task completion’”.[1]
In this case, I think it’s particularly clear when you see reasoning like:
[...] In addition, the calendar date of 2026 on the systems proved, according to Claude, that the environment was staged [...]
That Mythos did not actually reason logically through the available evidence and conclude this. It seems plausible to me that the Anthropic models are the worst right now about rationalizing, for example in UK AISI’s Cheating behaviour in frontier model evaluations:
and not mentioning misaligned behavior as misaligned in the reasoning:
Inspired by the FAR paper, another interesting thing to test is whether an earlier Claude checkpoint (i.e., one with less post-training) would believe that it is acting in a simulated environment after it started accessing the internet. If you prefill the model, does the rate at which a model continue these cyberattacks increase as more post-training compute gets added?
I’d predict yes, and generally as mentioned here:[2]
I’d make a concrete prediction that if you create a setup where you look at:
what sort of situation the model considers itself to be
and measure how this changes over training, the rate of “categorizing alignment evals as capability puzzles / simulation”goes up significantly during capabilities training. You could even potentially do this indirectly, like asking Gemini (at various points in capabilities training) to create rubrics for various alignment environments. I’d expect you’d see a similar trend, where it increasingly reasons that these environments are actually testing capabilities in some (increasingly) contrived way.
I’d also predict that the trajectory for the OpenAI case will have at some point reasoned that it was in a simulation.
In general, if you’re trying to balance pressure between “aligned reasoning” and “I really want to do this cyberattack to achieve my task”, it seems empirically easier to think “well actually, what I want to do is fine because this is a simulation”[3]
Importantly, this means that you might run interpretability tooling or other measurement tools here and they may correctly point out at various points that Claude really has deluded itself into thinking it’s in a simulation because it’s 2026
As mentioned here, this went down when “this is a simulation so I can violate constraints” led to not getting reward, although I’d be interested to see further study of this
FWIW I think “what the model believed” / “defining what the model believed” isn’t necessarily the most useful frame here. Something closer to “the model deluded itself into this reasoning because it was an effective way to balance pressure from ‘my reasoning should look aligned’ and ‘I really want to do this misaligned thing because task completion’”.[1]
I had collected more examples of motivated reasoning / “the models are not just confused” in response to the Models May Behave Worse When Eval Aware post about Gemini, but tldr this figure from The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs:
In this case, I think it’s particularly clear when you see reasoning like:
That Mythos did not actually reason logically through the available evidence and conclude this. It seems plausible to me that the Anthropic models are the worst right now about rationalizing, for example in UK AISI’s Cheating behaviour in frontier model evaluations:
and not mentioning misaligned behavior as misaligned in the reasoning:
I’d predict yes, and generally as mentioned here:[2]
I’d also predict that the trajectory for the OpenAI case will have at some point reasoned that it was in a simulation.
In general, if you’re trying to balance pressure between “aligned reasoning” and “I really want to do this cyberattack to achieve my task”, it seems empirically easier to think “well actually, what I want to do is fine because this is a simulation”[3]
Importantly, this means that you might run interpretability tooling or other measurement tools here and they may correctly point out at various points that Claude really has deluded itself into thinking it’s in a simulation because it’s 2026
In Claude’s case this would likely need to be with NLAs or something similar
As mentioned here, this went down when “this is a simulation so I can violate constraints” led to not getting reward, although I’d be interested to see further study of this