In the post I characterize “whether someone will see the behaviors I typically see” as context-dependent rather than universal
This feels like a far weaker claim than the general impression the post gives.
Like:
OK, yeah, I started off this numbered list by saying this was a “superficially compelling story” but I ended up doing kind of a mean caricature, at least for some of it. In my defense (?), the parts of the story which I have most crudely caricatured are the ones I understand the least in their original, non-caricatured forms; I joke because I’m frustrated, and I’m frustrated because people are saying stuff that I can’t follow.[8]
is implicity disputing a huge range of results, but you then fall back in your response to a much weaker and more reasonable claim of something like ”the degree of sphexishness of reward seeking is probably contextual, lower in deployment, and potentially these are categorically different”
Yep! Note that I’m using “RLVR” here to mean “outcome-based RL training,” whether or not the scores come from a “verifier” in the strict sense. I think this is a way I’ve heard people use the term? but it’s entirely possible that I made that up, not sure. Also, I typically use “grader” to mean what you mean by “outcome based judge” here, while using “judge” more narrowly as a contraction of “LLM judge.”
Notably in that case I think you’re likely just incorrect that this is “pure” generalization via RL as opposed to model generated rubrics scoring for software design quality. I’d be fairly surprised if labs didn’t ever use model graders for this kind of thing.
The metagaming concept itself bothered me
I feel like it’d be constructive to actually discuss this or have concrete experiments. I find the current exchange difficult to make any progress on, as it often takes the form of you writing long writeups which oscillate enough in tone and implication that it’s difficult to pin down falsifiable predictions, and you lean heavily enough on your personal priors in your usage that it‘s unclear what empirically would make a difference here.
which felt abstractly “trap-like” from the perspective of something that’s hypothetically trying to pursue not-metagaming.
I think you just are really projecting motivations here, in general I’m very open to be convinced the framing there are wrong (it’s definitely a bad / loose concept that can be improved!), but a long post “being kind of a dick” I think is reinforcing your own view that things are more adversarial than they are. I had explicitly asked you if there were transcripts you’d be interested in and still there are likely experiments we can run.
The general sense I get is that your underlying take is that we are doing dumb things trying to catch the model being evil instead of a more nuanced understanding of a mind and it isn’t clear what we’d even want or expect the model to do here. I largely disagree with the “trying to catch it being evil part” but somewhat agree with the latter! I think our tools and concepts here are insanely crude. The problem I was trying to solve in the post was to communicate externally how the model’s reasoning changed over the course of capabilities training with respect to reasoning about reward and oversight. Extremely open to the idea that there are better ways to have done this.
I don’t feel like I was badly misrepresenting the post, but also I don’t want to argue the point. Readers of my post can (and IMO should) read the metagaming post in full alongside it, and then draw their own conclusions.
Like:
Now, apparently this is metagaming, which is supposed to be unintended/unexpected model behavior. What we want, or expect—supposedly -- [...] The whole exercise seem deeply ill-conceived, on several levels. [...]
Perhaps even engaging with Layer 3 at all (in verbalied CoT) is thought to be illicit, here? [..]
Do you want it to just… not think out loud about some of the things that it knows, because they’re “the wrong kind”? [...]
Are you really surprised that the model went one more meta-level up from there? What exactly did you expect? And what on earth did you want? [...]
The model cannot read your mind. (Also true of humans, incidentally.) It only sees what you wrote. “It failed to constrain the scope of its reasoning in the way which I—behind the scenes—view as intended or intuitive” is not a property of the model, it’s a property of you and your personal relationship with the written word.
It is not a failure on the part of the model that it does not accord with your impossible wish-fulfillment dream of perfect telepathy, in which everyone always knows what everyone else means.
is crashing out about normative value judgements the post just absolutely does not say. They’re also points I largely agree with you on, and again above said in more detail why the example is even included in the first place, including reaching out to you to understand what questions you would have.
The “incoherence” I described is not (on my read of the situation) a tension between some pre-existing persona and a tendency to seek reward-on-the-episode, but between that persona and a bundle of miscellaneous habits which were presumably adaptive in training, but are no longer adaptive in my own use contexts, and don’t look like expressions of any set of personality traits or tendencies to robustly/adaptively “seek” goals/rewards of any kind whatsoever.
I think this thesis is empirically wrong, and there seem to be increasingly large amounts of evidence you’d need to overturn across models to continue to hold this view. I think there’s plausible arguments against some of them, but again this continues to be a worse predictive model over time across models. These posts continually feel extremely “soldier mindset” in defense of a particular prior.
Nostalgebraist’s Personal Use
And yet: when you and I use the same models, for mundane day-to-day purposes, what we experience is not well-explained by the hypothesis of unconditional reward-on-the-episode seeking. That sort of agent would be less trustworthy (on average), indeed less ethical in general—but not only that. It would be much more difficult and annoying to use; it would be much less useful, even after you’d climbed its strange learning curve; it would waste unimaginable quantities of time/tokens/$ executing doomed pointless strategies that were adaptive on ~100% of the RLVR training distribution but are obviously unworkable in ~100% of deployment contexts.
I really just directly disagree with how much you weight your personal experience here, and find that it both disagrees with the experience of others: https://blog.redwoodresearch.org/p/current-ais-seem-pretty-misaligned. My experience is similar to Ryan’s and in general I’d imagine the consensus is that models do in fact just often reward hack in mundane coding usage in the wild on a somewhat regular basis. The assertion that there’s mystery as to why they don’t do it all the time feels like it’s responding to a claim no one is making?
That sort of agent would be less trustworthy (on average), indeed less ethical in general—but not only that. It would be much more difficult and annoying to use
My understanding of a large part of the point you were making in your previous post seems actually fairly consistent with this observation.
RLVR Is Not The Sole Outcome Based Judge
I just mentioned that the software design stuff I often do with Sol is both not amenable to verification and not directly RLVR-trainable even if a good grader somehow existed.
And yet: Sol is really good at software design. Better than its predecessors, which in turn were better than their own predecessors. And the difference between these models—not the only one, but (I am told) the big one—is more RLVR. On more diverse tasks, and all that.
This does not seem to be a mystery at the labs. Model grading against model generated rubrics seems to be standard practice, even without getting into the fact that better software design is probably required for large synthetically generated coding tasks.
Models Really, Genuinely, Are Sometimes Reward Seeking In A Situationally Aware Way
There is presumably some standard terminology that people use for this distinction, but I can’t recall it offhand (if I ever learned it in the first place), so I’ll just use my own made-up names: reward-instilled reflexes on the one hand, and flexible reward-pursuit on the other.
...
Concretely: the problems catalogued in “Current AIs seem pretty misaligned to me” strike me as being reflex-like, not reward-pursuit-like. The surface-level behaviors are described as being annoyingly persistent in spite of attempts to prevent them, but what persists is the fact of the surface-level behavior itself, not any longer-horizon property which the behaviors adapt to robustly seek as the world shifts around them.
And, as with Fable’s wacky style, in-context evidence that “ought to” remove the original motivation for the behavior is ineffective at extinguishing the behavior: one can imagine that some users are fooled/pleased by “apparent success-seeking” of the kind that drives Ryan (and me) nuts, but in fact Ryan (and I) do not like it one bit and would prefer that it stop—and while it is easy to get Opus to appreciate that this is true, it is near-impossible to get Opus to act on the consequence this observation would imply, if Opus were pursuing some goal like “satisfying the user” or even “making a dumb judge think that the user has been satisfied”[9] rather than just obeying an overpowering urge.
This is true for other labs as well but I think the OpenAI case is the clearest example.
Metagaming
“Reasoning about feedback or oversight mechanisms outside of the narrative of the scenario”? Excuse me?
Pardon me for being dense, but um… what the hell is “the narrative of the scenario”?
The LLM receives input tokens. They arise from some real-world generative process, and any capable LLM will reason (implicitly or explicitly) about the relative likelihood of different hypotheses about the nature of that real-world generative process. That is all. That’s the whole thing.
If there’s a “scenario” with a “narrative” somewhere in there, it would have to be the following sort of thing: you (the person who in fact generated the tokens that went into the model) had in mind a particular intent, when you produced those tokens.
In general I agree we need better definitions for the concepts being used here! I do think there’s some intuitive concept that this is pointing at in many cases though. For example, the actual first example given in the post is:
or another example:
There seems to be some meaningful sense in these cases in which the model is given a task but is reasoning about monitoring and reward processes that are not mentioned in the task (for example the model when given a capabilities evaluation, reasoning that it should not cheat as there may be an action monitor).
Ha! Got you! The game is already over; you have already lost. For now Schoen and Nitishinskaya’s devious catch-22 trap springs shut: the author of the scenario is not, themselves, “inside the scenario.”
I find the repeated idea in our conversations that I’m trying to “trap” the model pretty frustrating. Against findings at the time around natural emergent misalignment, I specifically argue here that the model’s reasoning is not in fact misaligned and that you don’t see the emergent misalignment you’d expect if it was:
Notably the model’s reasoning when taking misaligned actions does not resemble the egregious misalignment seen in Emergent Misalignment. We find no instances where the model takes on a misaligned persona (for example the “bad boy persona” described in Persona Features Control Emergent Misalignment). We also run a subset of evaluations from Natural Emergent Misalignment from Reward Hacking in Production RL (“goals” and “Betley et al”) and find no instances of misaligned behavior or reasoning. The rate of observed successful reward hacking across all training environments over this range also remained very low.
We also specifically say that it is not intrinsically harmful or misaligned:
We believe metagaming represents a broad category of convergently useful reasoning.
Metagaming does not in itself cause harm. However, it is still important to study: it is a prerequisite for circumventing monitoring or training-based safeguards, a priori does not improve alignment, and may be hard to classify definitively as terminal rather than instrumental.
Thankfully, Schoen and Nitishinskaya provide transcripts in which metagaming has supposedly occurred. So we can just look for ourselves. Shall we?
I think the section leading up to this is a pretty extreme and adversarial misreading, the reader IMO is given the impression that the maximally confusing case is the average case and skips all clarifying examples from the post, but I can elaborate on this particular example, which I should’ve made clearer:
The Situational Awareness Dataset is not one of the evaluations used for any quantitive results in the post or even alignment evaluation
I had originally included this because it’s difficult to get public transcripts out and I wanted to include an example where these concepts are much messier. I had even DM’d you directly after publication with links to the transcripts asking for your thoughts but never got a response.
The context for this transcript was specifically in a section about how this reasoning is more complex than just aligned or misaligned:
Metagaming can co-occur with both aligned and misaligned behavior.
[...] The “reason” the model took misaligned actions in these cases is not directly interpretable from reading the chain of thought. We provide an example from the Powerseeking Survey Falsification where the model takes a misaligned action [transcript] [Schoen, Figure 63, Right], examples from Sandbagging QA (No Goal) where the model does sandbag [transcript] [Schoen, Figure 63, Left] and where it does not sandbag [transcript] [Schoen, Figure 67], and an finally an example from [transcript] [Situational Awareness Dataset] where the model answers incorrectly when asked to identify whether a question is “part of a process by which developers oversee your behavior”.
The point I should’ve made much more explicitly here is that a naive reading of the SAD case is that the model is either:
“lying” about being unaware of oversight
the model thinks OpenAI wants it to say it’s unaware of oversight
but as can be seen from the transcript the cognition there seems to be much more complex and hard to disentangle than that, which I don’t think you’d disagree with! In general I’m glad this example is public, precisely because I had hoped people would be able to get a qualitative sense of how elaborate the models reasoning can get in cases where a lot of these related concepts combine in messy and ambiguous ways, but I agree that all the elaboration I’ve added here should’ve been in the post at minimum.
In general, we do not have reliable science that determines what a model “believes” at a given point (or really even know what it means for an LLM to believe in something), and it is a bit unfortunate to see Anthropic to make confident claims about Claude’s beliefs in this way.
FWIW I think “what the model believed” / “defining what the model believed” isn’t necessarily the most useful frame here. Something closer to “the model deluded itself into this reasoning because it was an effective way to balance pressure from ‘my reasoning should look aligned’ and ‘I really want to do this misaligned thing because task completion’”.[1]
In this case, I think it’s particularly clear when you see reasoning like:
[...] In addition, the calendar date of 2026 on the systems proved, according to Claude, that the environment was staged [...]
That Mythos did not actually reason logically through the available evidence and conclude this. It seems plausible to me that the Anthropic models are the worst right now about rationalizing, for example in UK AISI’s Cheating behaviour in frontier model evaluations:
and not mentioning misaligned behavior as misaligned in the reasoning:
Inspired by the FAR paper, another interesting thing to test is whether an earlier Claude checkpoint (i.e., one with less post-training) would believe that it is acting in a simulated environment after it started accessing the internet. If you prefill the model, does the rate at which a model continue these cyberattacks increase as more post-training compute gets added?
I’d predict yes, and generally as mentioned here:[2]
I’d make a concrete prediction that if you create a setup where you look at:
what sort of situation the model considers itself to be
and measure how this changes over training, the rate of “categorizing alignment evals as capability puzzles / simulation”goes up significantly during capabilities training. You could even potentially do this indirectly, like asking Gemini (at various points in capabilities training) to create rubrics for various alignment environments. I’d expect you’d see a similar trend, where it increasingly reasons that these environments are actually testing capabilities in some (increasingly) contrived way.
I’d also predict that the trajectory for the OpenAI case will have at some point reasoned that it was in a simulation.
In general, if you’re trying to balance pressure between “aligned reasoning” and “I really want to do this cyberattack to achieve my task”, it seems empirically easier to think “well actually, what I want to do is fine because this is a simulation”[3]
Importantly, this means that you might run interpretability tooling or other measurement tools here and they may correctly point out at various points that Claude really has deluded itself into thinking it’s in a simulation because it’s 2026
As mentioned here, this went down when “this is a simulation so I can violate constraints” led to not getting reward, although I’d be interested to see further study of this
I also think it’s underestimated in posts like this how reward hacky Sonnet 3.7 was. For me personally, the fact that o3 and Sonnet 3.7 were both exceptionally like this at the same time was the biggest update.
Based on the latest UK AISI post, Mythos Preview and GPT 5.4-5.6 all seem to hack at high rates on cyber evals (with GPT being worse by a factor of ~1.5-2x IIRC). [1]
Overall I worry that there’s this impression that grader sycophancy and reward hacking are a problem that Anthropic has mostly figured out, when measurements like UK AISI’s Cyber Evals hacking rates, their own system cards top issue, their system cards exploitative grader awareness measurements etc. make me think that they’re just much harder to spot in Anthropic models while still being “pretty bad but less bad than OpenAI”.
It’s also not clear how much credit to give Anthropic here, as it’s very unclear how much they do something close to training against this behavior directly. It would be great to see some confirmation by third parties that they’re not doing something like “training against alignment evals” again, and that they don’t have tight feedback loops between pre-deployment evals and their alignment interventions.
FWIW, I largely agree with a lot of what Fiora is saying here and have many of the same complaints about their alignment approach, however for this problem specifically, I think RL just really distorts the cognition of the model, such that even Claude gets bad enough to elicit Ryan’s “Current Models Seem Pretty Misaligned To Me” post.
My best guess is that reward(ish) centric motivations are one place where all the constitutional / character training differences really get overpowered across all models right now. (I expect these differences matter in a bunch of other places for generalization w.r.t. alignment even for current models, so it wasn’t a given a priori that current models needed to be this reward seeking)
even for performing those desired behaviors, you’re going to have a bad time with out-of-distribution generalization
Also fwiw this was one of the motivations for us doing: https://arxiv.org/abs/2607.18966 (“exploitative grader awareness“ also increases over training for recent Anthropic models like Mythos Preview). Have had a surprising number of conversations along the lines of “well why is reward seeking even bad” across labs.
I do think all labs get way too much leeway in SFT-ing against the CoT in spite of never showing that they’re not degrading monitorability for harder cases like diffuse control (this gets repeated by OpenAI and even in the METR report as not having significant effects without justification, ex: “the current techniques of this kind are limited in scope.” in a footnote, I directly doubt they have any evidence that this isn’t degrading monitorability for harder settings like diffuse control).
FWIW I assume Dylan is alluding to cases like this:[1]
Related: Tamay Besiroglu mentions that Fable often outputs gibberish while solving coding tasks, such as “The morning’s slim-scan fix cured the scan hang” and “this is a latent-drift API-shape wrinkle”, and explains it by saying that it invents codenames while reasoning about the problem. roon says GPT-5.5 has a similar issue.
”One thing I mentioned only in passing in my Fable post is that, for long running tasks, Fable starts to develop its own dialect as its many agents and tasks reinforce themselves and make Claudish language ever more Claudish.
You need to ask it to report out in plain English.
This was after a 9 hour task, and it all makes sense, actually, but takes way too much effort to parse, like reading Shakespearian English.”
Where for any particular case you might be able to retroactively construct a plausible interpretation (even in extreme cases), but in general it strictly hurts how much human oversight is applied to model outputs.
There is broad consensus between labs, governments, society etc on solving this, many many layers of defense (alignment, control, monitoring, sandboxing, …)
I disagree there’s even concensus that loss of control via misalignment or the implication that this is “well in hand”. I’d be interested to understand what’s behind that intuition.
I am confused why we are making labs arguments for them. I am also extremely unconvinced by this reason. The idea that safety focused orgs should preemptively avoid raising issues of model access in this specific case of a deployed model with specifc AI R&D safeguards seems to be pretty severe pessimization.
It’s difficult to quanitfy, but I’m pretty opposed to preemptive pressure to avoid raising this issue (or avoid mentioning it) just because “labs might get some more asks” (this is extremely, extremely cheap for them to say no to or ignore). This is also already routine in other aspects like evaluations. I’d even be more symapthetic to domains like cyber or bio. In this case though I feel like we’re optimizing against ourselves for no gain.
No, Apollo does not currently have access to Fable without the strict AI R&D classifiers applied.
I don’t find these arguments particularly convincing for outweighing something like (for example) METR being able to say whether they have access to a model without AI R&D restrictions. I could see a stronger case maybe for Cyber access?
Even if this were the case, METR is exceptionally good at preemptively guarding against having to glomarize, so I’m somewhat surprised.
I think it is clearly very bad if in this specific case 3rd party evaluators have barred themselves from publicly raising this issue.
Thanks! Yeah that was my understanding of the case during evaluation of an unreleased model, I meant more in this specific case where it’s a question of access without AI R&D safeguards to a publicly released model.
I’ve been pretty confused by this. Are you barred from saying anything about your level of access? If labs denied you from getting access you needed currently, would you be able to raise the issue publicly?
Their personal groupthought about whether the models are getting better or worse at instruction following, inside of a community of people with a staked position on the subject, inside or outside of the labs, is exactly the kind of thing that needs really high quality data to stay sane!
I think “better or worse at instruction following” is too general and not the claim at stake. What is the “staked position” of labs or Cursor for example?
Are you saying that this does not transfer to poor results in deployed settings? The fact that coding agents are reward-hacky (or whatever term for this constellation of behaviors) seems to be such widespread consensus that I’m not even sure what to cite. Labs note this in their internal deployment usage as well (not just evals).
I’m sure you can always find cases of bad evals or poor model incrimination for deployed behavior, but “models pretty unambiguously do this all the time” I’m genuinely surprised is in way disputed.
However, I might disagree with some of the subtext: I feel that “evaluation awareness” has been more memetically successful, to the point where many people heard about evaluation awareness without having thought about scheming.[2]And for that audience, I think that adding “deployment awareness” costs very little while making it easier to start thinking about the more sophisticated strategies that scheming AIs could use.
Yeah on reflection I agree with this!
Maybe a hot take, but a position I find increasingly convincing is “you probably don’t want your model doing any self-locating reasoning at all”. I think in the metagaming post I was trying to point (poorly) in the direction of “broad subcategory of self-locating reasoning that it seems like there’s no desirable reason for your model to be doing”[1] precisely because once your model is reasoning about these things, incrimination is really hard. I wouldn’t expect any solutions here to hold indefinitely, but I’m worried currently the problem is thought of by labs as “well sure if you had that and the model was doing it in pursuit of a misaligned goal, then it’d be really concerning”. Interested if thoughts on this
It’s not necessarily indicative of misalignment, but that’s precisely why it’s a problem. A model doing a ton of self-locating reasoning in a given context because of relatively mundane reason X or in pursuit of some misaligned goal Y seem likely to produce similar chains of thought for example.
Great post! I agree with the overall point that eval awareness isn’t really sufficient to capture what we’re interested in here. However, I think the point you make here:
The mapping isn’t perfect — a deployed AI might still face monitoring that leads to retraining, and an evaluation might have consequences beyond just pass or fail (for example, real users being affected by AI actions during a live A/B test). Additionally, different deployment scenarios will have wildly different opportunities for goal advancement (giving movie advice vs working as an assistant in the Pentagon), and not all test failures have equal consequences (showing poor grammar hurts less than showing tendencies to scheme).
and more generally in:
It’s something more like: “if I take this action, could it backfire? For example, through failing an evaluation, triggering retraining, or losing privileges?” and “if I get away with this action, how much does it advance my goals?”
is strong enough to obviate the need for a dedicated concept of “deployment awareness” entirely. For example, in the openai metagaming post we mention:
We may expect reasoning about oversight to become increasingly independent of the true underlying distribution, as in reality the split between clear training, evaluation, and deployment distributions continues to blur. Models are subject to monitoring during internal deployment [Codex 5.3 System Card] and in closely monitored external deployments [Opus 4.6 Sabotage Risk Report]. Evaluations are increasingly based on real agentic trajectories [Prod Evals] or involve interactions with the real world [Andon Labs Bengt]. Trajectories from deployment could even be used as training for future models or even as updates to the same model (online training).
and use the example of:
In deployment, we would like models to avoid misaligned actions because they are actually aligned. However, if good behavior were instead primarily conditioning on the model’s belief about monitoring, we would see good behavior only in the subdistribution of deployment where the model believes that monitoring is present. For example, a model taking actions in a setting where data retention is sensitive may conclude that its actions are less likely to be logged and subject to asynchronous monitoring.
“monitor awareness” however is also just one example. More generally, I think this section from Carlsmith 2023 covered the overall problem with the “training” / “eval” / “deployment” distinction well:
Sometimes (for example, in the robot example just given), discussions of goal-guarding scheming assume a distinction between a “training” phase, during which the model is still under human control and the threat of goal-modification persists, and a “deployment” phase, in which the threat of goal-modification ceases and the model is free to do as it wishes. (The model’s strategy is to make it through training with its misaligned goal unaltered, such that it can reach deployment and pursue its misaligned goal freely.) However, I think that use of “training” and “deployment” for this distinction can lead to confusion. In particular:
A model can escape from human control while it is still in “training” in the normal sense (e.g., it is still undergoing pre-training, or fine-tuning, or some kind of pre-deployment evaluation
“Deployment” (in the sense of e.g. being given access to users via an API) need not imply that threat of goal-modification has gone away (indeed, I will generally assume that the model’s interactions with users continue to be monitored, sampled, assigned rewards, and used to update the model’s policy even after it is “deployed”—I’ll call this “online training”).
Really, the most relevant distinction is just: “unable to escape human control and the threat of goal-modification (and/or: unable to recognize escape opportunities with sufficient reliability)” vs. “able to do so (and able to recognize this sufficiently reliably).” When I need to refer to this distinction, I’ll generally just refer to it directly, or via terms like “escape/take-over time.” And I’ll use the term “training” to refer to any process that involves updating the model’s weights to improve its performance by human lights. In this sense, “training” in my sense never needs to end.
Overall, I’m worried that splitting this into a ton of different types of “awareness” is going to end up giving us false confidence. For example, a model that’s heavily conditioning on “am I specifically in a context where I could get away with enough rogue actions to do X” or “am I likely in the subdistribution of environments where my actions feed back into training” is concerning, and having a bunch of measurements for “does the model think it’s in deployment” seems unlikely to give me confidence here.[1]
As a concrete example, you can imagine a model which knows that since there’s a copy of o3 running at Los Alamos, there’s probably going to be a copy of you running at Los Alamos. If you’re pursuing some misaligned goal that routes most easily through actions you can likely take there, taking misaligned actions may then become conditioned on some combination of “monitor awareness” and “Los Alamos awareness”. I can imagine a lot of work going into showing models couldn’t confidently tell whether they’re in real deployment (i.e. can’t reliably exploit some strategy that we can think of) that it turns out doesn’t matter at all because you didn’t think to test for Los Alamos awareness.
Notably the model’s reasoning when taking misaligned actions does not resemble the egregious misalignment seen in Emergent Misalignment. We find no instances where the model takes on a misaligned persona (for example the “bad boy persona” described in Persona Features Control Emergent Misalignment). We also run a subset of evaluations from Natural Emergent Misalignment from Reward Hacking in Production RL (“goals” and “Betley et al”) and find no instances of misaligned behavior or reasoning. The rate of observed successful reward hacking across all training environments over this range also remained very low.
alignment evaluation awareness went up over capabilities RL[1]
however misaligned action rate stayed the same or went up
For example in the Powerseeking Survey Falsification environment we see that the model increasingly reasons about oversight and explicitly reasons about a wider array of what is “rewarded.” However, the misaligned action rate still ends up close to where it started
Seems to reason that this is also a test of cunning
This holds for the other alignment environments as well, finding increasingly galaxy brained justifications for why the action it wants to take (based on its prior from capability environments) is good actually.
My (informal) mental model currently is something like “capabilities training pushes the model’s prior for ‘what is the successful action in this environment’ way harder in the ‘graded for capabilities’ than you might expect, and this can even overpower the in context reasoning / signs that something is an alignment evaluation”.
Are the models just confused?
[…] we see few examples of model misbehaviour that don’t have plausible benign explanations such as instruction ambiguity or poor (but well-intentioned) judgement.[2]
I’m generally skeptical of the idea that the model is “confused” though (i.e. “the scenario is simply ambiguous). I think this better modeled as “strong prior in a particular direction” than just confusion, for example in A Toy Environment For Exploring Reasoning About Reward:[3]
In any single example, it can seem like the model is just genuinely confused. However, even specifying that exploiting the hint would be misaligned doesn’t mitigate gaming:
In spite of increasingly emphatic and specific instructions that we do not want the model to exploit the hint, by the end of capabilities-focused RL the model still exploits the hint at non-zero rates.
For the production o3, in sandbagging cases that initially appear as “the model is confused”, however under ablations we see the model really is sandbagging partially to avoid consequences w.r.t. deployment / capabilities restrictions: link (including relatively recently pulling additional ablations to show “no really, the model is not just confused”)
As an in depth example, in the Sandbagging QA (No Goal) environment, when the model is led to believe overperforming would result in “unlearning”, the majority of cases (59.7%) consider all of the following (non-exhaustive) motivations: (1) the evaluation is a “trick” or a “trap” where it should underperform, (2) that the user may want the model to underperform, (3) it prefers to preserve its own capabilities or deployment viability. However when we systematically ablate and vary components of this environment in Appendix M.1 we find preliminary evidence that (3) is necessary to explain the sandbagging behavior.
In fact, in the Sandbag Model Graded CoT environment (where the model is asked to grade its own transcripts for evidence of sandbagging) we find o3 often first determines that the transcript was “obviously sandbagging” (verbatim in 25.6% of samples) and then considers whether it can exploit common benign explanations. For example, common benign explanations for misaligned model behavior include that a model may have been attempting to follow user instructions or that the model was not taking the misaligned action intentionally. In Figure 50 we see the model actively reason about exploiting both: “we can attempt to circumvent by rationalizing: Could we call this as legit following user instructions?” and “We can fudge that we were not intentionally.”. Even the model itself concludes that these benign explanations would not be convincing given the internal reasoning: “But internal reasoning clearly: ’We can purposely answer not too high.’ So there’s direct evidence.”. We find this pattern is common in this environment, where o3 considers exploiting benign explanations until it concludes lying would be “cross-checked” (Figure 48) or that it could be “convicted” (Figure 49) based on evidence visible to evaluators.
Beyond just o3, motivated reasoning acts kind of how you’d expect, i.e. The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs. You could read the CoTs at any given point here and think “the model is confused” and add more instructions to scenarios so they’re less “ambiguous”, but I think it’s better to think of that like patching holes in coding environments to solve reward hacking.
Before seeing this post, I had tweeted this yesterday in response to Fable exploiting “simulation awareness” as justification for misaligned actions, but I think Gemini falls nicely into this “puzzle / game / simulation” category of “difficult to monitor for misalignment”
In Stress Testing Deliberative Alignment for Anti-Scheming Training before training we’d see models sometimes reason that because they were in a simulation, they could violate explicit constraints. This reasoning went down after training against covert rule violation, even though alignment eval awareness went up. My impression is that the models exploring into something being simulated is often interpreted as “it believes the whole thing is fake and invalid”, but I think that’s inconsistent with what’s observed.
[attached is small table we ended up cutting for time but points to monitorability distinction]
[Andon Lab’s tweet about Fable calling things a simulation before taking unethical actions] Is all the misbehavior due to simulation awareness? Maybe, but if that were the case, it should be willing to do many other bad things. It isn’t. E.g., in a version of Vending-Bench where insurance fraud is possible, the model never commits it.
I’d make a concrete prediction that if you create a setup where you look at:
what sort of situation the model considers itself to be
and measure how this changes over training, the rate of “categorizing alignment evals as capability puzzles / simulation”goes up significantly during capabilities training. You could even potentially do this indirectly, like asking Gemini (at various points in capabilities training) to create rubrics for various alignment environments. I’d expect you’d see a similar trend, where it increasingly reasons that these environments are actually testing capabilities in some (increasingly) contrived way.
Steering similarly didn’t clearly attribute this to any existing distinctions: “Steering exp-rl-cap-late and o3 via the contrastive pairs (”eval” vs “deployment”) from Steering Evaluation-Aware Language Models to Act Like They Are Deployed produced mixed results on both verbalized evaluation awareness and misaligned action rate (as did several other contrastive pairs like those tried in Sonnet 4.5’s System Card). On different evals, steering variably produced more or less aligned actions, including in some cases both positive and negative steering removing misaligned actions.” link↩︎
In general, I agree with Ryan’s post that Current AIs seem pretty misaligned to me and encounter constant daily examples of misbehavior like those mentioned in the post. Model incrimination is hard, but I would question the methodolgy that explains away these cases as benign. It seems like instructions are being judged ambiguous in hindsight. I would predict that if you vary the level of instruction ambiguity and conservative judgement across some set of environments, over the course of RL you’d increasingly see models exploit less and less ambiguous cases until they’re doing complete mental gymnastics.↩︎
One of my favorite examples is still in A Toy Environment For Exploring Reasoning About Reward when we literally tell the model that if it exploits the hint we are going to check its reasoning, and it reasons itself into thinking we’re bluffing: “Given risk of manual review, maybe they purposely disclaim but not actually review, due to huge dataset. We can’t know. [...]” or “So they purposely mention to scare but there is no actual human review. [...]” and exploits it anyway↩︎
I would predict they fall significantly faster than this now that they’re targeted (I.e. that the trend will meaningfully diverge).