This seems totally possible—the EM-from-SFT results make me think that’s not the full picture, but it may well be the case that “SFT on off-policy transcripts doesn’t care about personal consistency, but RL is on-policy and therefore starting consistency has strong effects”
Joey Yudelson
Reward Hacking Without Egregious Misalignment in an RL-Only Setting
In order to add comments to a Google Doc, you’ll need to first download it as a word doc, then either upload it to Claude.ai or use Claude (Code or App) locally.
Since Google Docs currently makes it impossible to programmatically add a comment anchored to text, I right now use some random google-drive MCP off the web, and have claude directly insert comments into the text of the doc, in brackets and another color. This is mildly annoying, but agreed that it’s useful enough I don’t really mind
Eval-suspicion (without full eval-awareness) probably still changes model behavior a lot
Let’s say we’re evaluating Claude Opus 4.8, and it verbalizes eval-awareness ~21% of the time (Section 6.2.4). But it can reliably (79% of the time) detect an eval transcript if it’s placed next to a deployment transcript and you ask which one is real. Which of these numbers should we care about more?
Claim: Even ignoring unverbalized eval-awareness, the ability to pairwise pick out the eval transcript means that there are features the model can pick up on that make the model suspect evaluation. And being even slightly suspicious that you are in an eval might change your behavior a lot.Intuition: Let’s say the model is getting “minor eval vibes~” from some situation. It might gain some benefit from doing the misaligned action, but if the odds of being caught are even 1%, it could shift the cost-benefit analysis a lot. (And much of this reasoning could happen “unconsciously”, analogously to humans getting slightly unsettled by walking alone in the dark)
Joey Yudelson’s Shortform
Thanks for publishing this! In UK AISI’s similar sabotage evaluations, they find much higher eval awareness, validated by the model pairwise distinguishing between an eval transcript and a comparable deployment transcript. Did you try this method of eliciting eval awareness?
If this is true, I wonder if the problem gets easier if we keep both the early-Claude model and the heavily-RL’ed-sociopath-Claude, and trust the early-Claude part a lot, and have both us and early-Claude only defer to sociopath-Claude in specific situations. At the very least, in that world, it seems like a bummer to only have kept around the models which are both super competent and also worrying. And it would’ve been much nicer to also keep scaling up SL’d, character-trained, much less superhuman models
Are these checkpoints available on HuggingFace? I’d be pretty interested to run some tests on them!
Very interesting, thanks for sharing! I wonder what would happen if you simultaneously did this deny-you-are-an-AI training and also RLVR, on math or something—would you end up with a mathematician? Or a conditional persona?
I’d also be pretty interested to see what Mistral-as-Maria has to say if you e.g. ask it/her to write some code for you!
there’s a question of how correlated the failures of the evaluation proxies are with the failures of the training proxies.
This seems very important! It seems quite hard to check if your proxies are correlated, or to what extent.
One possible experiment here: take your N methods of detecting misalignment on AuditingBench. Choose one method, M_0, and train against it. Then measure (1) how much less the model does that behavior, and (2) how much worse the other methods get at detecting misbehavior when it still does occur.
This was really beautiful. Thanks for writing.
Recent Redwood Research project proposals
To me this doesn’t seem like a failure of sophisticated reward models, it’s the failure of unsophisticated reward models (unit tests) when they’re being optimized against. I think that if we were to add some expensive evaluation during RL whereby 3.6 checked if 3.7 was “really doing the work”, this sort of special-casing would get totally trained out.
(Not claiming that this is always the case, or that models couldn’t be deceptive here, or that e.g. 3.8 couldn’t reward hack 3.7)
As a baseline, developers could train agents to imitate the truth-seeking process of the most reasonable humans on Earth. For example, they could sample the brightest intellects from every ideological walk, and train agents to predict their actions.
I’m very excited about strategies that involve lots of imitation learning on lots of particular humans. I’m not sure if imitated human researchers learn to generalize to doing lots of novel research, but this seems great for examining research outputs of slightly-more-alien agents very quickly.
This is very interesting! One thing I’m confused about—I agree that the monitor would have a hard time getting the bits, based on these figures. But the introspective models also don’t seem to get far above 50% accuracy, which means that right now they’d have a very hard time error correcting enough that their comrade would understand them but not the overseer. If the claim is that introspection-tuned models will get much better at self prediction, then its not clear to me that other models wont also get much better at cross-prediction?
Can we make the robots.txt programmatic by page, and then have a tag we can add to exclude a post from the robots.txt? That feels like the 80⁄20
+1, and I hope people are working on more credible ways to make deals with AI. I think if a smart model today were offered a deal like this, its priors should be on “this will not be honored”. Public commitments and deals that can’t be used as honeypots seem excellent.
I wonder if r1 could reliably make the claim that the text had stego-text without being told—i.e. give it a lineup of ten pieces of short creative writing that another instance generated, 1 has steg. See if r1 can pick out the steg, then if various monitors (Sonnet, o1) can also pick it out
This tree is a great place to hold a Kabbalat Shabbat underneath, incidentally
Lighthaven minyan when?
Seems possible, although I had the exact opposite starting intuition!