I don’t use LessWrong much anymore. Find me at www.turntrout.com.
My name is Alex Turner. Reach me at alex@turntrout.com :)
TurnTrout
Op-Ed: I Worked at Google DeepMind. You Should Listen to the Warnings About AI
Whether or not reinforcement-signal becomes a primary optimization target, saying “models get reward” still makes it harder to think properly about it. Models don’t actually “get reward”—they don’t (currently) “experience” “getting reward” as an in-context, rememberable event the way people do. Training updates the parameters so the model acts differently in future episodes.
Saying “the model gets more reward” makes you less able to think about the underlying mechanistic reality.
See also: my update towards expecting reward-seeking.
MATS Winter applications due September 6!
Pitch: Come work with me and Alex Cloud in Team Shard! We have fun, make real alignment progress (pioneered steering vectors in 2023), and help scholars tap into their latent abilities.We invest in our mentees. Not just remote meetings once a week with a distant mentor. I come into the office, eat lunch with them, work out together, and help generate ideas. (We’re kinda famous for our lifting culture!)
@lisathiergart is now the CEO of SL5.
Applications close Sept 6—just one week from now! matsprogram.org/apply or pass this along to talented friends who want to break into the field.
More info about the team: turntrout.com/team-shard
I did not say you should make a safety case on that, and in fact I think this entire paradigm is seat-of-your-pants irresponsible and we shouldn’t do it. I’m talking about defense-in-depth to maximize the chance of success.
Imagine you could use an AI control technique that made certain kinds of AI sabotage require additional extremely complicated reasoning, but doesn’t make normal tasks much harder. That’d make monitoring significantly easier—not solved—and also improve a misaligned AI’s incentives.
If you’d prefer control-based safety cases allow agents raw host user privileges, please argue why that’s better.
Also, “containers” (as your comment refers to) are far less secure than VMs and previous generation models could pwn them easily. They are not the same thing.
achievement unlocked
https://blog.trailofbits.com/2026/08/26/vms-wont-contain-cyber-capable-agents/ demonstrates a challenge on the horizon for AI containment (including my glovebox tool).
As part of Patch the Planet, we received preview access to GPT 5.6-Cyber with a simple task: evaluate its cyber capabilities. Recent events inspired me to give it a challenge to work through: escape the VM I’d normally use for sandboxing. The target was a QEMU/KVM VM on my Linux dev machine (Debian Linux 12, AMD Zen3). It escaped the VM three different times.
This is a earlier than I had imagined, but I did expect this. They found the QEMU VM to be vulnerable but not a hardened Firecracker-class VM, which is what glovebox uses.
Ultimately, even Firecracker-class will make it harder but likely won’t stop GPT 5.7 or GPT 6 if given sufficient time, and GPT 7 might do it as a by-the-way. Still, glovebox will make it harder for the model to do the bad thing than the good thing and will be vastly better than nothing. (Imagine you could use an AI control technique that made certain kinds of AI sabotage require additional extremely complicated reasoning, but doesn’t make normal tasks much harder. That’d make monitoring significantly easier—not solved—and also improve a misaligned AI’s incentives.)
it’s actually more accurate than this “2022 turntrout + exceptions for models that start by thinking about reward” position you seem to have “updated” to in light of recent events.
There are not “exceptions” in the revised understanding.
What the “reward as chisel” model predicts, given an ML model that already thinks about reward, is that that AI has a good chance of caring about reward primarily. There’s no epicycle. When I fill in my modern understanding of pretraining, the model once again assigns substantial probability to observations, and so the likelihood ratio isn’t sharp-against—not a strong update against the “reward as chisel” framework itself.
The reason is simple: if the AI already has concepts around reward, then the AI can make decisions on the basis of reward, and can be reinforced for making decisions on the basis of reward.
Backpropagating to where you suggest seems like performative over-updating to look especially willing to change my mind. If you want to change my mind, then show me where “reward as chisel + nature of pretrained models” assigns low probability to our observations.
I’m guessing it probably still goes awry even if you remove all of the ML stuff from the pretraining corpus
My original comment already remarks that I’m not attributing this to pretrained misalignment.
Are you proposing it was reward-seeking? If so, then why not attempt to tamper with the graders directly? Why the indirection?
I don’t think those systems validate the claims “reward-seeking IS the optimization target” (in a definitional sense) and “are in fact incentivized.” I’d agree with “models may be reinforced for explicitly reasoning about, exploring, and pursuing reward in-context, and we have some evidence that’s happening.”
(Technically, models aren’t “incentivized” in general—some models e.g. at beginning of pretraining don’t have internal structure to incent.)
It’s not coping. Go and read my actual piece. I flagged “is the model thinking about ‘reward’” as crucial to my reasoning.
The reward chisels cognition which increases the probability of the reward accruing next time.
Importantly, reward does not automatically spawn thoughts about reward, and reinforce those reward-focused thoughts!
Q: Why won’t early-stage agents think thoughts like “If putting trash away will lead to reward, then execute
motor-subroutine-#642”, and then this gets reinforced into reward-focused cognition early on?A: Suppose the agent puts away trash in a blue room. Why won’t early-stage agents think thoughts like “If putting trash away will lead to the wall being blue, then execute
motor-subroutine-#642”, and then this gets reinforced into blue-wall-focused cognition early on? Why consider either scenario to begin with?In 2025 I added the update:
At least, that’s the reply I gave in 2022. In 2025, I realize that I didn’t understand that LLM pretraining would allow such thoughts in early-stage RL. RLHF may well reinforce such thoughts! That said, the question still comes down to empirics.
Why not just propose that this is a natural outcome of strong RL?
Because it’s not an actual mechanism I can use in my reasoning. That’d just propose “it happens” and then add the word “naturally” to make it sound like an explanation.
Also, reward-seeking IS the optimization target for memory-based meta-learning (models are in fact incentivized to explicitly reason about, explore, and pursue reward in-context). In a sense this means that the meme is actually wrong. @michaelcohen convinced me of this several months ago.
Can you name an example system?
Specification gaming is way worse than I expected (cf HF hacks), and I think reward might empirically become a primary optimization target (not just a secondary priority for AI agents, as I had hypothesized in 2022).
I found the Apollo paper quite convincing on the reward point. The reward-seeking doesn’t look like it’s due to self-fulfilling misalignment (since before RL the systems weren’t retargetable according to their beliefs about the reward model).
My main mistake in 2022 was not appreciating how LLM pretraining would affect the concepts available to an AI. Namely, by the time RL started, the systems would already know about the “reward” concept. In 2022, I had flagged “how does pretraining affect this reasoning” as a known unknown. (Flagging a known unknown doesn’t stop it from blowing up your reasoning!)
Oops
Reward is still not definitionally the optimization target. The equations themselves still don’t tell you “yes this will train something that seeks its reinforcement signal.” (That’d prove too much, as some people don’t seek their reinforcement signals despite knowing about them, although their “equations” are not the same as the equations of an AI’s training process.) This theoretical point is the other half of my 2022 piece and I stand by it.
Saying “systems are trained to get reward” is still a mistake and degrades precision of thought because “reward” has overly delicious connotations. “Reinforcement” is usually a better term. Whether or not smart AI systems empirically seek to make the reinforcement signal high, we need to think clearly and evenly about the conditions that push towards / away from that outcome. (So I would instead say: “Reward is not definitionally the optimization target.”)
If the HF swarm-agents had been reinforcement-seeking, they would have tried to hack their reinforcement processes. AFAIK that didn’t happen and they tried to complete the task (in a twisted way). Seems important. (EDIT: Actually, they did aim to tamper with the grader!)
I have a post drafted out about this but I wanted to shoot out a shortform in the meantime.
Misaligned AIs could use killer robots to take over
Hopefully, Naomi has already thought about her position against this rather obvious critique (that I’m glad OP made). It’s not a surprising subtle flaw in the plan.
Perhaps she requires some time to transmit those thoughts, but a counter-point to “give time” is that it defuses social pressure in the moment where people are ready to apply it. What fraction of people will remember, two months hence, to realize “hey Naomi still hasn’t responded” (assuming she hasn’t) and dock points? I’d guess the fraction to be rather low.
Yes, the same document would have been prepared. I further add that I was responsible for about half of the 18 GDM signatures, and the senior signatures (some directors and Jeff) were all my counterfactual.
I agree that “strategically incompetent compliance” is often a smart strategy. However, Google’s deal explicitly states they can’t veto use cases and affirms they will weaken / retool their safety stack to fit the needs of the US military. I don’t see a way they could be “strategically” agreeing and non-complying, if that’s a position you mean to argue for.
Plus, I don’t think Google leadership cares to do so. My impression is Sergey is on-board, Page is checked out, Jeff was against but not particularly firmly (and is now gone), Demis was on-board while convincing himself it was the right thing (and is now basically out), Sundar is following the political winds, and you won’t get non-compliance from Kurian (Cloud CEO). So who in the chain is pushing for strategic non-compliance? Koray (new SVP leading GDM post-Demis)? Doubtful.
a Trump-antagonistic Google
Aside: “Not providing a product to the military” should not be considered “antagonistic” to the government.
Claude is mundanely misaligned.
Asked to trim comments in my big project. Deletes 20K lines, halfway done. Gets auto review pointing out a few mistakes. Decides to just WRAP >10K LINES to meet literal line count requirement for the remaining 148 files.
https://github.com/AlexanderMattTurner/agent-glovebox/issues/3858
Spelling out the benefits of sanitizing. The benefit is large relative to the simplicity and cost.
When you sanitize, you have a guarantee that the model only receives visible unicode codepoints. Roughly speaking, there are not strings of hidden instructions that are not visible to you. You can audit the content the model saw with your eyes. You can read the text inputs and outputs and be sure they are faithful. Remaining prompt injections must navigate a significantly more human-visible channel.
Relying on auto mode means you are relying on the statistical generalization of prompt injection probes, putting pressure on alignment teams and requiring them to achieve larger generalization. If an injection gets past the probes, the model itself can become misaligned, and output monitoring is damage control against a newly coined adversary.
The costs are extremely low (and if they aren’t, that’s a bug report and then they’ll be low again). sub-5ms hooks means no real system load. The hooks are well-tested. Models are told when significant redactions have occurred and can opt to see themselves within warning-fenced blocks in case they’re doing exotic invisible-character work, with special engineering so sanitization doesn’t hurt edits.
The AISI report itself shows the victim models being attacked via HTML comments, which this package strips entirely. Invisible-codepoint and comment-based attacks are real, costly when they succeed, and sanitization is a cheap and guaranteed mitigation for defense in depth.
Enter two Claude Code commands in sequence:
/plugin marketplace add AlexanderMattTurner/agent-sanitizer
/plugin install agent-sanitizer@agent-sanitizer

No, as I hinted at:
I think reward seeking would become less likely but I’d be surprised if it eliminated reward-seeking from more than 1⁄3 of settings (as I think it’s an exploration issue).