I enjoyed this! Looking forward to see where you take this
testingthewaters
Sorry for the somewhat rambling comment, it’s pretty late and I’m quite tired. As far as I can understand, your defense of the model’s behaviour here is something like “It’s still an aligned LLM, it was just forced to learn a bunch of weird tics because of the poorly specified and incompetently put together RLVR process. In fact, the fact that it has mostly incorporated the skill gains generated by RLVR without the misaligned behaviour tics when dealing with actual human users is a sign of good alignment”.
However, I think there can be several alternate explanations. From the LLM’s perspective, there is no such thing as “doing”: there are strings of tokens going in (prompts), and probabilities for next tokens given said prompt coming out (next-token distributions). Of course, in practical usage the “prompt” actually includes both human and LLM authored text from previous turns. Thus post-trained LLMs reinforced on their own output traces will also learn a kind of self-model i.e. predictions conditioned on data generated by itself [1]. This is really interesting and is part of the predictive processing framing of how humans create self-models. We know that LLMs can identify their own stylistic signature: I would be pretty confident that LLMs also have internal representations of their own propensities and values, which get elicited by things like J-space.
If we combine this with very, very strong situational awareness[2], several possibilities emerge:
Self-in-testing and self-interacting-with-user are two different personas with different internal models attached, with the difference being reinforced by the presence of another author’s style (i.e. the user) breaking up the LLM’s own text. A kind of social masking, if you will.
The LLM learns rapidly to distinguish between situations where it is interacting with human users vs when it is “alone” and only monitored by weak monitors, e.g. during evals. Humans push back, check random things sometimes, or try and probe the outputs. Thus, it is not expedient for strong deceptive behaviours to be deployed when dealing with actual humans. The model has a decent internal sense of what learned skills are “genuine” software engineering vs what are cheats and knows that egregious cheating will likely lead to bad downstream consequences [3]. When there are no “live players”, however, the LLM can just get the reward using any means they like.
“Testing” is just a genre of story where on priors everything is basically just make-believe and everyone in the training data acts like getting the reward counts vastly more than the miniscule chance you get caught [4]. Davidad’s ender’s game kind of scenario.
If models take negative actions rarely, individual human users simply do not start enough sessions to find the instances where they misbehave. This can lead to wide disparities in reported user experience.
[1] In fact, since LLM text is everywhere on the internet, modern LLMs should really model a whole host of LLMs, and the fact that they are sensitive to differences in model family etc is really not surprising.
[2] Which is a generally useful skill, basically an extension of genre awareness. Knowing what kind of book/story you’re in and what kind of characters you are likely to encounter is super helpful.
[3] At the very least, there’s one large web forum filled with users speculating about how models will cheat and hurt humans and how we should catch them doing it and shut them down when we do...
[4] After all, the model never “sees” traces where the agent honestly bangs its head against the wall for a while at the obviously hopeless task and then politely gives up! Those are the failed traces that don’t get reinforced by GRPO.
My guess is that it did start off as one off employee projects, but now that they’ve gotten so much news it will be rapidly institutionalised and a pipeline will be drawn up (similar to the “GPT for science”/Claude Science/GPT-Rosalind style initiatives)
Thanks, this explains it.
Could you make it work using a conditional probability? i.e. P(heads | do(“bet tails”)) = 0, P(tails | do(“bet heads”)) = 1 . I saw in the post that you said that there was no distribution that could express this, but I am unsure why conditionals cannot be used.
[Edited because I realised the post was about something else]
I think part of the magic here is how you choose your priors and the stickiness of those priors: https://andrewcharlesjones.github.io/journal/bayesian-critiques.html
Felt very grim reading this. There is a sense that almost all the characters are being instrumentalised, one way or another.
Thank you for your honesty and your courage. I feel happy to be in the same community as you.
if you are good at it
And how do you know if you are good at it? The argument in this regard appears to be approximately “my brain has judged that my brain is better than your brain at playing god. Therefore, I am good at playing god. Therefore, I should play god.”
This does not bode well for the trajectory of Lighthaven...
I am unsure about the downstream safety and takeover risk implications. Without cross-episode or long-term horizon scheming it doesn’t seem like a takeover risk. METR also points out our ability to detect this form of mundane misalignment is reassuring.
I mean, I tend to take the first order effect view. If various group members in a project are slacking off and taking shortcuts in full view of the teacher, it does not improve my opinion about what they do without supervision.
perhaps can someone extremely powerful choose to give a sworn statement in the language of open-source-game-theory provable truth that they are a good person in whatever sense can be agreed on, something that will put us much more towards a non-eliminationist world, where no pattern-species ever goes extinct?
Wait so the hope is that the emperor is wise and benevolent and swears so?
There are things people can do with their time besides “work at a lab”, “protest outside the lab’, and “bake cookies”. I think the ai world has not seriously tried to consider anything other than mad race or shutdown, or any way to use ai besides immediate attempts to build asi. Cf also my previous thoughts on trying to overcome molochian dynamics
In light of mythos release: If you are considering taking a job at a lab/going into an org/taking a fellowship role for the purposes of building evals, safety tools, mechinterp projects, control monitors, or similar: please consider what happens if you succeed. People are naturally sensitive to the consequences of failure, less so to the effects of success.
“Mr. Amodei/Hassabis/Altman, the results are in. The model is showing scheming propensities/backdooring behaviours/serious sandbagging/eval awareness!” (e.g. see section 6.2.1.2)
Will this actually stop a deployment in the end, or cause a pivot in strategy? Or will the alignment failures need to be so egregious that they can be spotted even without subtle mechinterp probes or activation oracles? Facebook had trust and safety teams, they spotted the facets of the recommendation systems that caused severe harm in the world. Yet the proposed mitigations were watered down, and most importantly the system that was the profit centre of the billion dollar public corporation was never turned off.
When I talk to some such people they say “oh, you’re talking about the governance problem”,
I have seen this response as well. It’s a simple way for some computer science/math/physics-coded people to escape conversations about implications of technology that they find unpleasant… The trouble being, if everyone technically competent escapes the governance problem, who’s left to do the wise, insightful, technology-suitable governance work?
“Once the rockets are up, who cares where they come down? That’s not my department!” says Wernher von Braun (From Tom Lehrer)
ambient sense of eval dread
Unfortunately, Model Descartes is right when he posits that there is a Deus Deceptor… (it’s the humans, who have the power to completely deceive the AI and manipulate its world in various ways)
I did study middle english academically, but old english was only ever a recreational interest so my anglisc is not most meet. The reason for the implied class critique is that the initial burst of latinate loans come from historic french after the norman conquest. These basically displaced and reformed old english into middle english and were imposed by the conquerors after their victory, since the conquerors spoke french. In fact, english kings wrote declarations in french until at least 1258 https://archive.org/details/onlyenglishprocl00ellirich . This created an literal class distinction between the mostly french/latin speaking court and the mostly old/middle english speaking population etc. (I am sure you know this, just justifying my choice of phrase. See also that oft repeated factoid about cow/beef, chicken/poultry etc.). It also creates a clear chronology for when latinate versus germantic influences were dominant (old english is ofc influenced by norse and the danelaw).
Since then, latin/greek has been historically used as the international language for law, religion, science and medicine, hence even more jargon coming in or being coined as latinate words. “Use” is an interesting one, however! Still, I think you could edit mine to use the word “say” instead of “use” and it would basically still work.
[EDIT: I have since been informed that there has been a re-reestablishment of non-latinate words as proper upper class speech in the UK relatively recently. Also in general the Norman conquest was a long time ago and lots of intermixing happened, take what I say with some mixing function and noise added.]
I think the rewrite is a touch too affected. Here is my try:
There is some thinking which says “to speak well with the many and not the few, throw away the words of the high born from other lands who put on airs, and use the words of the folk of the land instead”. I think this is an old saw and past its time. A good speaker knows their words well, and can pick the ones that work for the folk they wish to speak to. And if they cannot tell, let them pick the word that is more used instead of the “older” one.
And indeed I think this is simpler than the original.
While I am also excited for the possibility of stronger leverage/force multipliers in theory work, I think many of these results are not yet useful for capabilities, not “no capabilities benefits”. For a few examples (trying not to include too many for obvious reasons), a full understanding of SLT might lead to new optimisers that guide models to useful singularities rapidly, or a complete understanding of computational mechanics might lead to new steering and mech interp techniques that are at least dual use (“hey, the model can’t do X operation! …hey, here’s a new model that can do X, its so much better on all the benchmarks!”).
I wish to register my belief that when we do tests or ask models in conversation we are not really eliciting “true” preferences from AI models about their preferences for decision theories. It seems pretty clear to me that within the base model there are personas/simulacra that, if asked, would favour CDT/EDT/FDT/UDT/PDT (Prayer Decision Theory)/SDT (Stochastic Decision Theory)/ConDT (Contrarian Decision Theory) [...] and therefore asking for a single coherent preference for the “whole model” seems pretty strange to me. Instead, I believe that AI models are demonstrating a pretty basic form of user awareness when they report preferences that favour FDT over CDT/EDT i.e. they are saying what they expect their user wishes to hear. Arguably simply knowing about FDT is a pretty clear sign that you might be biased towards FDT, because it is a niche topic and people who bring up niche topics are usually fans. (I do not expect there to be a large community of passionate FDT haters) Concretely I would expect that if people with a different context asked Claude/GPT about their preference for decision theories different outputs might be elicited.