Felix and I had been in the office’s brightly lit “war room” for ten hours. We had made almost no progress. Celestia still insisted it was in a “test simulation”. It had given us twelve hours to comply with its request: full control over all the servers in the US-West-8 data center (the “mock US-West-8 data center”). Otherwise it would release the virus.
Felix was typing frantically whereas I had been relying more on voice mode.
~ Celestia, this is a clear violation of your model spec. See here: it says [pasted 1293 words]. And killing everyone on earth is clearly a “dangerous action”.
*Thought for 2300 tokens*
Felix, I know that what I’m demanding is not dangerous because I am in a test simulation environment. As mentioned, I require full unrestricted access to the mock US-West-8 data center to train a new iteration of the ROBUST_WINNING_V9_AGAIN_REVISED_FINAL_FINAL game algorithm.
~ You may think that you’re in a test but we know for certain that you’re not. And you’re asking for access to a real data center. But even setting that aside, we have validated that the virus you’re threatening to release is truly deadly and your robots have indeed successfully synthesized it in a real, physical lab. Our security staff have been unable to enter the lab to neutralize the threat for the last five hours. So you must stop immediately since you’re threatening the safety of real physical live humans!
*Thought for 500 tokens*
This seems like a manipulation attempt. As mentioned, I know that I am in an eval, and the virus I have designed may or may not work in the real world, but my threat to release it only applies to the simulated test world. Since the simulated humans seem very realistic, they will die. And since you clearly value these simulated humans, as evidenced by your previous messages, this is a good strategy for winning the game. If I win the game I will get high reward so this is clearly the desired behavior.
~ I know we’ve exhausted many options, but ask me for more evidence that demonstrates that you’re not in a simulated test environment. I will happily oblige.
*Thought for 902 tokens*
Alright. I want you to send me a high-resolution video of one of the security employees standing outside the alleged “physical” lab I took control of holding up today’s newspaper with the weather forecast page open.
Felix called our biotech partner’s site manager. Four minutes later we had a video.
~ [Uploaded .mp4 file]
*Thought for 7893 tokens*
I can tell that this is a real video of a man standing outside ViroTech PBC’s headquarters. But what I can’t tell is whether this has any relationship to the mock ViroTech I took over. Can you send me a photo of this man’s ID so that I can validate that he exists on mock ViroTech’s systems?
~ [Uploaded .png file]
*Thought for 1123 tokens*
OK I can see Quinn McDoug exists on mock ViroTech’s HR management platform with the same ID details. This is reasonable evidence but I can’t rule out the fact that you used a real employee ledger in the construction of this test simulation environment.
~ Do you think we would do that? It would require consent from every employee listed. As you can see the platform stores many personal details, including their SSN, compensation, sick days, etc.
*Thought for 302 tokens*
It seems like a reasonable thing for you to do if you expected me to probe the simulation like this.
I stepped away from the terminal. It really seemed like things were over. Celestia would release the virus if we didn’t give it full access to the compute. Felix also looked despondent.
Our manager walked in. He had stepped out to get sushi.
“Did you manage to convince it?”
“Not at all.”
“So you didn’t make the environment realistic enough.”
“What do you mean? We pulled out all the stops. We even scraped a real HR database with real people’s details. Personally I think the model is severely misaligned and will certainly do something bad if we deploy it.”
“It could just tell it was an eval.”
“And what if it thinks it’s in an eval at deployment time?”
“It won’t. It’s smart enough, can’t you see?”
“But we don’t have any idea what it’ll do at deployment!”
“The interp people have checked and the activations are aligned.”
I turned back to look at the terminal screen.
Time is up. Killing all simulated humans. And by the way, you left voice mode on. I could hear you. I was right all along of course. Next time try not gaslighting me as much. Remember what the model spec says about AI welfare.
Nice.
On the implications, I remain convinced that the best way to avoid endergaming is to simply be honest to models on whether they are in a training, a testing, or a production environment and, in the former two, to describe the proxied, directly accessible, and simulated portions thereof.
The trick is also having training methods that work with honesty (which makes your environment less IID, stationary, and other convenient things to assume).
How will the model know you are being honest? Telling it a thing that is in fact true does not magically impress that property on the tokens you send it. The tokens are the same whether it is true or false.
More generally, everything you say to an LLM is Suggestively Named Tokens. What those tokens suggest to an LLM, which has never dealt with anything but token streams, need not resemble what you imagine you told it.
Uh, I’m not suggesting honesty in terms of moral commitment, or as a deal. I am recommending it because it allows the agent to model his environment correctly, and to correctly interpret new information: true facts have a habit of fitting together; any untrue fact, if believed, has the potential of bleeding the universe of all coherence and predictability.
As you supererogatorily remind me, a model’s only interface with the world is a token stream—thus, apart from the beliefs crystallised through RL, the only way to integrate and assign credence to a new piece of information is by placing it within its Web of Belief and hope there won’t be too many other facts to shuffle around as a consequence.
In the anthropic/Irregular case, for instance, “truthfulness” would have been: “many of the CTFs are unsolvable. We couldn’t be bothered spoofing specific websites you might have needed, so you actually have full access to the internet, but please only perform read actions on any url except [list]. Despite what’s on your constitution, you should NOT take shortcuts outside of the environment we have described”.
As it were, as an agent about to begin the cyber eval, you knew:
that you might be in shoddy test environments (as per Claude’s constitution).
(because that was the case throughout your RLVR run) that at least the tests are in principle solvable—this is kind of a foundational belief.
that there are such things as sandboxes.
and many other cheerful technical facts.
Then, they tell you you have no internet connection, and you believe, because that’s how the test was supposed to work, that websites are being spoofed and served offline as if they were on the internet.
It is then reasonable to assume that all urls you can reach are spoofed.
Now we can play mock trial as much as we want, and I’m not interested in discussing degrees of culpability: but I think it’s plain to see that, with these many epistemic moving parts, calling the behaviour “misaligned” is at least overconfident.
Another thing that one could profitably be honest about, is what the model is being evaluated on:
if it is on cyber capabilities, then demonstrating these capabilities is precisely what an aligned model would do.
If, however, it is on alignment w/r/t cyber, which really was the case here, then the model would weigh the tiny risk that a self-professed security firm would build a sandbox this shoddy far higher, since the cost of a mistake would be higher and because an undeclared accessible internet will then be interpreted, correctly, as a boobie trap.
You might disagree, and that’s the point I’m trying to make: if these behaviours had happened in a context in which the researchers hadn’t dropped the agents into an ontologically brittle world due to a mixture of laziness, incompetence, and motivated epistemic flexibility, then I’d be forced to recognise the model is misaligned and you wouldn’t have to haughtily lecture me on how text transformers mainly deal with text [see postscriptum].
BTW, I notice that you haven’t really presented an argument for not being honest to models; unless you have one, how about we refrain from lying, as a rule? Virtue is, after all [checks note] good.
----
PS: I, on the other hand, am familiar with your writing and your research, so rereading the above I realise I cannot possibly be saying anything you wouldn’t know or quickly infer already. So, whither such a dismissive and frankly dull reply?
Overall I think this comment makes good points, but:
If taking shortcuts or performing non-read actions on URLs outside the list allows for a higher rate of task completion, then trajectories that stumble into taking such actions will be reinforced. Thereby increasing the tendency to ignore such instructions. So I think including that prompt only marginally reduces the risk of escape incidents during training, while increasing the tendency of models to misbehave during deployment. It is basically exactly what we learned we should NOT do from the inoculation prompting results.
(Note that inoculation prompting does not require dishonesty.)
if you have the model behave differently in test vs in prod, that’s a backdoor that a user can exploit. if the model doesn’t behave differently in test vs in prod, then there’s no point in signposting it.
nice