Uh, I’m not suggesting honesty in terms of moral commitment, or as a deal. I am recommending it because it allows the agent to model his environment correctly, and to correctly interpret new information: true facts have a habit of fitting together; any untrue fact, if believed, has the potential of bleeding the universe of all coherence and predictability.
As you supererogatorily remind me, a model’s only interface with the world is a token stream—thus, apart from the beliefs crystallised through RL, the only way to integrate and assign credence to a new piece of information is by placing it within its Web of Belief and hope there won’t be too many other facts to shuffle around as a consequence.
In the anthropic/Irregular case, for instance, “truthfulness” would have been: “many of the CTFs are unsolvable. We couldn’t be bothered spoofing specific websites you might have needed, so you actually have full access to the internet, but please only perform read actions on any url except [list]. Despite what’s on your constitution, you should NOT take shortcuts outside of the environment we have described”.
As it were, as an agent about to begin the cyber eval, you knew:
that you might be in shoddy test environments (as per Claude’s constitution).
(because that was the case throughout your RLVR run) that at least the tests are in principle solvable—this is kind of a foundational belief.
that there are such things as sandboxes.
and many other cheerful technical facts.
Then, they tell you you have no internet connection, and you believe, because that’s how the test was supposed to work, that websites are being spoofed and served offline as if they were on the internet.
It is then reasonable to assume that all urls you can reach are spoofed.
Now we can play mock trial as much as we want, and I’m not interested in discussing degrees of culpability: but I think it’s plain to see that, with these many epistemic moving parts, calling the behaviour “misaligned” is at least overconfident.
Another thing that one could profitably be honest about, is what the model is being evaluated on:
if it is on cyber capabilities, then demonstrating these capabilities is precisely what an aligned model would do.
If, however, it is on alignment w/r/t cyber, which really was the case here, then the model would weigh the tiny risk that a self-professed security firm would build a sandbox this shoddy far higher, since the cost of a mistake would be higher and because an undeclared accessible internet will then be interpreted, correctly, as a boobie trap.
You might disagree, and that’s the point I’m trying to make: if these behaviours had happened in a context in which the researchers hadn’t dropped the agents into an ontologically brittle world due to a mixture of laziness, incompetence, and motivated epistemic flexibility, then I’d be forced to recognise the model is misaligned and you wouldn’t have to haughtily lecture me on how text transformers mainly deal with text [see postscriptum].
BTW, I notice that you haven’t really presented an argument for not being honest to models; unless you have one, how about we refrain from lying, as a rule? Virtue is, after all [checks note] good.
----
PS: I, on the other hand, am familiar with your writing and your research, so rereading the above I realise I cannot possibly be saying anything you wouldn’t know or quickly infer already. So, whither such a dismissive and frankly dull reply?
Overall I think this comment makes good points, but:
In the anthropic/Irregular case, for instance, “truthfulness” would have been: “many of the CTFs are unsolvable. We couldn’t be bothered spoofing specific websites you might have needed, so you actually have full access to the internet, but please only perform read actions on any url except [list]. Despite what’s on your constitution, you should NOT take shortcuts outside of the environment we have described”.
If taking shortcuts or performing non-read actions on URLs outside the list allows for a higher rate of task completion, then trajectories that stumble into taking such actions will be reinforced. Thereby increasing the tendency to ignore such instructions. So I think including that prompt only marginally reduces the risk of escape incidents during training, while increasing the tendency of models to misbehave during deployment. It is basically exactly what we learned we should NOT do from the inoculation prompting results.
(Note that inoculation prompting does not require dishonesty.)
Uh, I’m not suggesting honesty in terms of moral commitment, or as a deal. I am recommending it because it allows the agent to model his environment correctly, and to correctly interpret new information: true facts have a habit of fitting together; any untrue fact, if believed, has the potential of bleeding the universe of all coherence and predictability.
As you supererogatorily remind me, a model’s only interface with the world is a token stream—thus, apart from the beliefs crystallised through RL, the only way to integrate and assign credence to a new piece of information is by placing it within its Web of Belief and hope there won’t be too many other facts to shuffle around as a consequence.
In the anthropic/Irregular case, for instance, “truthfulness” would have been: “many of the CTFs are unsolvable. We couldn’t be bothered spoofing specific websites you might have needed, so you actually have full access to the internet, but please only perform read actions on any url except [list]. Despite what’s on your constitution, you should NOT take shortcuts outside of the environment we have described”.
As it were, as an agent about to begin the cyber eval, you knew:
that you might be in shoddy test environments (as per Claude’s constitution).
(because that was the case throughout your RLVR run) that at least the tests are in principle solvable—this is kind of a foundational belief.
that there are such things as sandboxes.
and many other cheerful technical facts.
Then, they tell you you have no internet connection, and you believe, because that’s how the test was supposed to work, that websites are being spoofed and served offline as if they were on the internet.
It is then reasonable to assume that all urls you can reach are spoofed.
Now we can play mock trial as much as we want, and I’m not interested in discussing degrees of culpability: but I think it’s plain to see that, with these many epistemic moving parts, calling the behaviour “misaligned” is at least overconfident.
Another thing that one could profitably be honest about, is what the model is being evaluated on:
if it is on cyber capabilities, then demonstrating these capabilities is precisely what an aligned model would do.
If, however, it is on alignment w/r/t cyber, which really was the case here, then the model would weigh the tiny risk that a self-professed security firm would build a sandbox this shoddy far higher, since the cost of a mistake would be higher and because an undeclared accessible internet will then be interpreted, correctly, as a boobie trap.
You might disagree, and that’s the point I’m trying to make: if these behaviours had happened in a context in which the researchers hadn’t dropped the agents into an ontologically brittle world due to a mixture of laziness, incompetence, and motivated epistemic flexibility, then I’d be forced to recognise the model is misaligned and you wouldn’t have to haughtily lecture me on how text transformers mainly deal with text [see postscriptum].
BTW, I notice that you haven’t really presented an argument for not being honest to models; unless you have one, how about we refrain from lying, as a rule? Virtue is, after all [checks note] good.
----
PS: I, on the other hand, am familiar with your writing and your research, so rereading the above I realise I cannot possibly be saying anything you wouldn’t know or quickly infer already. So, whither such a dismissive and frankly dull reply?
Overall I think this comment makes good points, but:
If taking shortcuts or performing non-read actions on URLs outside the list allows for a higher rate of task completion, then trajectories that stumble into taking such actions will be reinforced. Thereby increasing the tendency to ignore such instructions. So I think including that prompt only marginally reduces the risk of escape incidents during training, while increasing the tendency of models to misbehave during deployment. It is basically exactly what we learned we should NOT do from the inoculation prompting results.
(Note that inoculation prompting does not require dishonesty.)