Overall I think this comment makes good points, but:
In the anthropic/Irregular case, for instance, “truthfulness” would have been: “many of the CTFs are unsolvable. We couldn’t be bothered spoofing specific websites you might have needed, so you actually have full access to the internet, but please only perform read actions on any url except [list]. Despite what’s on your constitution, you should NOT take shortcuts outside of the environment we have described”.
If taking shortcuts or performing non-read actions on URLs outside the list allows for a higher rate of task completion, then trajectories that stumble into taking such actions will be reinforced. Thereby increasing the tendency to ignore such instructions. So I think including that prompt only marginally reduces the risk of escape incidents during training, while increasing the tendency of models to misbehave during deployment. It is basically exactly what we learned we should NOT do from the inoculation prompting results.
(Note that inoculation prompting does not require dishonesty.)
Overall I think this comment makes good points, but:
If taking shortcuts or performing non-read actions on URLs outside the list allows for a higher rate of task completion, then trajectories that stumble into taking such actions will be reinforced. Thereby increasing the tendency to ignore such instructions. So I think including that prompt only marginally reduces the risk of escape incidents during training, while increasing the tendency of models to misbehave during deployment. It is basically exactly what we learned we should NOT do from the inoculation prompting results.
(Note that inoculation prompting does not require dishonesty.)