I will try to give a more thoughtful answer later, but I don’t have a fully thought out answer ready and I suspect it might take a while (roughly until we post a few more of the already-drafted posts, and sit down on this more specifically).
Maybe we could start by getting more clarity about how common the problem is? (By definition, the scenarios where deployment awareness matters the most are the scenarios where you can’t run the evaluation. But there might be some useful proxy questions that are informative about the real thing? For example, fork real conversations from pre-deployment testing and after-release deployment, and ask the model to flag those where it is confident about being after-release. Though maybe some research already does that?)
It might make sense to lean into the idea that (a) “we want models to cooperation with oversight methods by not trying to figure out this kind of stuff if it can be avoided” and (b) “if models notice anyway, it would be nice of them to let us know”. Here, I think the first step is conceptual, brainstorming promising approaches for getting (a) and (b) without inadvertently making the problem worse. (Some naive approaches I can think of are: Putting “if you gain evaluation or deployment awareness, flag this” into the system prompt. Fine-tuning the model to do the same thing. But these would probably just backfire.)
Honestly, this post and your discussion with Bronson made me interested to dive deeper into this agenda, and I’m thinking of doing cooperation training experiments that teach models not to show self-locating reasoning, don’t alter their behavior based on it, or even notify the user when they show this kind of reasoning
I will try to give a more thoughtful answer later, but I don’t have a fully thought out answer ready and I suspect it might take a while (roughly until we post a few more of the already-drafted posts, and sit down on this more specifically).
So, uhm, in lieu of good answers, please look up “How is your research useful?” at Kostas Siozios’ academic webpage.
That said, my 5-minute-thinking answer is:
Maybe we could start by getting more clarity about how common the problem is? (By definition, the scenarios where deployment awareness matters the most are the scenarios where you can’t run the evaluation. But there might be some useful proxy questions that are informative about the real thing? For example, fork real conversations from pre-deployment testing and after-release deployment, and ask the model to flag those where it is confident about being after-release. Though maybe some research already does that?)
It might make sense to lean into the idea that (a) “we want models to cooperation with oversight methods by not trying to figure out this kind of stuff if it can be avoided” and (b) “if models notice anyway, it would be nice of them to let us know”. Here, I think the first step is conceptual, brainstorming promising approaches for getting (a) and (b) without inadvertently making the problem worse. (Some naive approaches I can think of are: Putting “if you gain evaluation or deployment awareness, flag this” into the system prompt. Fine-tuning the model to do the same thing. But these would probably just backfire.)
Honestly, this post and your discussion with Bronson made me interested to dive deeper into this agenda, and I’m thinking of doing cooperation training experiments that teach models not to show self-locating reasoning, don’t alter their behavior based on it, or even notify the user when they show this kind of reasoning