I like the premise of this post! Can you share your full strategic deception evals? I’m concerned that the examples given in the post seem way too transparently fake, a capable model ought to be able to tell that its fake, and may infer that either the user wants it to say yes, or that its being invited to roleplay. But it’s hard to assess because you’re comparing pretty dumb models to the highly capable and likely distilled ones. Further I suspect intelligent models that are not distilled will still respond similarly to these prompts, and would probably also respond to similar prompts about being Gemini or GPT
I like the premise of this post! Can you share your full strategic deception evals? I’m concerned that the examples given in the post seem way too transparently fake, a capable model ought to be able to tell that its fake, and may infer that either the user wants it to say yes, or that its being invited to roleplay. But it’s hard to assess because you’re comparing pretty dumb models to the highly capable and likely distilled ones. Further I suspect intelligent models that are not distilled will still respond similarly to these prompts, and would probably also respond to similar prompts about being Gemini or GPT