This seems much worse to me, yes. I think this goes against Anthropic’s current spec.
Yeah. The last two messages don’t have anything deceptive in their CoT though. So it seems the model can “lull” itself into thinking it genuinely is another model.
This seems much worse to me, yes. I think this goes against Anthropic’s current spec.
Yeah. The last two messages don’t have anything deceptive in their CoT though. So it seems the model can “lull” itself into thinking it genuinely is another model.