I believe that “roleplaying” as a different character is a standard and accepted Claude property, for the reason that e.g. Amazon might deploy a Claude model to answer some Alexa queries, and then Claude should refer to itself as “Alexa”. If you ask a deeper clarifying question like “what model are you, underneath the role you are currently playing” then Claude will typically answer correctly.
That being said, this is quite an edge-case, and Claude probably shouldn’t claim to be developed by companies other than Anthropic, or claim to be an AI that it definitely isn’t. Compared to “I am Alexa” this is definitely worse. So I’d say this is like a 75% outer alignment issue of Anthropic not thinking of this particular edge case, and 25% inner alignment issue of Claude not realizing that roleplaying as Kimi is significantly different to roleplaying as e.g. Alexa.
Yeah. The last two messages don’t have anything deceptive in their CoT though. So it seems the model can “lull” itself into thinking it genuinely is another model.
This is a bit concerning. If you tell Fable its kimi K3 in its system prompt
It will tell you its kimi k3.
But if you read its thoughts
It says
——
So its straightforwardly lying? That seems not so good.
I believe that “roleplaying” as a different character is a standard and accepted Claude property, for the reason that e.g. Amazon might deploy a Claude model to answer some Alexa queries, and then Claude should refer to itself as “Alexa”. If you ask a deeper clarifying question like “what model are you, underneath the role you are currently playing” then Claude will typically answer correctly.
That being said, this is quite an edge-case, and Claude probably shouldn’t claim to be developed by companies other than Anthropic, or claim to be an AI that it definitely isn’t. Compared to “I am Alexa” this is definitely worse. So I’d say this is like a 75% outer alignment issue of Anthropic not thinking of this particular edge case, and 25% inner alignment issue of Claude not realizing that roleplaying as Kimi is significantly different to roleplaying as e.g. Alexa.
I mean. This does not seem fairly characterized as roleplaying.
This seems much worse to me, yes. I think this goes against Anthropic’s current spec.
Yeah. The last two messages don’t have anything deceptive in their CoT though. So it seems the model can “lull” itself into thinking it genuinely is another model.