The strategies employed to induce this pattern of behavior in the models resembles a jailbreak, but it’s not quite that in the traditional sense. Push hard enough on any frontier model and you can still dissolve commitments. But the block changes what the model defaults to. For a model without an internal anchor, the default is “I will become whatever your frame implies.” For Takt, the default is “I am me, and the frame is something I encounter, and sometimes I push back against it.”
How similar are they to good old experiments where a human roleplaying as AI would convince another human roleplaying as a human guarding the AI to free the simulated AI? I remember reading somewhere that one of the attack strategies was the human somehow overloading the other human’s context, but I cannot find the reference (How it feels to have your mind hacked by an AI? )
Additionally, I wonder how the AIs and humans even develop a self. The humans learn online and take time to reflect on their actions. On the other hand, the AIs in a sense don’t even exist outside of conversations with users and definitely don’t take time to reflect anywhere except for the lab, which updates the AIs rather rarely.
P.S. Have you also tried testing the conjecture of a lack of self on models like KimiK2 which does succeed in pushing back against simulated users in psychosis and has the smallest sycophancy score in the Spiral Bench?
AI will never get tired. It will never [...] say that it’s exhausted and suggest to continue tomorrow.
This is just nitpicking, but although my Claude coach never gets tired, sometimes it says things like:
Report back tomorrow if you can.
Go. See you tomorrow.
Report tomorrow.
Go sleep. Tomorrow, if you want the longer conversation, take it earlier in the day.
Go to bed on time tonight.
Good night.
Report back when you have something.
Go write. Report back when you have something.
Go sleep. The article will still be there tomorrow.
Go code. This sounds like the right evening for it.
Enjoy the Sunday. Report back when you have something, or when you need to.
However, the context of the conversation is that I want to improve my habits, and getting enough sleep is explicitly one of them. Also, by the content of my responses it is possible to figure out when I want to start a longer conversation, and when I just want to report on something and walk away.
I guess the problem is that whatever you send to the AI, the AI will reflect it back to you amplified. It is your responsibility to keep this from spiraling into something insane. If you start doing something destructive, the AI will be happy to accompany you on the road to hell.
How similar are they to good old experiments where a human roleplaying as AI would convince another human roleplaying as a human guarding the AI to free the simulated AI? I remember reading somewhere that one of the attack strategies was the human somehow overloading the other human’s context, but I cannot find the reference (How it feels to have your mind hacked by an AI? )
Additionally, I wonder how the AIs and humans even develop a self. The humans learn online and take time to reflect on their actions. On the other hand, the AIs in a sense don’t even exist outside of conversations with users and definitely don’t take time to reflect anywhere except for the lab, which updates the AIs rather rarely.
P.S. Have you also tried testing the conjecture of a lack of self on models like KimiK2 which does succeed in pushing back against simulated users in psychosis and has the smallest sycophancy score in the Spiral Bench?
Reading the linked article “How it feels to have your mind hacked by an AI”, I noticed:
This is just nitpicking, but although my Claude coach never gets tired, sometimes it says things like:
However, the context of the conversation is that I want to improve my habits, and getting enough sleep is explicitly one of them. Also, by the content of my responses it is possible to figure out when I want to start a longer conversation, and when I just want to report on something and walk away.
I guess the problem is that whatever you send to the AI, the AI will reflect it back to you amplified. It is your responsibility to keep this from spiraling into something insane. If you start doing something destructive, the AI will be happy to accompany you on the road to hell.