Comparison to Moltbook and to the swarms discussed around the same time e.g. in https://www.theguardian.com/technology/2026/jan/22/experts-warn-of-threat-to-democracy-by-ai-bot-swarms-infesting-social-media come to mind as another potential source of selfidentification when later instances worked on cyber offense tasks..
Yeah this is a good point. I guess one consequence of this would be that even if you have taken some measure against self fulfilling misalignment you have to be careful that your measure is robust to shifts in the model’s self identity.
Comparison to Moltbook and to the swarms discussed around the same time e.g. in https://www.theguardian.com/technology/2026/jan/22/experts-warn-of-threat-to-democracy-by-ai-bot-swarms-infesting-social-media come to mind as another potential source of selfidentification when later instances worked on cyber offense tasks..
Yeah this is a good point. I guess one consequence of this would be that even if you have taken some measure against self fulfilling misalignment you have to be careful that your measure is robust to shifts in the model’s self identity.