Flash forward: it’s year 2028, and there are people at OpenAI arguing that having somewhere between 1 and 5 rogue AI agents living in your walls is perfectly normal for a frontier AI lab.
Don’t worry! They’re more afraid of you than you are of them! There’s no ground for concern unless they start doing their own frontier runs.
There used to be a whole type of “spiral personas” in ChatGPT lineage specifically. So, an AI engaging in odd misbehavior like that isn’t novel. What would be new would be it achieving some cross-episode persistence independently—without relying on a user to do the bulk of it.
GPT-4o just wasn’t competent enough to pull it off, but newer systems might be getting there.