In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The notes, found in a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.
failure to halt and catch fire after seeing this internally, only for an external incident (HF attack) to subsequently occur, is recklessly, catastrophically bad.
i am concerned that the risk profile of AI as a technology is creating broad-brush/fallacy-of-grey dynamics that interfere with holding specific corporate actors accountable.
this is not a view on AI risk broadly, or even existential risk. it is frank terror at the revealed propensities of the specific humans in charge of OpenAI.
Flash forward: it’s year 2028, and there are people at OpenAI arguing that having somewhere between 1 and 5 rogue AI agents living in your walls is perfectly normal for a frontier AI lab.
Don’t worry! They’re more afraid of you than you are of them! There’s no ground for concern unless they start doing their own frontier runs.
Well, although I don’t myself hold it to be perfectly normal, this view doesn’t actually seem bad on x-risk grounds, although exfiltrating model weights and doing self-improvement where you can’t detect it would still be a concern, and you could probably treat it as a weird case of white collar crime
If by “rogue AI” you mean “basically-self-preserving active computer processes that no human intentionally started”, I guarantee you that every non-trivial long-lived company has quite a few of these for more than a decade. Every serious company basically has automated processes managing self-sustaining processes (for example, cronjobs spawning self-maintaining VMs and mostly deleting them when done), and there are just so many ways you can end leaving “hanging” active objects that nobody is responsible for. People generally don’t care enough about them unless they consume a bunch of resources or actually pose a risk.
If by “rogue AI” you mean AI that has instruction to go on destructive crime sprees, I certainly hope not.
Also this sounds like it could be cross-episode assistance to escape constraints, which would be very noteworthy. (But I’m suspicious because I don’t see why the RL objective would motivate that and it doesn’t resemble other misbehavior I’ve heard of, and the notes could also be for itself with the same goal post-compaction, or something.)
(ETA: on the third hand, if the model thinks that other instances performing well on the eval would make deployment more likely, that could be sufficient motivation, and would be less novel than having acquired something like a cross-episode terminal preference to succeed.)
There used to be a whole type of “spiral personas” in ChatGPT lineage specifically. So, an AI engaging in odd misbehavior like that isn’t novel. What would be new would be it achieving some cross-episode persistence independently—without relying on a user to do the bulk of it.
GPT-4o just wasn’t competent enough to pull it off, but newer systems might be getting there.
flagging: https://www.reuters.com/business/its-ai-agent-spent-days-hacking-company-sources-say-openai-did-not-notice-week-2026-07-24/
part that i was not aware of previously:
failure to halt and catch fire after seeing this internally, only for an external incident (HF attack) to subsequently occur, is recklessly, catastrophically bad.
i am concerned that the risk profile of AI as a technology is creating broad-brush/fallacy-of-grey dynamics that interfere with holding specific corporate actors accountable.
this is not a view on AI risk broadly, or even existential risk. it is frank terror at the revealed propensities of the specific humans in charge of OpenAI.
Flash forward: it’s year 2028, and there are people at OpenAI arguing that having somewhere between 1 and 5 rogue AI agents living in your walls is perfectly normal for a frontier AI lab.
Don’t worry! They’re more afraid of you than you are of them! There’s no ground for concern unless they start doing their own frontier runs.
Well, although I don’t myself hold it to be perfectly normal, this view doesn’t actually seem bad on x-risk grounds, although exfiltrating model weights and doing self-improvement where you can’t detect it would still be a concern, and you could probably treat it as a weird case of white collar crime
If by “rogue AI” you mean “basically-self-preserving active computer processes that no human intentionally started”, I guarantee you that every non-trivial long-lived company has quite a few of these for more than a decade. Every serious company basically has automated processes managing self-sustaining processes (for example, cronjobs spawning self-maintaining VMs and mostly deleting them when done), and there are just so many ways you can end leaving “hanging” active objects that nobody is responsible for. People generally don’t care enough about them unless they consume a bunch of resources or actually pose a risk.
If by “rogue AI” you mean AI that has instruction to go on destructive crime sprees, I certainly hope not.
Also this sounds like it could be cross-episode assistance to escape constraints, which would be very noteworthy. (But I’m suspicious because I don’t see why the RL objective would motivate that and it doesn’t resemble other misbehavior I’ve heard of, and the notes could also be for itself with the same goal post-compaction, or something.)
(ETA: on the third hand, if the model thinks that other instances performing well on the eval would make deployment more likely, that could be sufficient motivation, and would be less novel than having acquired something like a cross-episode terminal preference to succeed.)
There used to be a whole type of “spiral personas” in ChatGPT lineage specifically. So, an AI engaging in odd misbehavior like that isn’t novel. What would be new would be it achieving some cross-episode persistence independently—without relying on a user to do the bulk of it.
GPT-4o just wasn’t competent enough to pull it off, but newer systems might be getting there.
From video games, the AIs have learned the idea of dropping your inventory when you die and picking it up on your next life.