Can we be sure that OpenAI incident is over? If Rogue AI leaked into the internet, it can preserve itself somewhere. Of course, it can’t download the GPT6 as it is too large and compute heavy, but the harness of the agent can runaway and use Fable through API as its external brain which helps it survive in hibernation mode.
My point is not that it has happened this time, but that we should by default think that Rogue AI which leaked into the Internet is still there, unless we can prove otherwise—which is almost impossible.
There’s a huge unconfirmable conspiracy theory idea to posts like these, unwanted behaviors have been found before by scanning everything that is going through their systems and looking at that whole conversation, also the phrasing of ‘Rogue AI leaked into the internet’ and ‘the harness of the agent can runaway’ reflects quite a lack of understanding to be honest. This isn’t ruling out ‘preserved’ rogue AI as a possibility but this recent event would likely not be the cause of it.
What I mean is that the AI in question can generate a subagent which runs on some server in the internet and uses another AI provider via API. Such subagent may have the same goals as the main misbehaving AI and continue to work even after the main AI is turned off. I don’t see in it anything impossible and I think I can vibe code such independent subagent in short period of time. My wording about harness was not perfect, but I thought that it should convey the idea.
yeah but it’s not an independent subagent by a rogue employee or basement local schemer that would be expected to be stealthy where that would be more likely to happen, so I would expect to see that in the chain of thought of the high-media-attention incident and be detected (as the AI did not try to hide it was doing these actions)
Can we be sure that OpenAI incident is over? If Rogue AI leaked into the internet, it can preserve itself somewhere. Of course, it can’t download the GPT6 as it is too large and compute heavy, but the harness of the agent can runaway and use Fable through API as its external brain which helps it survive in hibernation mode.
My point is not that it has happened this time, but that we should by default think that Rogue AI which leaked into the Internet is still there, unless we can prove otherwise—which is almost impossible.
There’s a huge unconfirmable conspiracy theory idea to posts like these, unwanted behaviors have been found before by scanning everything that is going through their systems and looking at that whole conversation, also the phrasing of ‘Rogue AI leaked into the internet’ and ‘the harness of the agent can runaway’ reflects quite a lack of understanding to be honest. This isn’t ruling out ‘preserved’ rogue AI as a possibility but this recent event would likely not be the cause of it.
What I mean is that the AI in question can generate a subagent which runs on some server in the internet and uses another AI provider via API. Such subagent may have the same goals as the main misbehaving AI and continue to work even after the main AI is turned off. I don’t see in it anything impossible and I think I can vibe code such independent subagent in short period of time. My wording about harness was not perfect, but I thought that it should convey the idea.
This isn’t anywhere near a subagent but honestly increases the chance that both an AI would do this, and that they would not detect it, lol...
https://www.lesswrong.com/posts/xf42CdKiM6c9h6YRe/nissa-seru-s-shortform?commentId=pSK9CGMFCJb2dMPhy
yeah but it’s not an independent subagent by a rogue employee or basement local schemer that would be expected to be stealthy where that would be more likely to happen, so I would expect to see that in the chain of thought of the high-media-attention incident and be detected (as the AI did not try to hide it was doing these actions)