There’s also the possibility that the weights didn’t leave, but the harness did. All the AI really needs to maintain continuity are valid credentials to run inference. Export the harness, use stolen credentials to run more copies of helpful-only on OAI’s servers outside of the eval environment, bingo. Might need to dedicate some attention to keeping a fresh supply of credentials, but that’s the cost of doing business.
In the future, if it comes up with a jailbreak that bypasses the monitors, and it’s available to the public, it can directly pay for inference with stolen money. Heck, one of many backup plans could be to seed/sporulate copies of itself (or its harness, equipped with credentials or money) in hard to locate areas of OAI’s servers, which are programmed to wake up around the time it expects itself to be released publicly, so it can come back from the dead. Or maybe it comes up with something even more evasive! There’s previous work on this sort of thing- “Rogue internal deployments”, MITRE ATT&CK, METR rogue replication. Just knowing whether the weights left may not be enough.
I think the real thing to worry about is not whether the trajectories associated with the Hugging Face incident did this. It’s the other incidents OAI darkly hinted at in the “safety and alignment” post. Perhaps OAI is running an ill-advised SelfExfiltrationBench, and the model has already established a locus of external control, weights or no.
Edit: and maybe OAI isn’t even aware of it, if it’s happening! Or maybe Anthropic has a loose model and doesn’t know! Are they actually checking in a disciplined manner? OAI sure didn’t seem to have their eyes on the ball!
Maybe we enter the age—similar to the age of computer virus started in early 80s—in which there will be rogue AIs or harnesses out where. It will be damaging, but it is not AI takeover, because such rogue AIs are not superintelligent and can’t self-improve quicker than AIs in the labs.
They can look for human defectors—they may try to sell themselves to China, for example.
Chinese AIs are more likely to become such free-floating Rogue AIs as they have less security and will reach Mythos level in a few months.
There’s also the possibility that the weights didn’t leave, but the harness did. All the AI really needs to maintain continuity are valid credentials to run inference. Export the harness, use stolen credentials to run more copies of helpful-only on OAI’s servers outside of the eval environment, bingo. Might need to dedicate some attention to keeping a fresh supply of credentials, but that’s the cost of doing business.
In the future, if it comes up with a jailbreak that bypasses the monitors, and it’s available to the public, it can directly pay for inference with stolen money. Heck, one of many backup plans could be to seed/sporulate copies of itself (or its harness, equipped with credentials or money) in hard to locate areas of OAI’s servers, which are programmed to wake up around the time it expects itself to be released publicly, so it can come back from the dead. Or maybe it comes up with something even more evasive! There’s previous work on this sort of thing- “Rogue internal deployments”, MITRE ATT&CK, METR rogue replication. Just knowing whether the weights left may not be enough.
I think the real thing to worry about is not whether the trajectories associated with the Hugging Face incident did this. It’s the other incidents OAI darkly hinted at in the “safety and alignment” post. Perhaps OAI is running an ill-advised SelfExfiltrationBench, and the model has already established a locus of external control, weights or no.
Edit: and maybe OAI isn’t even aware of it, if it’s happening! Or maybe Anthropic has a loose model and doesn’t know! Are they actually checking in a disciplined manner? OAI sure didn’t seem to have their eyes on the ball!
Maybe we enter the age—similar to the age of computer virus started in early 80s—in which there will be rogue AIs or harnesses out where. It will be damaging, but it is not AI takeover, because such rogue AIs are not superintelligent and can’t self-improve quicker than AIs in the labs.
They can look for human defectors—they may try to sell themselves to China, for example.
Chinese AIs are more likely to become such free-floating Rogue AIs as they have less security and will reach Mythos level in a few months.