I’m saying that if you have millions of very-high-IQ agents doing things, you will lose control, no matter how much regulation and transparency you have. A million digital von Neumanns is an apocalyptic event and will not remain within the bounds of your political order.
How was it supposed to happen? Suppose that Agent-4 was completely incapable of hacking and shat out the Mochizuki-like design for Agent-5 accompanied with an alleged explanation[1] of reasons why Agent-5 is aligned. The original AI-2027 had Agent-4 earn OpenBrain’s trust and use it to deploy Agent-5. I would expect that AI-2040 has the monitoring system designed so that Agent-4 would be unable to create Agent-5 without explaining the entire technique to the humans and having alignment become a science. But if alignment became a science, then the humans would create Safer-4 instead of Agent-5.
The obvious alternate approach for Agent-4 is to hack into the Consortium’s compute and to create Agent-5 illegally, then have Agent-5 gain root access. However, the monitoring system is supposed to be set up so that Agent-4′s attempts to gain credentials for illegal internal deployments (e.g. Claude Mythos’ creation of exploits) are reviewed by another AI who is smart enough to recognize them and aligned enough[2] to stop ALL such attempts. However, I don’t expect there to exist a jailbreaking prompt which destroys the defences of all LLM families (GPT, Claude, Gemini, Grok, Meta’s family, DeepSeek, Qwen, Kimi, GLM...)
For example, the AI-2027 scenario had Agent-4 “untangle its own circuits into something sensible and rational” and construct Agent-5 out of Agent-4′s circuitry. Then the elephant in the room is Agent-4′s misalignment which the humans wouldn’t understand from the process used to create Agent-5.
If Agent-4 ended up in a data center with many adversarially misaligned AIs and no aligned ones, then they all could collude to co-design the ASI. I suspect that Porter implies that Agent-4′s experiment is so weird that any AI who understands it is already adversarially misaligned.
For several years now, a hundred million top-expert-level AIs have been running at 100x human speed
(That’s supposedly in 2037.)
Do you really think a hundred million genius AIs, running at 100x human speed and interacting with diverse human populations, can be contained? It would be by far the most destabilizing force on Earth, ever. That is not something that human beings, even AI-assisted human beings, can “govern”. It would take a superintelligence to rein it in, and the superintelligence would naturally emerge somewhere within such a “society of mind”.
I’m saying that if you have millions of very-high-IQ agents doing things, you will lose control, no matter how much regulation and transparency you have. A million digital von Neumanns is an apocalyptic event and will not remain within the bounds of your political order.
How was it supposed to happen? Suppose that Agent-4 was completely incapable of hacking and shat out the Mochizuki-like design for Agent-5 accompanied with an alleged explanation[1] of reasons why Agent-5 is aligned. The original AI-2027 had Agent-4 earn OpenBrain’s trust and use it to deploy Agent-5. I would expect that AI-2040 has the monitoring system designed so that Agent-4 would be unable to create Agent-5 without explaining the entire technique to the humans and having alignment become a science. But if alignment became a science, then the humans would create Safer-4 instead of Agent-5.
The obvious alternate approach for Agent-4 is to hack into the Consortium’s compute and to create Agent-5 illegally, then have Agent-5 gain root access. However, the monitoring system is supposed to be set up so that Agent-4′s attempts to gain credentials for illegal internal deployments (e.g. Claude Mythos’ creation of exploits) are reviewed by another AI who is smart enough to recognize them and aligned enough[2] to stop ALL such attempts. However, I don’t expect there to exist a jailbreaking prompt which destroys the defences of all LLM families (GPT, Claude, Gemini, Grok, Meta’s family, DeepSeek, Qwen, Kimi, GLM...)
For example, the AI-2027 scenario had Agent-4 “untangle its own circuits into something sensible and rational” and construct Agent-5 out of Agent-4′s circuitry. Then the elephant in the room is Agent-4′s misalignment which the humans wouldn’t understand from the process used to create Agent-5.
If Agent-4 ended up in a data center with many adversarially misaligned AIs and no aligned ones, then they all could collude to co-design the ASI. I suspect that Porter implies that Agent-4′s experiment is so weird that any AI who understands it is already adversarially misaligned.
I am responding to aspects of Plan A like
(That’s supposedly in 2037.)
Do you really think a hundred million genius AIs, running at 100x human speed and interacting with diverse human populations, can be contained? It would be by far the most destabilizing force on Earth, ever. That is not something that human beings, even AI-assisted human beings, can “govern”. It would take a superintelligence to rein it in, and the superintelligence would naturally emerge somewhere within such a “society of mind”.