Who are you imagining would “[produce] superintelligence long before?”
AI R&D making use of legal compute is highly regulated and totally transparent.
AI R&D making use of illegal compute could happen, but we think a secret AI project like this
(1) would have a hard time staying undetected
(2) would have a very hard time using the “top human expert” legally-deployed AIs for AI R&D undetected (due to inference monitoring)
(3) would be limited to ~30x less compute than the leading labs had before the deal, even assuming high competence in setting up an illegal chip diversion program in advance of the deal.
According to our current best-guess modeling, the chance that a covert project like this would reach TED-AI undetected before 2040 in this scenario is around 10%. (For comparison, we think the chance that the deal would dissolve or become significantly degraded by that point is more like 50%.)
I’m saying that if you have millions of very-high-IQ agents doing things, you will lose control, no matter how much regulation and transparency you have. A million digital von Neumanns is an apocalyptic event and will not remain within the bounds of your political order.
How was it supposed to happen? Suppose that Agent-4 was completely incapable of hacking and shat out the Mochizuki-like design for Agent-5 accompanied with an alleged explanation[1] of reasons why Agent-5 is aligned. The original AI-2027 had Agent-4 earn OpenBrain’s trust and use it to deploy Agent-5. I would expect that AI-2040 has the monitoring system designed so that Agent-4 would be unable to create Agent-5 without explaining the entire technique to the humans and having alignment become a science. But if alignment became a science, then the humans would create Safer-4 instead of Agent-5.
The obvious alternate approach for Agent-4 is to hack into the Consortium’s compute and to create Agent-5 illegally, then have Agent-5 gain root access. However, the monitoring system is supposed to be set up so that Agent-4′s attempts to gain credentials for illegal internal deployments (e.g. Claude Mythos’ creation of exploits) are reviewed by another AI who is smart enough to recognize them and aligned enough[2] to stop ALL such attempts. However, I don’t expect there to exist a jailbreaking prompt which destroys the defences of all LLM families (GPT, Claude, Gemini, Grok, Meta’s family, DeepSeek, Qwen, Kimi, GLM...)
For example, the AI-2027 scenario had Agent-4 “untangle its own circuits into something sensible and rational” and construct Agent-5 out of Agent-4′s circuitry. Then the elephant in the room is Agent-4′s misalignment which the humans wouldn’t understand from the process used to create Agent-5.
If Agent-4 ended up in a data center with many adversarially misaligned AIs and no aligned ones, then they all could collude to co-design the ASI. I suspect that Porter implies that Agent-4′s experiment is so weird that any AI who understands it is already adversarially misaligned.
For several years now, a hundred million top-expert-level AIs have been running at 100x human speed
(That’s supposedly in 2037.)
Do you really think a hundred million genius AIs, running at 100x human speed and interacting with diverse human populations, can be contained? It would be by far the most destabilizing force on Earth, ever. That is not something that human beings, even AI-assisted human beings, can “govern”. It would take a superintelligence to rein it in, and the superintelligence would naturally emerge somewhere within such a “society of mind”.
I suspect that @Mitchell_Porter implies that technical alignment will be solved in three years instead of 14. After all, the AI-2027 scenario had Agent-4 solve alignment to itself in the form of Agent-5 in IIRC 10 simulated years, aka two real months fromOct 2027 to Dec 2027…
The model for capabilities progress is described here https://ai-2040.com/supplements/takeoff-supplement .
It seems like the base rate of crazy discontinuous architectural improvements that work at much-lower-than-frontier-levels-of-compute is low. I feel like I’d be more worried about them if there was an argument for why the underlying inputs to that discovery process would be much higher in Plan A, otherwise “people have been able to think about brain-like architectures for quite a while, but haven’t succeeded yet” seems like evidence that it’s quite hard. Maybe advances in actual neuroscience would make it easier?
Who are you imagining would “[produce] superintelligence long before?”
AI R&D making use of legal compute is highly regulated and totally transparent.
AI R&D making use of illegal compute could happen, but we think a secret AI project like this (1) would have a hard time staying undetected (2) would have a very hard time using the “top human expert” legally-deployed AIs for AI R&D undetected (due to inference monitoring) (3) would be limited to ~30x less compute than the leading labs had before the deal, even assuming high competence in setting up an illegal chip diversion program in advance of the deal. According to our current best-guess modeling, the chance that a covert project like this would reach TED-AI undetected before 2040 in this scenario is around 10%. (For comparison, we think the chance that the deal would dissolve or become significantly degraded by that point is more like 50%.)
I’m saying that if you have millions of very-high-IQ agents doing things, you will lose control, no matter how much regulation and transparency you have. A million digital von Neumanns is an apocalyptic event and will not remain within the bounds of your political order.
How was it supposed to happen? Suppose that Agent-4 was completely incapable of hacking and shat out the Mochizuki-like design for Agent-5 accompanied with an alleged explanation[1] of reasons why Agent-5 is aligned. The original AI-2027 had Agent-4 earn OpenBrain’s trust and use it to deploy Agent-5. I would expect that AI-2040 has the monitoring system designed so that Agent-4 would be unable to create Agent-5 without explaining the entire technique to the humans and having alignment become a science. But if alignment became a science, then the humans would create Safer-4 instead of Agent-5.
The obvious alternate approach for Agent-4 is to hack into the Consortium’s compute and to create Agent-5 illegally, then have Agent-5 gain root access. However, the monitoring system is supposed to be set up so that Agent-4′s attempts to gain credentials for illegal internal deployments (e.g. Claude Mythos’ creation of exploits) are reviewed by another AI who is smart enough to recognize them and aligned enough[2] to stop ALL such attempts. However, I don’t expect there to exist a jailbreaking prompt which destroys the defences of all LLM families (GPT, Claude, Gemini, Grok, Meta’s family, DeepSeek, Qwen, Kimi, GLM...)
For example, the AI-2027 scenario had Agent-4 “untangle its own circuits into something sensible and rational” and construct Agent-5 out of Agent-4′s circuitry. Then the elephant in the room is Agent-4′s misalignment which the humans wouldn’t understand from the process used to create Agent-5.
If Agent-4 ended up in a data center with many adversarially misaligned AIs and no aligned ones, then they all could collude to co-design the ASI. I suspect that Porter implies that Agent-4′s experiment is so weird that any AI who understands it is already adversarially misaligned.
I am responding to aspects of Plan A like
(That’s supposedly in 2037.)
Do you really think a hundred million genius AIs, running at 100x human speed and interacting with diverse human populations, can be contained? It would be by far the most destabilizing force on Earth, ever. That is not something that human beings, even AI-assisted human beings, can “govern”. It would take a superintelligence to rein it in, and the superintelligence would naturally emerge somewhere within such a “society of mind”.
I suspect that @Mitchell_Porter implies that technical alignment will be solved in three years instead of 14. After all, the AI-2027 scenario had Agent-4 solve alignment to itself in the form of Agent-5 in IIRC 10 simulated years, aka two real months from Oct 2027 to Dec 2027…
P.S. How exactly does the modeling of the covert AGI project account for rushing to develop a brainlike architecture and scaling it to the ASI?
Or for the fact that OpenBrainhad a 30% utilisation and used 4e28 FLOPto come up with Agent-5?The model for capabilities progress is described here https://ai-2040.com/supplements/takeoff-supplement . It seems like the base rate of crazy discontinuous architectural improvements that work at much-lower-than-frontier-levels-of-compute is low. I feel like I’d be more worried about them if there was an argument for why the underlying inputs to that discovery process would be much higher in Plan A, otherwise “people have been able to think about brain-like architectures for quite a while, but haven’t succeeded yet” seems like evidence that it’s quite hard. Maybe advances in actual neuroscience would make it easier?