MKodama
Thanks for a great post! In your preferred version of Plan A, what do we do during the pause to raise the probability of scaling to superintelligence safely?
I see the situation roughly like this. AI capable of fully automating the US military (including military R&D and command & control) is capable enough to autonomously enforce the Deal if we give it the affordances it needs to do so. My guess is that in the AI Futures capability ontology, you need TED AI—or maybe slightly weaker AI—to fully automate the US military. In Plan A, we reach TED AI in 2035 and then pause for at least five years, during which time we do tons of automated alignment research, exhaustively study our TED AIs to confirm they’re in the basin of good deference, have different TED AIs adversarially audit one another for signs of misalignment, and so on. My guess is that if we spend our five years of pausing at TED AI well, we can have high confidence by the end of the pause that our AIs want to follow human instructions and are not scheming to disempower us.
Now, I take you to be saying that even in this situation, we would not have solved all the problems we need to solve to scale safely all the way to the limits of intelligence. For instance, we wouldn’t have solved the problem of ASI inventing an alien ontology or becoming so good at persuasion that the concept of instruction following breaks down. These seem like real problems to me. But it’s not clear to me that we could solve them with a longer pause at TED AI. They seem like fundamental problems with aligning an incomprehensibly greater intelligence to a human intelligence.
Our response to Séb Krier on Plan A
I find the data in the left half of figure 2 surprising. It’s saying that without CoT, models are good at very short horizon tasks and at somewhat long horizon tasks but bad at intermediate tasks. For instance, Opus 4.7 does well on <1 minute tasks and well on 1 hour tasks, but it does poorly on 5 minute tasks.
Do you have an idea why that might be the case? It’s not what you would expect under a survival model (like Toby Ord proposed to explain the METR TH study).
Most of the way good thinking happens IMO is by finding and using a good ontology for thinking about some situation, not by probabilistic calculation.
As a side point, this is a trendy view in epistemology. Most of our learning in real life is not a matter of reallocating credence among hypotheses we were already aware of according to Bayes’s theorem. Rather, most of our learning is becoming aware of new hypotheses that weren’t even in the domain of our prior credence function.
Beyond Uncertainty by Steele & Stefánsson is a good (and short) overview of approaches to awareness growth in formal epistemology.
If you do take on this project, I’d especially like to know whether the decline in parasite prevalence coincides with GPT-4o’s retirement. I suspect that much of the chatbot psychosis & parasitic AI wave of 2025 was due to GPT-4o being an egregiously bad model, and we should therefore expect a lot fewer cases now that it’s gone.
Interesting! Thanks for the link.
Emergency Response Measures for Catastrophic AI Risk
My understanding is that the US Marshals are not only accountable to the Court either. They take their day-to-day commands from the Director of the US Marshals Service, a presidential appointee, who in turn reports to the US Attorney General, another presidential appointee.
This makes me even more doubtful that the US Marshals would side with SCOTUS and, eg, arrest the President in a worst-case constitutional crisis. Both the Marshals’ boss and their boss’s boss would likely side with the President, having been chosen by him for their loyalty.
It’s sometimes suggested that as a defense against superpersuasion, one could get a trusted AI to paraphrase all of one’s incoming messages. That way an attacker can’t subliminally steer the target by choosing exactly the right words and phrases to push their buttons.
Maybe one could also defend against superhuman persuasion and manipulation by getting a trusted AI to paraphrase all of one’s outgoing messages. That way an attacker can’t pick up on subtle clues about the target’s psychology embedded in their writing style. If this paraphrasing strategy works, the attacker is denied the training data they need to figure out how to push the target’s buttons.
The world’s first frontier AI regulation is surprisingly thoughtful: the EU’s Code of Practice
Introducing SB53.info
If your two assumptions hold, takeover by misaligned AGIs that go on to create space-faring civilizations looks much worse than existential disasters that simply wipe humanity out (eg, nuclear extinction). In the latter case, an alien civilization will soon claim the resources humanity would have taken over had we become a space-faring civ, and they’ll create just as much value with these resources as we would have created. Assuming that all civs eventually come to an end at some fixed time (see this article by Toby Ord), some amount of potential value will be lost, but not as much as would be lost if a misaligned AI used all the resources in our future light cone for something valueless while excluding aliens from using them. So the main strategic implication is that we should try to build fail-safes into AGI to prevent it from becoming grabby in the event of alignment failure.
It sounds like your high level plan is to make artificial intelligence smart enough to stabilize the situation and no smarter. Then do relatively chill human intelligence augmentation & cyborgism until eventually the augmented humans/cyborgs are ready to build superintelligence.
I don’t think this is a crazy plan. It might be better than Plan A. I guess my main worries would be
Exogneous xrisks that TED AIs can’t manage down to zero get us while we’re doing chill intelligence augmentation.
The TED AIs somehow drift into misalignment.
The limits of machine intelligence prove to be so far from the limits of augmented human intelligence that we’re never ready to bulid superintelligence, and maybe we miss out on some unique benefits that superintelligence would have provided.