Existential Risk Is An Extraordinary Claim
This is a cross-post from my blog post.
Over the past decade, the effective altruism movement has become increasingly focused on “existential risk,” the idea that, this century, there’s a significant chance that humanity will go extinct, become permanently disempowered, or otherwise lose almost all of its future potential.
In this post, I’m not going to argue that, given what we know, we should believe that existential risk is high or low. Instead, what I’m going to argue is that existential risk is an extraordinary claim and that, as such, we must have extraordinary evidence to be concerned about it. Whether we have such evidence I’ll leave to the reader.
My most general argument for this is simply that existential risk represents a dramatic state change in the world and that, the more dramatic of a state change someone is arguing for, the more evidence they must supply to convince us of it. If your friend tells you that one of your windows could get broken this year because of the upcoming hail storm, you have good reason to believe them. But, if they tell you that every single house in your entire state will be completely destroyed, they need to provide extreme evidence to substantiate such a claim.
Extinction
The dominant existential risk that effective altruists worry about is human extinction.
But, notably, in order for this to happen, there has to be an extremely specific set of conditions. It can’t merely be that some event, say a nuclear war or a pandemic, causes a large number of people to die. Instead, it has to be that, for some reason, every single last human is killed. The only mechanisms I can really think of that would cause this is either that some entity actively seeks out and kills every human or that every part of the Earth becomes uninhabitable for a significant duration of time.
One way some argue this could occur is that humanity will develop a superintelligent AI that will systematically seek out and kill every last human. But, to believe that this could happen, we have to believe the claims that humanity can make such an intelligence, that humanity will make such an intelligence, and that, if we do, it will seek to destroy us.
Another way some argue that this could occur is that humanity will develop engineered viruses that infect and then kill everyone. But, to believe that this could happen, we need to have good evidence to think that such a virus (one that spreads rapidly and has 100% fatality) can be engineered, that people who are unreachable (because they’re in bunkers or outer space or Antartica) will also die, and that humanity will be unable to create a cure for such a virus in time.
The last way that some argue this could occur is that two great powers will engage in a nuclear war that causes a nuclear winter. This could cause crop failures around the world, and billions would certainly die. But, it’s important to remember that nuclear winters are not expected to last more than a decade. Given this, we need to believe not only that total war between nuclear powers will occur and that this war will result in a nuclear winter, but also that those living underground or in protected areas will also perish before the surface of the Earth becomes capable of agriculture again.
Some may take the claims of a few experts or their own “vibes” as evidence that we should be very worried about existential risk, but it’s important to remember that humans have a natural proclivity towards believing that their generation could be the last. (According to Pew, “39% of adults believe that humanity is ‘living in the end times.’”) Given this, we can’t merely rely on the beliefs of a few experts or our personal opinions. Instead, we need the concern of a significant portion of experts to really justify concern. Some may believe they have strong enough evidence to worry, but it’s important to remember the mountain of learning that has to be done to be at all confident about these things. (For instance, it seems pretty hard to say whether extremely deadly viruses could be engineered without knowing a lot about virology.)
Disempowerment
Another existential risk effective altruists worry about is disempowerment.
The basic concern is that, if superintelligence is developed, it actually may not eliminate humanity because that would be a waste of resources. Instead, it would simply disempower us by eliminating our ability to make choices about the future of our species. Whether this is more or less likely than the extinction scenario I’ll leave to the reader.
Irreversible Civilizational Collapse
The last existential risk effective altruists worry about is irreversible civilizational collapse.
Some effective altruists believe that, even if the risks in the extinction section wouldn’t actually kill everyone, they would still kill enough people that humanity would be unable to re-industrialize afterwards, possibly because it has gone through all of the easily accessible fossil fuels on Earth. (The thought is that, if you can’t get coal, you might not be able to make a steam engine and then not be able to make almost anything else.)
I think this is also an extraordinary claim because it supposes that there’s a significant risk of almost everyone dying, that, if almost everyone died, civilization would become set back to pre-industrial development, and that, if this were to occur, humans would never be able to re-industrialize despite having vast amounts of knowledge, technology, and creativity.
Okay, by that definition it is an extraordinary claim and demands extraordinary evidence.
But that evidence exists in abundance. We are in fact creating AI rapidly; there are piles of evidence to that effect.
In short: yes it’s an extroardinary claim, and unfortunately we’ve got extraordinary evidence supporting that claim, in such abundance that it now demands some very weird targeted demands for rigor/evidence to conclude that we’re not at fairly high (~10%) to very high (~50%) risk of extinction.
You might demand extraordinary evidence that it will get much farther, but I doubt you’d go there; there’s tons of evidence that it is steadily and rapidly increasing in intelligence, so it seems more sensible to ask for evidence that this will stop some time soon or some place safe.
You might also demand evidence that we will make future AI autonomous, more like a peer species than the human-directed tools we see in publicly available systems. Even before the recent Hugging Face incident, I don’t think you should have demanded extraordinary evidence for that one; there’s abundant evidence that people want work done for them, and no evidence that making those systems autonomous has a separate barrier from making them smarter.
Finally, you might say that we’re not at risk even though the evidence is that we’re going to pretty soon have created a set of peer species smarter than we are in all the ways that matter. To that, I think the most obvious reference class is the history of intelligent species competing for resources with a smarter or even peer species: the extinction of all other humanoid species when sapiens arrived. That’s not quite fair, because we’re building this successor; that makes it a truly unprecedented event to which prior evidence doesn’t directly apply, and you’ve got to reason by particulars.
That is what we call the alignment problem, and we don’t even have broad agreement on how difficult it is. Summing opinions weighted by rough time-on-task would give a very wide distribution of doom estimates, but I think realistically with a center of gravity far too high for comfort, somewhere in the 30%-60% range.
So yes it is an extraordinary claim, but that has already been addressed and wrapped into the thinking of the alignment/x-risk community.
Thanks for the thoughtful response!
Could I ask why you think that AI’s intelligence is scaling? It seems to me like there are a range of capabilities on which it is rapidly advancing but others on which it is barely or not at all improving. This, at least to me, doesn’t really suggest we’re on the edge of ASI. I know people commonly argue that AI will do AI research creating a feedback loop, but I can’t see how this feedback loop would start if the AI doesn’t have a good research sense.
I personally feel that their research sense/taste is often underestimated/underelicited, at least if we keep the multiagent scenarios in mind (where agents can discuss the matters of taste among themselves).
In particular, some experiments tried to measure how good the models were at imitating the particular choices of specific human researchers, but without evidence that that particular set of human researchers themselves had good tastes worthy of imitation.
Basically, what matters is not if the agents are very good at imitating the “mediocre middle” (which is what I believe people have been measuring), but whether they have or will have enough taste for a good rate of real breakthroughs. It’s difficult to measure that in a controlled experiment.
In this sense, in my “single data point”, GPT-6 Astra found for me a solution of a not very difficult open math problem I was not able to solve for decades and previous models were not able to solve, and that solution was stunningly beautiful and unexpected in its approach. It’s not anything conclusive, but it does update me towards believing that their taste is underestimated and underelicited at the moment, that the ability to reliably imitate choices human researchers make in an average ML paper does not tell us much (and that teaching that ability too strongly might even be detrimental to those capabilities which matter).
What was the problem?
There was a conjecture that every complete lattice is isomorphic to the set of fixed points of a Scott continuous transformation of a powerset.
Astra found a very neat counterexample that solved most of it [1] . Happy to post a link if people would like to look at the details.
The case where the Scott topology on the complete lattice in question is required to have countable basis is still open. I am 99.9% sure that the solution is correct, with Lean 4 verifications and other cross-checks (like discussing the solution with other models and writing the solution in my own words and contributing an improvement to it myself). But I am still putting some further efforts into this.
It doesn’t seem like the timing is really an element of your argument, though?
But to address it anyway, because I think it’s crucial for how fast we have to move on X-risk: I agree that models aren’t scaling rapidly to acquire research taste, but I don’t think they need it to help substantially with AI progress. I do think they are scaling rapidly towards being able to absorb and make sense of the claims in complex research literature, in a way that will help human researchers apply their research taste much more efficiently. And merely the help with efficiently coding experiments and processing results will let human researchers spend much more of their time applying their research taste locally.
As for why I think we are on the verge of human expert-level abilities in most of the relevant areas, and then will rapidly move beyond:
It’s complicated. Much of my work addresses this. You could look at either Capabilities and alignment of LLM cognitive architectures which is old but has held up or the newer Human-like metacognitive skills will reduce LLM slop and aid alignment and capabilities for specific reasons I think we should be very concerned with rapid scaling to at least weak ASI.
But I’d call the claim that we are making progress now, and that progress is just going to stop at some point, much more extraordinary than the claim that progress will continue at some pace until human abilities are exceeded.
I know people who have pretty good arguments for expecting that LLMs won’t scale to ASI until another method surpasses them, and I think this is plausible. But I don’t really know of any credible claims that LLMs are making zero progress in any area of cognition. If you have, I’d be interested to hear.
It’s enough for LLMs to scale enough to invent that other method, it’s not relevant if they scale to ASI themselves. For me, the crux is continual learning in a strong sense (enabling accumulation of serially deep collections of new ideas). With this year’s qualitative results, it’s quite plausible continual learning is a sufficiently shallow problem that in 2026-2028 LLMs (as they scale) figure it out, and then LLMs-with-continual-learning directly scale to ASI (within a few months, reaching a point where they can invent enough stuff for physical compute to start growing on its own, with doubling times counted in days).
I wish I didn’t, but I very much agree.
I think the question on continual learning isn’t whether but how good it gets how fast. Unfortunately, LLM AGI will have memory, and memory changes alignment By memory I meant what we now call continual leaning and by will have I mean probably soon. I’m now a little pleased by every month I don’t see a breakthrough claimed.
Ah, thanks for sharing your research!
I only asked about the timing thing out of curiosity. I think that, if there’s no clear path to AGI, that should stretch out our timelines significantly as it seems hard to predict when research breakthroughs will occur (although, obviously, you’d know better than I.
The portion of experts that are concerned is definitely significant! This statement saying
was signed by two of the three “godfathers of AI” (Geoffrey Hinton and Yoshua Bengio) the CEOs of all three top US AI companies, (Sam Altman, Dario Amodei and Demis Hassabis), as well as lots of other experts. In this 2024 survey, 744 researchers were asked to put a probability on “AI causes human extinction or permanent disempowerment”. Half answered 10% or above, and the mean answer was 18%.
When the observer’s very existence is at stake, the maxim “extraordinary claims require extraordinary evidence” becomes paradoxical, because the most obvious extraordinary evidence one could imagine would involve being killed. The anthropic principle thus acts as a censor, living observers will never have access to that level of evidence. A lawyer would call it a probatio diabolica.
You may see this as a double standard, but there is something rational in relaxing the demand for extraordinary evidence in this case, and accepting ordinary evidence (alignment research, OAI/HF incident...).
Alternatively, you could revisit your prior. Is this really an extraordinary claim ? Homo neanderthalensis fared rather poorly after encountering Homo sapiens. I won’t elaborate here : there is an entire literature on the subject.
Also extinction need not be sudden, either. It could just as well be gradual, through disempowerment and slow fading.
Hi Raphael,
I think you think that I view extraordinary evidence as being unattainable. I don’t. For instance, if 10% of all virologists signed a document stating that we should permanently ban all gain-of-function research due to risks of human extinction, I would think we should ban such research.
I didn’t mention gradual extinction risks since people don’t seem to focus on them. But, if you’re referring to the essay “Gradual Disempowerment”, I generally find it highly unlikely since it supposes that society wouldn’t push back as these forces are occurring. It could make sense to worry about in the future, but I think there would be much more evidence to justify such worry then compared to now.
How would those signatories have come by their “extraordinary evidence”? How would the very first virologist to notice such a risk have come by it?
How did Eliezer come by his supposedly “extraordinary evidence” for x-risk?
The word “extraordinary” is doing the wrong sort of work here. What is an “extraordinary” claim? The Bayesian meaning would be “a day of very low probability”, requiring by the standard Bayesian calculation a corresponding number of bits of evidence to elevate it. But you instead glossed “extraordinary” as “dramatic”, which is about feels. You want evidence that “feels dramatic” to counteract it.
The “preference cascade” is already in motion. At what point would you join with it, and why?
By the time he thinks “there is a risk” there should be a next virus researcher who thinks we’re 95% of the way to what may be a serious risk, and another researcher who thinks we’re 90% of the way, etc. The extra leap past 95% would then not count as extraordinary.
And the standards for what counts as extraordinary between a virus researcher and a layman are different anyway.