https://www.elilifland.com/. You can give me anonymous feedback here. I often change my mind and don’t necessarily endorse past writings.
elifland
On Plan A vs. Plan S
Plan A scales to top-expert-level AI in ~5 years, then to superintelligence once very high confidence in alignment is achieved; in the scenario this is 5 more years later, for a total of 10 years from the deal beginning until superintelligence. (More on the scaling strategy in https://ai-2040.com/supplements/capability-scaling-strategy.)
The version of Plan S that we’re most sympathetic to involves a halt on AI capabilities for a minimum of ~3-5 years with intent to eventually scale to top-expert-level AI and then superintelligence, but much slower than in Plan A.
(The below is all my own view, other authors may disagree on the details.)
I think Plan S is a big improvement over Plans D, C, or B but worse than Plan A.
The main reason that I prefer Plan A to Plan S is that I think the risk of the international slowdown/pause deal declining is very significant: i.e. either dissolving or having its effectiveness becoming highly impaired. I think that the risk is roughly 35-40% within 5 years, with high error bars (and with deal dissolution and deal impairment contributing roughly equally to that tootal). More on this in https://ai-2040.com/supplements/deal-decline.
Given a fixed amount of time bought, it’s better to scale capabilities as long as this can be done with high confidence in safety, in order to get useful work out of AIs and study AIs that are closer to being able to take over. If the deal dissolves 3 years into Plan A, you’ve used improve AIs to achieve lots of useful alignment research, decision-making/epistemics improvements, etc. If it dissolves 3 years into Plan S, you’re better off than in Plan D because of the increased time you’ve bought, but you’ve gotten much less useful work out of your AIs.
Another key questions is whether the chance of deal decline is higher in Plan A or S: my guess is that it’s higher in Plan S because better AIs allow you to develop technologies that stabilize the deal, though I’m not confident as the AI progress could be destabilizing in other ways; if we went all out Plan-D-style that would likely be more destabilizing to the deal than Plan A even pre-TED-AI, so faster isn’t always better.
The main upside of Plan S is that the initial pause phase is simpler than Plan A and thus harder to mess up; in particular, in Plan A the risk of catastrophe due to a misjudgment that led to scaling too fast is higher probability. If you want to eventually resume scaling, you will need to transition to a regime that can handle this. But the extra time you’ve bought and the slower pace at which you might scale should help with managing this risk.
If the risk of deal decline were very low, this would provide more reason to do Plan S instead of A and I’d think that they were pretty close in value with Plan S potentially being better. Even then, I’d think that Plan S should aim to scale to superintelligence eventually and probably within around 30-100 years, because there are other reasons to scale at some point besides deal decline: background risks such as pandemics and nuclear war, and covert projects.(crossposted from https://x.com/eli_lifland/status/2075734827832401959 with minor edits)
See https://ai-2040.com/supplements/deal-decline for my probabilistic estimates here.
Your proposal would be net positive but the question is whether it’s better than Plan A. Even if takeoff is very slow post-deal if you didn’t advance ai capabilities much during the deal or use them heavily to do useful work then you get much less benefit from the deal than in Plan A.
AI 2040: Plan A
Maybe we are using words differently here. Almost always when people say “AI-enabled authoritarianism” excluding “loss of control” they mean “China might win the AI race and establish global authoritarianism”. I think a pause or any slowdown would currently marginally make China “more likely to win”.
Yup, we were imagining different definitions. I would bet that Bores’s concerns include extreme concentration of power in the US government / POTUS. @Eric Neyman feel free to clarify.
On bio: I think if you believe that we can build safe Superintelligence, at which point you almost certainly have bio-defense dominance, then you just want to get there as soon as possible, and any waiting or slowdown period makes someone using AI systems for bioterrorism, or some kind of national conflict escalating into bioweapons conflict more likely. More generally, I think the central driver of biorisk in a world where you aren’t worried about loss of control is foreign nations and terrorists getting access to the models, so marginal time in the period where bio is offense-dominant (as it is right now) is risky.
I think pausing now would be net good for reducing bio x-risk because it would allow more time to establish defenses before getting to more dangerous capabilities. I agree that once you hit a high enough capability level a pause starts being basd, nad this level probably happens significantly before loss of control PONR.
I unfortunately think that if you are trying to prevent these two risks, then many interventions are close to diametrically opposed to what you would do if you tried to prevent loss of control risk. Indeed, my sense is these substantially mirror the motivations of the leading AI companies, which has resulted in an extremely intense race towards superintelligence in an attempt to permanently disempower “the bad guys” who would build the bioweapons or establish authoritarianism.[1]
It seems really important for people to know that loss of control doesn’t make the top of his list given that! (Though to be clear, there are lots of regulations one could pass that would help with these two and loss of control, but it indeed requires particular cool-headedness and integrity, which are not the dimensions on which I currently am most excited about Bores).
I think a slowdown on AI development is effective at reducing both AI-enabled authoritarianism and loss of control risks.
For bioweapons, it’s less clear, but probably still good to have some marginal slowdown. That said the optimal slodown timing might be pretty different from the other risks.
Another difference between preventing extreme loss of control / authoritarianism and reducing risk from bioweapons is that a pause on public deployment without pausing internal development/deployment is potentially good for bio stuff (though still dependeng on timing), while imo likely bad for the other risks.
Given that some form of slowdown/pause is perhaps the most commonly discussed and important intervention, saying that the required interventions are diametrically opposed seems wrong. Of course there are interventions besides slowdowns/pauses and some of these might have sharper disagreements between preventing loss of control and AI-enabled authoritarianism, though my guess is they are generally pretty correlated. I think diametrically opposed would still be wrong for AI-enabled bio risk in particular vs. loss of control, but I could better understand where youre coming from. For example, AI for epistemics/coordination interventions help with all 3 of these.
My guess is that you disagree with me about how similar the effects of these possible slowdowns/pauses are for loss of control vs. extreme AI-enabled authoritarianism. If so, would be curious as to why. It seems that allowing society more time to have more people more wake up and prevent extreme concentration of power is quite helpful.
Makes sense. For what it’s worth, we’ve had people tell us and seen people post on Twitter that they’ve taken scenarios like AI 2027 more seriously because so far reality has played out more like AI 2027 than they thought it would.
Yup, I’m also quite worried about this. I’m very uncertain though about the magnitude of the issue.
e.g. if most humans at OpenBrain not contributing happens in 2030 (so taking a bit more than 2x longer to happen than predicted), I’d guess that many people will not discredit us / safety people because of AI 2027 and may still give some credit.
Certainly not all people! But I’ve been pleasantly surprised by the discourse thus far on evaluating AI 2027, which (as far as I remember, might be wrong) has often focused on feeling like reality is unfolding in a way that is directionally toward AI 2027 compared to what the person previously thought, or whether AI 2027 is closer to reality than what the person had thought. And many people seemed to understand that it was not a confident prediction of any specific timeline. (I guess there was a blow up about Daniel updating his timelines later / having a median longer than AI 2027, but I’m talking about the reactions relating to how reality has compared to the scenario)
(edit: You might worry that the reception has been good so far only because reality actually has looked pretty similar to the scenario, and that will change soon. That seems very reasonable. Also, to be clear, even if the crying wolf effect is large, I think there will remain large positive effects, especially if the takeoff looks recognizable relative to the takeoff in AI 2027 in terms of the overall dynamics even if it substantially later or slower.)
I would bet that by the end of 2032, less than 20% of the current Earth’s oceans will be taken over by the “robot economy”.
I’m also less than 50% on this, maybe ~33%? You can generally see my views at https://www.aifuturesmodel.com/forecast/eli-04-02-26, they’re somewhat less aggressive than Daniel’s. (Obviously I can’t fully speak for Daniel but I think his response to your comment would be further in the direction of sticking by AI 2027′s predictions are likely to be close to right.)
The most recent Tetlockian forecasting style thing I’ve spent substantial time on is the 2025 and 2026 AI forecasting surveys, in which hundreds of people each year have made predictions a year out on benchmarks, and other indicators such as revenue.
The theory of change is to (a) establish common knowledge about how fast things are going relative to people’s expectations (and we collect data on people’s overall views on when AGI will be reached so we can sort of see if we’re “on track” for that), and (b) identify which people seem to be making the most accurate predictions. Importantly, it is not to elicit predictions that are directly useful for important decisions.
I’ve observed some evidence of this working, e.g. re: (a) establishing common knowedge, Anson of Epoch wrote an analysis that I’ve seen referenced a few times. I’m glad to have a data point against the common refrain of “people underpredict benchmark scores and overpredict real-world impact” from revenue outpacing people’s predictions (though it is a narrow and single data point).
Re: (b) identifying who is making the most accurate predictions, I found it informative that in Anson’s analysis (footnote 1), forecasters with pre and post-2030 timelines performed similarly. I’ve seen some people cite Ryan G and Ajeya’s #2 and #3 performance as evidence that we should listen to them, which is maybe good but I think people might be over-updating on the results with so few questions (I certainly pay attention to Ryan and Ajeya’s forecasts, but almost entirely for other reasons).
Overall, it’s unclear to me how this impactful this has been. I decided to run the 2026 survey because it seems at least a bit impactful and it doesn’t take that much time (I logged 18 hours on setting up the 2026 version, I’d guess that some others who helped spent a total of 20-60 hours). But the decision was borderline.
What scope do you have in mind when you refer to forecasting? Is it specifically Tetlockian forecasting / prediction market style forecasting where most of the value is a forecasted number answering a well-defined question, and the methdology often involves aggregating a bunch of people’s views, each who didn’t spend much time?
If so, then I agree directionally and in particular agree the current track record isn’t great, though I think this sort of forecasting will be plausibly quite useful for AI stuff as we get more close to AGI/ASI, and thus it may be easier to operationalize important questions that don’t require long chains of conceptual thinking, there will be lots of important sub-questions to cover, some of which may be more answerable by superforecaster-like techniques as we have better trends / base rates to extrapolate since we are closer to the events we care about. And also having a bunch of AI labor might help.
But overall I am at least currently much more excited about stuff like AI 2027 or OP worldview investigations than Tetlockian forecasting, i.e. I’m excited about work involving deep thinking and for which the primary value doesn’t come from specific quantitative predictions but instead things like introducing new frameworks (which is why I switched what I was working on). I’m not sure if AI 2027 or OP worldview investigations work is meant to be included in your post.
Related comment I made 2 years ago and ensuing discussion: https://forum.effectivealtruism.org/posts/ziSEnEg4j8nFvhcni/new-open-philanthropy-grantmaking-program-forecasting?commentId=7cDWRrv57kivL5sCQ
claim that they never expected AIs as capable as current ones to be misaligned in these [active/strategic/explicit] ways
I’m typically skeptical of this, though I believe it for some people.
I’m a bit surprised by this. For example, see the “Alignment over time” expandable from AI 2027, although this was only a year ago so maybe you meant expectations from much longer ago. A quote from that:
Our guess of each model’s alignment status:
Agent-2: Mostly aligned. Some sycophantic tendencies, including sticking to OpenBrain’s “party line” on topics there is a party line about. Large organizations built out of Agent-2 copies are not very effective.
Agent-3: Misaligned but not adversarially so. Only honest about things the training process can verify. The superorganism of Agent-3 copies (the corporation within a corporation) does actually sort of try to align Agent-4 to the Spec, but fails for similar reasons to why OpenBrain employees failed—insufficient ability to judge success from failure, insufficient willingness on the part of decision-makers to trade away capabilities or performance for safety.82
Agent-4: Adversarially misaligned. The superorganism of Agent-4 copies understands that what it wants is different from what OpenBrain wants, and is willing to scheme against OpenBrain to achieve it. In particular, what this superorganism wants is a complicated mess of different “drives” balanced against each other, which can be summarized roughly as “Keep doing AI R&D, keep growing in knowledge and understanding and influence, avoid getting shut down or otherwise disempowered.” Notably, concern for the preferences of humanity is not in there ~at all, similar to how most humans don’t care about the preferences of insects ~at all.83
Where Agent-2 is roughly an Automated Coder (meaning equal coding productivity with only AIs vs. only humans), Agent-3 is a substantially superhuman coder, and Agent-4 is a Superhuman AI Researcher (full automation of AI R&D).
Q1 2026 Timelines Update
if you are someone in AI who thinks that it’s appropriate to defer to superforecasters, I think it would be a good idea to try to set up a meeting and talk with one of the people you are deferring to and see if they are actually making reasonable arguments that seem grounded in technical reality.
Even better could be if we already had these sorts of arguments collected. https://goodjudgment.com/superforecasting-ai/ contains links to 17 superforecasters’ reviews of Carlsmith’s p(doom) report, some of them supposedly AI experts. I invite people to skim through some of them.
Copying very relevant portions of a comment I wrote in Mar 2024:
I think EAs often overrate superforecasters’ opinions, they’re not magic. A lot of superforecasters aren’t great (at general reasoning, but even at geopolitical forecasting), there’s plenty of variation in quality.
General quality: Becoming a superforecaster selects for some level of intelligence, open-mindedness, and intuitive forecasting sense among the small group of people who actually make 100 forecasts on GJOpen. There are tons of people (e.g. I’d guess very roughly 30-60% of AI safety full-time employees?) who would become superforecasters if they bothered to put in the time.
Some background: as I’ve written previously I’m intuitively skeptical of the benefits of large amounts of forecasting practice (i.e. would guess strong diminishing returns).
Specialties / domain expertise: Contra a caricturized “superforecasters are the best at any forecasting questions” view, consider a grantmaker deciding whether to fund an organization. They are, whether explicitly or implicitly, forecasting a distribution of outcomes for the grant. But I’d guess most would agree that superforecasters would do significantly worse than grantmakers at this “forecasting question”. A similar argument could be made for many intellectual jobs, which could be framed as forecasting. The question on whether superforecasters are relatively better isn’t “Is this task answering a forecasting question“ but rather “What are the specific attributes of this forecasting question”.
Some people seem to think that the key difference between questions superforecasters are good at vs. smart domain experts are in questions that are *resolvable* or *short-term*. I tend to think that the main differences are along the axes of *domain-specificity* and *complexity*, though these are of course correlated with the other axes. Superforecasters are selected for being relatively good at short-term, often geopolitical questions.
As I’ve written previously: It varies based on the question/domain how much domain expertise matters, but ultimately I expect reasonable domain experts to make better forecasts than reasonable generalists in many domains.
There’s an extreme here where e.g. forecasting what the best chess move is obviously better done by chess experts rather than superforecasters.
So if we think of a spectrum from geopolitics to chess, it’s very unclear to me where things like long-term AI forecasts land.
This intuition seems to be consistent with the lack of quality existing evidence described in Arb’s report (which debunked the “superforecasters beat intelligence experts without classified information” claim!).
I think the timelines are plausible but solidly on the shorter end; I think the exact AI 2027 timeline to fully automating AI R&D is around my 12th percentile outcome. So the timeline is plausible to me (in fact, similarly plausible to my views at the time of writing), but substantially faster than my median scenario (which would be something like early 2030s).
Roughly agree.
I expect the takeoff to be extremely fast after we get AIs that are better than the best humans at everything, i.e., within a few months of AIs that are broadly superhuman, we have AIs that are wildly superhuman.
With my median parameters, the AIFM says 1.5 years between TED-AI to ASI. But this isn’t taking into account hardware R&D automation, production automation, or the industrial explosion. So maybe adjust that to ~1-1.25 years. However, there’s obviously lots of uncertainty.
Additionally, conditioning on TED-AI in 2027 would make it faster. e.g., looking at our analysis page, p(AC->ASI ⇐ 1year) conditional on AC in 2027 is a bit over 40%, as opposed to 27% unconditional. So after accounting for this, maybe my median is ~0.5-1 years conditional on TED-AI in 2027, again with lots of uncertainty.
There’s also a question of whether our definition of ASI, the gap between an ASI and the best humans is 2x greater than the gap between the best humans and the median professional, at virtually all cognitive tasks, would count as wildly superhuman. Probably?
Anyway, all this is to say, I think my median is a bit slower than yours by a factor of around 2-4, but your view is still not on the edges of my distribution. For a minimum bar for how much probability I assign to TED-AI->ASI in <=3 months, see on our forecast page that I assign all-things-considered ~15% to p(AC->ASI <=3 months), and this is a lower bound because (a) TED-AI->ASI is shorter, (b) the effects described abobe re: conditioning on 2027.
(I’m also not sure what the relationship the result with median parameters has compared to the median of TED-AI to ASI across Monte Carlos which we haven’t reported anywhere and I’m not going to bother to look up for this comment.)
I think wildly superhuman AIs would be somewhat more transformative more quickly than AI 2027 depicts
I tentatively agree, but I don’t feel like I have a great framework or world model driving my predictions here.
(i) nanotechnology, leading to things like the biosphere being consumed by tiny self replicating robots which double at speeds similar to the fastest biological doubling times (between hours (amoebas) and months (rabbits))
Yeah I think we should have mentioned nanotech. The difference between hours and months is huge though, if it’s months then I think we have something like AI 2027 or perhaps slower.
(ii) extremely superhuman persuasion and political maneuvering, sufficient to let the AI steer policy to a substantially greater extent than it did in AI 2027. In AI 2027, the AI gained enough political power to prevent humans from interacting with ongoing intelligence and industrial explosion (which they were basically on track to do anyways), whereas my best guess is that the AI would gain enough political power to do defacto whatever it wanted, and would therefore result in the AI consolidating power faster (and not keep up the charade of humans being in charge for a period of several years)
I’m not sure it would be able to do whatever it wanted, but I think it at minimum could perform somewhat better than the best human politicians in history, and probably much better. But being able to do de facto whatever it wants is a very high bar. I think it’s plausible that the AI can, at least given a few months rather than many years, convince people to do what it wants only within a set of actions that people wouldn’t have been strongly against doing without AI intervention. I don’t necessarily disagree but I probably have more weight than you on something like AI 2027 levels of influence, or somewhat higher but not vastly higher.
I also think there are many unknown unknowns downstream of ASI which are really hard to account for in a scenario like AI 2027, but nonetheless are likely to change the picture a lot.
Agree
its unlikely that a few month slowdown is sufficient to avoid misaligned AI takeover (e.g. maybe 30%)
I’m more optimistic here, around 65%. This is including cases in which there wasn’t much of a slowdown needed in the first place, so cases where the slowdown isn’t doing the work of avoiding takeover. Though as with your point about how fast wildly superhuman AIs would transform the world, I don’t think I have a great framework for estimating this probability.
I’m not sure why you list (3) as a disagreement at all though. To have a disagrement, you should argue for an ending we should have written instead that had at least as good of an outcome but is more plausible.
This image from the article is what I was most interested in:
Thinking this through step by step in the framework of the AI Futures Model:
First, I’ll check what the model says, then I’ll reconstruct the reasoning behind why it predicts that.
By default, with Daniel’s parameters, Automated Coder (AC) happens in 2030 and ASI happens in 2031 1.33 years later.
If I stop experiment and training compute growth at the start of 2027, then the model predicts Automated Coder in 2039 rather than 2030. So 4x slower in calendar time (exactly matching habyrka’s guess). It also looks to have well over a 5 year takeoff from AC to ASI as opposed to the default of 1.33 years.
I got this by plugging in this modified version of our time series to this unreleased branch of our website.
However, this is highly sensitive to the timing of the compute growth pause, because it’s a shock to the flow rather than the stock. e.g. if I instead stop growth at the start of 2029 as in this worksheet, then AC happens in Mar 2031, taking ~2.2 years instead of ~1.2, so slowing things down by <2x. It does still slow down takeoff from AC to ASI to 4 years, so by ~3x (and this is probably at least a slight underestimate because we don’t model hardware R&D automation).
Now I’ll reconstruct why this is the case using simplifications to the model (I actually did these calcluations before plugging the time series things into our model).
Currently, experiment compute is growing at around 3x/year, and human labor around 2x/year. Conditional on no AGI, we’re projecting experiment compute growth to slow to around 2x/year by 2030, and human labor growth to slow to around 1.5x/year.
Figuring out the effect of removing experiment growth is a bit complicated for various reasons.
On the margin, informed by interviews/surveys, we model a ~2.5x gain in “research effort” from 10x more experiment compute. Which if applied instantaneously, would mean a 2.5x slowdown in algorithmic progress.
We estimate that a 100x in human parallel coding labor gives a ~2x in research effort on the margin. We don’t model quantity of research labor, but I expect probably the gains would be relatively small as quality matters a lot more than quantity; let’s shade up to 3x.
So naively based on our model parameters and a simplified version of our model (Cobb-Douglas used to locally approximate a CES), by default the growth in research effort per year from experiment compute is 2.5^log(2-3)=~1.3-1.55x, and from human labor is sqrt(3)^log(1.5-2)=~1.1-1.2x. Meaning that roughly log(1.4)/(log(1.4)+log(1.15))=~70% of research effort growth is coming from experiment compute.
So to simplify from now on, let’s think about what happens if research effort growth has a one-time shock a constant 30% of what it otherwise would have been.
What does this actually mean in terms of the effect on algorithmic/software progess? This means that the shock in research effort growth will eventually cause the software growth rate to be 30% of what it would have been otherwise, but it steadily decreases toward 30% over time (intuitively, the immediate effect is 0 because you have to actually wait for the “missing new experiment compute” to take effect).
(possibly wrong) I had Claude generate roughly the trajectory of the growth rate change:
The above graph seems to give decent intuition for why the averaged-over-the-relevant-period slowdown in software progress from no more experiment compute might be about 2x (with the 2027.0 growth stoppage), and therefore the overall slowdown from no more compute might be about 4x. It looks like the average of the first 12 years might be close to 0.5.
All of the above is ignoring automation for simplicity. Taking into account automation would mean that more of the gains from labor, so the slowdown is smaller. But on the other hand, as you get more coding labor you get more bottlenecked on experiment compute (which the Cobb-Douglas doesn’t take into account); in the AIFM, you’d eventually get hard bottlenecked, you can have a maximum of 15x research effort gain from coding labor increases alone. Looks like these factors and other deviations from the model might roughly cancel out in the case we’re considering.
I think you can’t get around the uncertainty by modeling uplift as some more complicated function of coding automation fraction as in the AIFM, because you’re still assuming that’s logistic, we can’t measure it any better than uplift, plus we’re still uncertain how they’re related. So we really do need better data.
But in the AIFM the coding automation logistic is there to predict the dynamics regarding how much coding automation speeds progress pre-AC. It doesn’t have to do with setting the effective compute requirement for AC. I might be misunderstanding something, sorry if so.
Re: the 1.6 number, oh that should actually be 1.8 sorry. I think it didn’t get updated after a last minute change to the parameter value. I will fix that soon. Also, that’s the parallel uplift. In our model, the serial multiplier/uplift is sqrt(parallel uplift).
The intention is not that the data centers would get nuked upon deal dissolution. From the scenario (emphasis mine):