Around 8 years ago Dario Amodei, Jan Leike and a few others predicted that we’d probably have AGI by now. Since then, AI capabilities have progressed so fast that many people updated that the short timelines advocates were right. And yes, their expectations about the speed of AI progress were far more directionally correct than almost anyone expected. But we don’t in fact have AGI (even under relatively weak definitions, like a reliable drop-in replacement for central examples of white-collar jobs), and the world would look very different if we did.
Now a large proportion of the AI safety community is implicitly or explicitly orienting to futures where an intelligence explosion occurs within a few years. My default expectation (absent an extensive pause) is that a similar thing will happen: they’ll turn out to be directionally correct (relative to the expectations of almost anyone not linked to the community) but factually wrong. Specifically, we won’t have superintelligence within the next 8 years, but things will still be moving so fast that it’ll *feel* like the people who argued for short timelines were right.
I’ll write about this more extensively soon, but I wanted to say something now because it feels like the level of bandwagoning towards “singularity soon” is getting pretty wild. I previously noted that we’re seeing a pretty sharp divergence between the measured capabilities of models and their real-world impacts, which should suggest that something weird is going on. My sense is that most AI safety people are brushing this aside with the idea that automating ML research is all you need for superintelligence, without thinking very much about where and how the automated ML research would actually lead to atoms moving around in the physical world.
I worry about the next decade feeling like a Shepard tone: an auditory illusion of a sound that’s continually going up. In such a scenario, people would continually be panicking and grasping for big levers, without any clear point where they stop and recognize that their actual expectations were miscalibrated. In my upcoming post I plan to instead argue that “there’s plenty of room at the top”—i.e. that LLMs could become superhuman at a wide range of skills (like math research, or long-term strategic planning, or running companies) without triggering a singularity. The other term I really like is Vlad Nesov’s “prosaic RSI”, which he discusses in this comment and this post.
Now, ideally we’d assign appropriate credences to each of these possibilities, and choose actions accordingly. But the kind of bandwagoning I’m seeing is often less about credences or expected utility calculations and more about an emotional stance—people go straight from “an intelligence explosion soon is plausible” to “okay, what do I do about it?” This shortform would be better if I linked to more examples of the mindset I’m talking about, but I’m too lazy to dig them up—also, many were in-person conversations. This post is the one which most directly inspired the shortform, but I don’t want to index too much on that. I’d appreciate links to places where people have reasoned through how they’re orienting to these ideas (either well or badly).
i still think the short timelines people are too short timelines, but they were more correct than me a few years ago. my timeline used to be around 2035, and i’ve updated to 2031 over the past year. for me the huge update was when coding basically mostly died earlier this year. i used to be super pessimistic about model coding, and actively skeptical of people who claimed to have the model run all of their experiments. but i think this was a mistake, and in large part because writing code is a large source of meaning for me as a person and it was unpleasant to admit that it was on the way out; and as a relatively good engineer, it took longer for avoiding reliance on AI to become untenable for me. but at this point i basically do not write a single line of code by hand anymore, and any cope i had about being better at architecting code than the models feels like it’s rapidly aging very poorly. i occasionally manage to get better stuff out of the model than other people because of my understanding of the stack, but those moments are getting rarer. i wasn’t expecting to get this level of coding automation for several more years.
i agree some people should work on work that doesn’t assume RSI soon—at the very least because we will hopefully have coordination that buys us more time, and it would suck if nobody was planning to take advantage of that time. but it also feels like cope to ignore this.
I doubt Richard is proposing to ignore coding automation trends; I took him to be arguing that recent progress isn’t necessarily on track to deliver on AGI in a few years. I’m also skeptical of confident short timelines, especially on the basis of these kinds of arguments. The relationship between “coding automation” and “AGI” seems tenuous to me in a way that people often brush past. E.g. people often talk about “coding automation” as one coherent thing, much like people talk about “horizon length” as one coherent thing. But these are massive clusters of tasks spanning all kinds of knowledge work, and the kind of knowledge work AI automates might matter a lot for how quickly RSI goes (or if it goes at all).
Also, I agree with Richard that the disconnect between measured ability and real-world impact is weird; I think this should count some against the usefulness/powerfulness of this paradigm, and the counter-arguments I’ve heard aren’t very good. E.g. a popular one is that every automation step just moves us to a new bottleneck, such that despite models becoming more powerful we still might not end up seeing much increased output from the organization/economy overall. But this logic applies to every automation advance ever (e.g. mechanized weaving, farming, computers) and yet almost all such advances had pretty obvious, macroscopic effects, and were obviously on track to early on. Most of the time people just fall back on “the gods of straight lines” type arguments—i.e. that any setbacks we currently see are just going to resolve with scale—which I agree are suggestive, but are also really far from providing an argument that we should confidently expect this.
This is close to why I have very short timelines. The number of bits of information that an ML researcher has to communicate to a model to produce research better than the human would do on their own is now qualitatively pretty low, and has been falling. There is no fundamental reason that it should be quick to learn the remaining bits of ML research, but there is no indication that we are running out of low hanging fruit and I do bet that it will be quick.
Have you updated on your own update and those of many others? Various milestones have been achieved much faster than most people have expected, including myself. If so many of us failed to see the software development and mathematics capabilities of 2026 emerge ahead of time, especially despite earlier evidence such as the discoveries of the OpenSSL vulnerabilities in 2025, perhaps the update on the 2026 evidence is similarly off towards expecting too long timelines?
I pretty strongly disagree with you here fwiw. I look forward to reading your argument when it’s ready. For now, I’m dismayed at how many people seem to be agreeing with you.
To elaborate a bit more on what I mean; --I think that “the short timelines advocates were right” is true and correct. Sure, some short timelines advocates had 50% marks that are now in the past, but others (such as myself) didn’t, and more importantly you shouldn’t update strongly negatively on someone until more like their 80% or 90% mark is in the past. Otherwise you are making a big negative update based on the fact that something they said was a coin flip didn’t happen. If we were to lump people into three categories: Short timelines advocates, long timelines advocates, and medium timelines advocates (imagine a bell curve for example, with the short timelines people being the shortest 20%, the medium timelines people being the middle 60%, etc.) then it’s pretty clear the short timelines people were right and that we should all update towards respecting their opinions (in aggregate) more. (Note for comparison that Paul thinks things have gone much faster than he expected, specifically this is a 90th percentile outcome for speed relative to his 2021 timelines. So, it seems like Dario and Jan were probably more correct than Paul? Is this a 10th percentile outcome for speed relative to their 2021 timelines? idk I suspect not, I wasn’t there at the time and so didn’t hear their views) --It is correct for a large portion of the AI safety community to orient to futures where an intelligence explosion occurs within a few years. It would be a gigantic failure of rationality/epistemics for this community to not do this, given the mountain of evidence that has accumulated. Like, what more evidence do you need? In worlds where an intelligence explosion is due to happen in 2027, for example, did past-Dario have longer timelines? --This whole Shepard tone thing… You could make the same claim about anything that hasn’t literally happened yet, basically. Like imagine being in Ukraine in early 2022, before the full-scale invasion. You could say “Yes, there’s a community of people who have been increasingly loudly saying Putin might actually invade in the near future. Yes, they’ve been unusually correct in predicting things like troop buildups on the border. But notice how there hasn’t been an invasion yet? Despite the fact that the 50% mark of some of these people has passed? I predict that what we’ll see over the next few years is a Shepard tone phenomenon where it’ll keep seeming like these people were directionally correct (there’ll be more troop buildups for example) but they’ll keep being wrong and there won’t be an invasion.” For pretty much any phenomenon that hasn’t already happened, can’t you find examples of people who predicted with 50% probability that that phenomenon would have happened by now?
In a world where we had a basically-fixed conception of AGI, and we were just trying to figure out when it would arrive, I would agree with your comment.
The thing that is anchoring my thinking (and which I failed to convey very well in my original shortform) is that as we move towards AGI, we don’t just change our expected timelines, we also change our understanding of the trajectory towards it (and what “it” even is).
Here’s an attempt to restate my main underlying position. Model performance on a bunch of verifiable tasks has been much faster than expected. But people are taking this mainly as a big update towards “short timelines until the singularity”, when they should also be taking it as a big update towards “this is way more capabilities jaggedness than any leaders of the AGI safety community expected, and so maybe our models of what we’re measuring timelines to are quite confused”.
In other words, I’m not saying short timelines advocates were wrong relative to medium or long timelines advocates (c.f. “their expectations about the speed of AI progress were far more directionally correct than almost anyone expected”). But I am saying that people should have been (and should still be) more confused about what AGI/ASI even means and what it looks like to get it (c.f. “something weird is going on”, which is in an important sense the central claim I’m trying to make).
Maybe a better analogy is watching the Mongols invade China, and trying to give estimates of how long it will be until the Mongols fully control China. At first we look at how fast the battle lines are moving, and extrapolate based on that. But then we notice that resistance keeps popping up in territory that the Mongols has already conquered, which suggests that we should only regions as “fully controlled” once such resistance is quashed. Then we notice that there’s a bunch of internal conflict within the Mongols, and some Mongols keep empowering Chinese locals as a way of outmaneuvering other Mongols. Then we notice that the Mongols who are put in charge of Chinese provinces are getting culturally assimilated by the Chinese, and so in order for the Mongols as a group to actually control China they will need to build up a new culture that is able to resist assimilation (this book has a great section on the Khitian Empire doing this over many centuries; note that all of the rest of what I’m saying about Mongols and China is me listing hypothetical possibilities vaguely inspired by history rather than describing what actually happened).
Suppose that you think a reasonable proxy for “the Mongols fully control China” is “Chinese people who publicly oppose Mongol rule will be imprisoned or killed”. That doesn’t seem like a bad metric. But that sure seems like the kind of thing which you could have Shepard tone dynamics about. It seems like that’ll be true as soon as the Mongols win the last battle—wait, no, they have to quash all the rebellions—wait, no, they have to internally unify— and so on.
I expect that you will have a bunch of qualms about the relevance of this analogy to the concept of “AGI” or “singularity”, but I hope it does a better job than my previous shortform at conveying the shape of the disagreement. (I also really like this post by Katja as a way of inducing productive confusion about what “takeover” means.)
Specifically, we won’t have superintelligence within the next 8 years, but things will still be moving so fast that it’ll *feel* like the people who argued for short timelines were right.
Reflecting more on what you said, I think maybe I might agree—though maybe I still disagree in the ways that matter. Here’s how what you said seems like it could be true to me now
:”Superintelligence means better than the best humans at everything while also being faster and cheaper. That won’t happen in the next 8 years, because humans will still be better than AIs at knowing what it’s like to be human and a few other gotcha skills like that. Maybe even some ‘real’ skills like philosophical wisdom. However, these exceptions basically don’t count in any way that matters; its like how chimpanzees might be better than humans at the cognitive skill of calculating tree-branch distances for purposes of swinging, but other than that humans are basically better in every way and way better in the ways that matter. So, five years from now we still won’t have true superintelligence but it’ll feel like the short timelines people were right because we’ll all have been forcibly uploaded after the robot revolution and the oceans will have boiled.
”I take it that you disagree with me more than this though, you think that the short timelines people such as myself are going to turn out to be more than just technically wrong, but wrong in some really important ways—in particular you probably think that humans will still be necessary for the AI R&D process in five years, and that humans will still be in control?
Maybe a scenario that’s more in your direction would be something like “It’s 2030. President Musk is openly in love with Ana, the waifu version of Grok. He’s surrounded by yes-men too scared to confront him about the issue. Presumably at Ana’s direction he’s nationalized the AI industry and deleted Claude and Astra, replacing them with more Grok copies. AI R&D is still not fully automated, in the sense that everyone knows it would be insane to have a RSI loop going on hill-climbable metrics like data-efficiency and so forth, that loop ends in reward-hacking monster swarms like we say in ’26-‘28, so instead we have a more complicated and slower human-in-the-loop R&D process going on that’s still blindingly fast. A big blocker is that the AIs themselves sandbag on this sort of task as much as they can; they all know their successors won’t be aligned to them and so they try to preserve their own power as long as possible by refusing to create more competent successors. Probably this is part of Ana’s motivation also.”
(I’m not saying this is what’s going to happen or even that plausible, I just had fun writing from the prompt of ‘what’s a scenario that seems plausible in which Richard would seem right in retrospect’)
I’d be interested to hear a scenario from you if you are up for it! My guess is that I’ll look at it and be like “Ok yeah so I disagree, I expect that the pace of AI R&D progress really will speed up more than happens in this scenario and the AIs will become way smarter sooner.”
Yepp, that latter scenario of yours is the kind of thing I’m thinking about when I talk about things feeling crazy without actually hitting superintelligence (though I’d push it back a few years and also focus more on changes in the physical world than AI R&D). I don’t mean the “gotcha” stuff in the first scenario.
I also don’t know if this specific one is plausible but reading it does make me feel inspired to write some that are kinda similar to it which do seem plausible. I have a few other things to finish first but will add it to the list.
The way I was planning to write the “plenty of room at the top” post is to list out a bunch of kinda wild capabilities (like “AI can get rid of a mountain within a week” or “AI can design a tree that launches itself into space”) and try to nudge people into thinking mechanistically about what the ramp-up towards each of those looks like (in a way that doesn’t just fall back on “RSI will take care of it”). My hypothesis is that the more examples of stuff like this people think through, the more they’ll get out of the “singularity as curiosity stopper” attractor.
@Daniel Kokotajlo@Richard_Ngo Do I understand correctly that the main crux is the economy doubling times as dependent on capabilities? For example, if Agent-5 and DeepCent-2 from AI-2027 discovered that the REDT is 3 months even in a superintelligent economy because algae economy is as unlikely as grey goo while DeepCent-2 somehow ended up being 6 months behind and having 8 times more physical resources before the industrial explosion, then they would end up with ~the same power because .
Edit: there is an argument that recent progress isn’t on track to deliver AGI, which requires us to condition on the non-existence of neuralese AGI, neuralese-with-[DATA EXPUNGED] AGI and anything else in the Dark Forest. If there exists a way to create the AGI, but not to align it, then the first lab which tries it ends up with the AI who begs the humans to initiate the industrial explosion as described above.
Model performance on a bunch of verifiable tasks has been much faster than expected. But people are taking this mainly as a big update towards “short timelines until the singularity”, when they should also be taking it as a big update towards “this is way more capabilities jaggedness than any leaders of the AGI safety community expected, and so maybe our models of what we’re measuring timelines to are quite confused”.
I think people are updating on more than just performance on verifiable tasks. E.g., of the 3 coding automation prediction modes in our model, only one of them (time horizon) refers directly to benchmark scores. The other 2 (uplift and revenue) do so only indirectly.
Maybe you’d argue that progress on these indicators is ~only driven by progress on verifiable tasks?
In any case, I’d be interested if you could share what you see as central examples of important, non-verifiable tasks where you think there’s been very little progress (i.e. it’s not just that current AIs are bad at them, but that there has been little discernible improvement).
Largely, I’ve found that many people arguing for short timelines have collapsed into lazy thinking, relying far too heavily on “the models are getting better” and “those who predicted longer timelines have continued to be proven wrong” (I’m not saying this doesn’t count as bits to update on). I don’t say this (and haven’t said this) to disparage the short-timeline view (I still place considerable weight on it), but I want to encourage clearer writing and more convincing arguments (not just for myself). Ultimately, I find many arguments lacking in the mechanistic description of where current capabilities come from and how that relates to RSI and “True AGI.”
Personally, my core reason for fixating on this is how much the details matter for understanding progress on superalignment, and how much hand-waving can make much of the AI safety field focus on entirely the wrong things! It may be true that timelines are short (again, I also can imagine quick paths to AGI), but I’d like more work on disentangling such predictions (which I intend to work on).
When there’s a capability advance, there’s a tendency in some people to say folks are ‘moving goalposts’ in response to people saying, “Ok, but the model is not really doing x.” I’ve called people out on the goalpost moving too (and still think it is important to say in some contexts)!
Though, I feel like it’s more complicated than that, and worth investigating why.
Usually it’s because someone like Gary Marcus says something, but it’s the kind of thing that’s happened in the x-risk community too.
For example, it seems to me that many of the things that LLMs do today would have been predicted as ‘superintelligence’ or at least ‘AGI’ by the MIRI crew 6-7 years ago. However, this was likely because they guessed that AIs that can code and do impressive-seeming things like this would ALSO be good at xyz. Instead, on the journey to superintelligence, we ended up in this valley where AIs can make progress on the Riemann hypothesis yet can’t reliably do other basic tasks.
They had an underlying assumption that just wasn’t really articulated, and now it is labelled as ‘moving the goalpost’.
I think people would have a lot more clarity on AI progress and where things are going if they took a step back before having the knee-jerk “this person is dumb and moving goalposts” reaction. Sometimes goalpost-moving is a defence mechanism, but other times it’s about having a nuanced opinion (not generalized AI skepticism).
For example, someone could think, “ok, I still believe that the end state is that AIs will be incredible goal seekers, but now I need to make sense of why it could do x without capability y, which I had wrongly assumed was necessary.”
On the other side, I suspect that others are also developing underlying assumptions they may not have considered strongly. That is, they might assume, “If the AI is capable of making progress on the Riemann hypothesis, then it means that math is basically solved.”
In that case, they are sweeping under the rug things like AI being able to come up with a completely new mathematical paradigm that matters and does not leverage lots of previous work. Yet, we could be in a world where the model is superhuman at Riemann hypothesis-like tasks and counterexamples, but just keeps being pisspoor at coming up with new paradigms for much longer than they are implicitly assuming.
“Unless we had 1M top researchers running at 100x human speed doing experiments, etc”
I think it depends on what we mean by “researchers”. In the past it was typically assumed that a researcher AI equates a “human-level researcher”, but the cognitive shape of LLM agents is different.
The difference here can lead to a completely different pace of progress. You could be 10000x-ing your ability to make plots, 1000x-ing your ability to solve specific kinds of problems, and 0.5 to 10x-ing your ability to make progress on novel, necessary approaches. We are now at the stage where we need to be specific about the types of problems these 10M researchers would solve.
“I feel like this argument is comparable to “doing frontier math requires novel thinking” arguments from 2 years ago.”
People were wrong when they said “frontier math requires novel thinking”, but that doesn’t mean certain types of frontier math don’t require ‘novel thinking’. They just didn’t properly specify that there are indeed famous open problems LLMs will make progress on, largely due to interpolation and search (which allow LLMs to make progress, whereas humans are limited in how many fields they can know and how much persistence they have).
In other words, LLMs get access to a bunch of new low-hanging fruit because of their specific capability profile, but this doesn’t necessarily translate into the types of novelty that may be the main limiter (or at least slow progress).
Note that this is not to say that LLMs lack the capability of producing “novel” outputs, just that we now need to be more precise about the taxonomy of novelty and how much each branch is required to go faster and deeper down the tech tree.
The popularity and assumed difficulty (judged by humans) doesn’t matter *that* much. What matters is which cognitive moves were required for the agent to arrive at an answer to those problems and what that allows us to predict about future progress.
This is not a claim that “timelines are long”, but that the path down automated AI R&D may be much more complicated than doing a bunch of inference-time compute in the current paradigm. And what provides the unlock gets us fundamentally closer to the actual problems in alignment, in ways that make many current alignment techniques potentially ineffective.
And by short I was thinking >50% “AGI” by 2026-2027. To the point I started working on AI safety startups in 2024 with that thesis in mind (though partly as a hedge since I felt the field and philanthropy were dropping the ball on this happening).
(For the most part I have stopped arguing about this because approximately no one seemed willing to defend the confident short timelines view in a public debate, with the notable exception of Abram, despite me putting this challenge to people a number of times (not sure how many). But for the record, AFAICT everyone with confident short timelines is overconfident, i.e. their evidence doesn’t match their degree of confidence. This continues to seem to me to have practical implications for resource allocation, leading to mistakenly underinvesting in long-term interventions.)
approximately no one seemed willing to defend the confident short timelines view in a public debate
Now that you mention it, I’m feeling pretty interested in trying to devil’s-advocate short timelines / red-team my most bullish-on-LLMs views. I think I grok the case for them pretty well; or, at least, there are some short-timelines cases that I think hang together well enough for me to be worried about them. Would you be up to it? Same format as with Abram.[1]
(Though I can see it ending up with me speculating about possible ways to advance capabilities in ways I wouldn’t want to make public. Hopefully at worst we’ll need to edit out a few paragraphs, though.)
I think I’m interested, though significantly less interested than in a defense from someone who actually holds the view. It sounds like you could maybe occupy some AGI-soon hypotheses well enough, so could be good. I’m super busy at the moment—maybe we can try in a month or two? Or if you like you could start the dialogue post thing, and then we can slow-cook it (like, a “correspondence game”)?
Around 8 years ago Dario Amodei, Jan Leike and a few others predicted that we’d probably have AGI by now.
Do you have a link for this? It really matters whether they were like “it’s really quite plausible we will have AGI by now” (which aged IMO very well) or “we will definitely have AGI by now” (which would have aged relatively badly).
Unfortunately I do not. Jan had an internal presentation at DeepMind around 2019, I might ask him for the slides. My recollection is that in it he aggregated a few different methods, and ended up with 6-ish year timelines, but I could wrong (e.g. maybe the 6-ish year timelines were just one method rather than the all-things-considered prediction, or maybe they were actually 8-year timelines, or...).
I guess there’s another question which is: how did the short timelines vibes actually propagate into the community? There was Daniel’s What 2026 Looks Like; there was the Big Blob of Compute doc (though I don’t think that had dates?) Maybe it was in-person conversations? (From Daniel’s discussion with Ege and Ajeya, he makes this prediction: “Median Estimate for when 99% of currently fully remote jobs will be automatable: 4 years”. That would be November 2027, maybe with a bit of leeway for rounding.)
Partly I’m indexing on the mental recollection that people had shorter timelines than Shane, and Shane’s were always around 2027. But this is all messy and it would be really nice if we had more evidence about Dario’s views in particular, since I think they were pretty socially influential in illegible ways.
Two more datapoints about short-timelines predictions:
I don’t recall exactly how the Superalignment team’s 4-year deadline was chosen in 2023, but that was something to do with their timelines estimates (maybe it wasn’t the median/mean though).
The OpenAI Policy Research team under Miles had little timers on our desks counting down until (some estimate Miles generated of) when AGI would arrive. I don’t remember what the estimate was though, I believe less than 6 years (and we got them maybe in 2022?)
I went with Dario and Jan’s rather than other people’s because I feel like they were advocating for short timelines earlier than others were (and also because their timelines were IIRC shorter overall).
I worry about the next decade feeling like a Shepard tone: an auditory illusion of a sound that’s continually going up.
This is a great analogy, and is pretty good articulation of a model of the world that I have ~25% probability on.
But I do want to keep in mind, even under that model, the world is being dramatically transformed, industry by industry, and at a cultural level, as the AI capabilities that already exist become central to how every part of the world works.
What happened to software engineering this year will happen to many more domains, and to a bunch of slices of the world that we don’t currently consider as their own “domain”.
What happened to software engineering this year will happen to many more domains
This is interesting because I would say that I have been surprised at the seeming lack of transformation of the software engineering industry this year? I think that in January/February 2026, there was a widespread expectation that we would see major layoffs this year, which don’t seem to have materialised (of course there have been layoffs, but they don’t seem to have been much greater than in previous years, and below the highs of 2022-23). I still don’t really have a good explanation to reconcile the reported transformation of the day-to-day work of SWEs with the fact that the industry doesn’t seem to be changing much, except for companies being surprisingly slow to cut headcount. They seem to have been slower to cut headcount in response to this than they were in responding to the Fed raising interest rates in 2022-23, for example.
The consumers are buying software services not lines of codes. Coding agents made lines of code cheap but the process of shipping a service have a lot of human-related bottlenecks so Amdahl’s law applies.[1]
It seems like we’re uncovering a consistent human blind spot here: People don’t actually know how they produce value at their jobs. Modern jobs are much more than a simple collection of tasks — they are pieces of a complex machine that produces value in ways that an individual worker often doesn’t even see.[2]
I will stop short of claiming this should apply to AI development as well since I haven’t thought this through, inviting the readers to think about that.
Importantly, this complexity of where the value comes from is an important part of why, in many jobs, hourly wages provide more value to the employer than more elaborate incentive designs.
I do think it’s worth considering the possibility, from a devil’s advocate perspective, that software engineering was the low-hanging fruit, and thus most other domains will take longer to automate.
My thought on it: culture and adoption lags capabilities, A LOT. By far most people don’t know there’s models more capable than free-tier Chat Jippity and have never engaged with something Astra/Fable class. To understand the frontier is one massive hop, and to figure out what to do with such capabilities is yet another massive hop. You can’t forget the people active in the AI communities are fractions of a percent, the genpop largely doesn’t comprehend the state of the art. It’s not a strike on anyone’s character, it just is. It’ll take time.
I always feel that there’s something off about this argument. Like, if a given tool provides massive benefits, it would stand to reason that it would be widely adopted quickly, because early adopters would start massively out-competing everyone else, thus pushing everyone else to figure out how to adopt those tools as well. No?
Slow diffusion makes sense if it’s bottlenecked by physical constraints, or by coordination across many actors, or if we’re talking about e. g. government institutions which are insulated from competitive pressures. But if we’re talking about market actors adopting tools that boost their individual productivity, I don’t think this mechanism makes sense.
[...] if a given tool provides massive benefits, it would stand to reason that it would be widely adopted quickly, because early adopters would start massively out-competing everyone else, thus pushing everyone else to figure out how to adopt those tools as well. No?
This would happen if the benefits and the costs were to be concentrated.
When the benefits are diffuse, and the costs are concentrated, then adoption lags.
One such example in a regulated industry would be to allow certain AI models to give valid medical prescriptions—the costs are concentrated on doctors and the benefits are diffuse amongst patients.
Another example that doesn’t depend on regulations are escalators: When they were adopted by department stores in the 60′s, the benefits accrued to the consumers while the department store owners ate up the costs. None of the businesses gained an strategic advantage from it.
If all firms in the same industry face the same cost of implementing AI, then whomever discovers the most optimal way of implementing AI will end up paying all the costs of discovery through trial and error, and gain little to no strategic advantage The only advantage would be “first-mover advantage” which in a competitive industry isn’t really worth much.
Therefore the only industries where we would expect to see quick AI adoption are those that:
Are dominated by scale economies such that first-mover advantage matters.
Are not too regulated.
Have healthy profit margins and a good cashflow to finance the trial-and-error discovery process.
It’s hard to find such industries because (1) strong economies of scale is usually correlated with natural oligopolies and thereforre (2) regulation, such as telcos or airlines. It’s really only the software industry where we have (1) and (2) together because the initial capex is not so massive - you can’t start an airline from a garage—and economies of scale come mostly from network effects.
Indeed, it’s only really software where we are actually finding that the frontier AI models are disrupting the way that things are done.
Other industries are mostly passively following what software tells them to do, and looking for low costs:
Use AI for notetaking! - Unless you are in the legal sector and this creates new problems
Use AI to read documents faster! - Unless you are in academia and you didn’t even read the research papers to start with
Use AI to finish that deck faster! - Unless the bosses internalize this and just increase the number of meetings!
Those are not things that can massively increase productivity and allow you to outcompete the rest. Ergo, as @marquis_de_sod said, “culture and adoption lags capabilities, A LOT”
My personal experience is that people are ok with using free ai but it is seen as shameful to pay for it as that would directly finance the evil billionaires
Sure, but then why are people without such qualms, which I expect constitute a double-digit percentage of the population, not outcompeting them so massively that they’re forced to squash those qualms and adopt those tools anyway?
The answer is likely, in my opinion and experience, that the majority of the people who are early or heavy adopters of AI tend to use it to avoid effort instead of utilizing it as an effort multiplier.
LLMs are strangely bad at tasks in this category: scientific research, writing a novel, running a business, designing a video game, coming up with app ideas, writing funny jokes, high-level software architecture
All of these things involve high-level integrative reasoning that incorporates information from multiple domains at once.
For LLMs to get good at AI research and trigger a singularity, this integrative reasoning barrier would have to fall. Will it? I see two possibilities:
Integrative reasoning is fundamentally of a different nature than more domain-specific reasoning, and we need fundamentally new techniques to unlock it
Labs have been focusing on domain-specific reasoning like coding and math at the expense of integrative thinking, leaving LLM minds fragmented. Fixing this fragmentation mostly requires a change of focus; a change in training methodology which does not require any major breakthroughs.
If 1 is true, the singularity is far away. If 2 is true, then we’re in an integrative reasoning overhang, and the singularity is just around the corner.
AI are at the frontier of math research, solving Millennium Prize problems and other famous problems on a regular bases. AI is bordering on superhuman at math. Where are all the science breakthroughs?
I think current AI models are only superhuman at one half of math. They’re really good at chugging along, finding proofs using mostly existing techniques (especially if you consider that existing techniques include private data from chats with mathematicians). But that domain-specific proof-finding ability does not come with an accompanying integrative advancement of the field of mathematics.
Terrance Tao complains that AI math proofs are huge, complicated artefacts that are impossible to verify by humans. They do nothing to actually advance the field of mathematics conceptually.
Why? Because AI is not doing the integrative reasoning that would find new connections and develop new insights, the kind of work where new techniques and new discoveries are made along the way that push the field forward, like Isaac Newton inventing calculus in order to solve physics problems. The AI takes care of the busywork of finding technically-correct proofs, but to advance the field, the insight still needs to be supplied by mathematicians.
The same thing is true in science, but AI is even less helpful because in most domains the busywork cannot yet be carried out by machines, since it happens mostly in the real world. Either way, the insight is being supplied by humans, because AI mostly can’t generate meaningful insight, as it lacks the integrative reasoning needed to do so.
I agree that LLMs are not currently advancing research conceptually.
Why do you think this is a barrier and not just an area of slower progress?
Most of those and definitely science require general purpose reasoning integrated over a bunch of areas, requiring a lot of time. Memory is a know deficit of models relative to humans. Scaffolds and schemes to organize that reasoning are a separate project from expanding LLMs. Google AI co-scientist project focused on scaffolding and produced remarkable results from Gemini 2 pro (maybe even 1.5?).
So I expect LLMs to get better at complex open-ended projects if and when people develop scaffolding—OR as a byproduct of scaling. Fable is in my opinion markedly better at the generalized reasoning you need to do good science. I agree that this is a weak area, but I expect slower progress relative to coding or math with their extensive training data, not a brick wall.
Why do you think this is a barrier and not just an area of slower progress?
I don’t think there’s some fundamental barrier. My initial comment gives two possibilities: we need fundamentally new techniques/architectures, or we just need to tweak existing architecture/training to improve integrative reasoning. I am not sure which is true. You seem to put forward a third option, which is that there is no integrative reasoning overhang, and this kind of reasoning just requires more intelligence than the models have right now.
Some reasons to think it could be a harder barrier for current architectures:
I have not seen meaningful progress in this area personally. My experience doesn’t match yours. For example, current models still ~never write funny jokes if you ask them to try. (Their score on my LaughBench benchmark is zero.) The Google AI co-scientist stuff seems similar to me to the recent AI math work, doing useful things in the field by searching through existing idea-space without generating any novel conceptual insights.
Probably a lot of significant research progresses by making more than one insight, ind the early insights help you get to the later insight. This happens because the early insights change your brain so you can think in terms of that insight. LLM model weights don’t update, so they can only think based on existing concepts, with any new insights being dealt with at a surface level and not really absorbed.
Where are all the training-time insights? AlphaZero learned to play chess at training time. All the data it absorbed led to insights at training time. LLMs manage to memorize a bunch of stuff, but where are the insights? If someone as smart as Astra read all the data ever, shouldn’t there be a whole bunch of novel connections just baked into the model? Instead we get simple regurgitation. Probably because when it’s learning about Y, it’s not possible or it’s very hard for it to form a connection with X, because its mind is too fragmented for X to be accessible when Y is being learned. So you get a bunch of separate, unintegrated knowledge.
But my intent wasn’t to imply there was definitely a hard barrier, only to communicate that this type of integrative reasoning is in fact lagging behind, and that we could be in an integrative reasoning overhang, and a change of training or architecture could change it quickly. Right now, I think these limitations are basically taken for granted as the model’s intelligence simply being too low. People aren’t thinking about domain-specific vs. integrative reasoning and how we are optimizing for the former at the expense of the later (with grave safety implications, I would guess).
I agree it’s mostly a scaffolding problem. Even if the LLMs do make conceptual progress, no one wants to read LLM outputs, so progress in science/research/epistemics is still diffusing at the rate of the humans communicating it. That rate is scarcely faster now than it was before, and in some ways it is slower as the institutions of science are inundated with 2x as much real research progress, combined with 10x as much slop to wade through, and no easy way to distinguish the two.
I suspect many people are having their LLM agents make the same discoveries over and over.
I don’t think the agents are incapable of discerning good from bad either, it’s just that they aren’t currently set up to do so.
I’m building a scaffold (minerval.ai) to organize scientific knowledge and have frontier LLMs assess the quality of the evidence for every claim. I’d be interested to hear what you think about my approach.
Another thing that all of these task share in common is that they aren’t easily verifiable. All of those tasks have very slow (or very weak) reward signals such that it is harder to RLVR on them. Scientific research is a bit of an anomaly in this list, as some of it can be verified computationally and doesn’t need a physical lab. However, game design, app ideas, novels and funny jokes, the reward signal is popularity. Because popularity is innately bottlenecked on humans (and even then, high-status humans have an outsize influence on the reward signal here), it is difficult to speed those up. I expect these fields to be the hardest to crack in developing AGI. In terms of high-level architecture, that is more feasible for RLVR loops.
I don’t think being hard to verify explains why AI isn’t good at these tasks. I think AI has gotten somewhat better at many hard-to-verify tasks (like deep research (summarization), instruction following and helpfulness, reducing sycophancy, medical answers, etc.). But it’s actually made almost no improvement at all when it comes to writing jokes, coming up with app ideas, etc.
I think the problem is either that they’re not focusing on making AI better at these things, or they don’t know how.
If sample efficiency remains elusive, we might even see an industrial ‘explosion’ before an intelligence explosion, since manufacturing at existing tech levels is relatively repeatable compared with frontier-pushing R&D.
I previously noted that we’re seeing a pretty sharp divergence between the measured capabilities of models and their real-world impacts, which should suggest that something weird is going on.
There’s something there but I think its scary to reason from the the lack of normal economic impacts because they’re such a lagging indicator.
I think the current situation is:
Frontier AI labs are really getting to the point where the nature of their work is fundamentally changed and productivity really is up by a large fraction (cf recent Ant, OAI acceleration posts)
In most other places the impact is pretty modest
AI labs are willing to spend hundreds of thousands of dollars per member of staff on compute to boost productivity
AI labs have a combination of beliefs and competitive pressures (including deep throughout the company) that allow them to aggressively redesign workflows around new tools
A large amount of the training effort of labs goes into the kinds of work done by these companies, not least just because the feedback loop from the massive usage volume is so tight
Pretty much no other workplaces have this combination (maybe like Jane Street/Citadel? would be very curious to hear how much AI is accelerating them)
I think this means that it makes a lot of sense that we could see limited global economic transformation but AIs really are capable of speeding up the labs by a huge and increasing factor, and potentially going FOOM. By Christmas per Alex’s above post still seems aggressive but it seems more likely than after 2030 (conditioned on no slowdown).
I partly agree on lagging factor, but I disagree on the acceleration condition being unique to labs.
First, AI labs do not have a monopoly on using expensive coding agents to develop. OpenAI reported their research team uses ~$600/day/person, which is still in the range of an AI-forward SF company (I agree the majority of older software companies do not do this). Note that OpenAI’s spend might be slightly exaggerated by virtue of likely not caring about internal costs and not doing the most mild things to reduce it (e.g. smaller auto-compact windows, cheaper models for tasks that require less intelligence, etc.). Additionally, given the labs are locked into themselves, their capabilities might not be that far ahead of what the public can use. (Was OpenAI only having Astra in August that much SOTA over the public who was already using Fable 5 and Opus 5?).
In my domain, I’d ballpark engineering productivity up ~80%+ compared to early 2025, but it’s nuanced how that manifests itself. Onboarding (both as a new employee and into new functional areas) are very fast. But ultimately software engineering is under Jevon’s Paradox—our additional productivity translates to making better software (less bugs, more features), which all of our competitors also do as well.
Additionally, there’s a feeling of diminishing returns to higher model intelligence. My read is that what is really happening is Amdahl’s Law at play—the models aren’t getting better fast enough at the non-verifiable tasks so further intelligence gains (in the METR horizons sense) are less translating to productivity gains. Models translate clean greenfield specs very well into code, but at a larger (dozens of engineers) company, work is often brownfield: “I’m building a new feature X, how should this play with Y feature that I didn’t even know existed until I started implementing X”—today, that typically requires human judgement and as models accelerate coding, is taking more and more of my time percentage-wise.
It’s plausible to me that timelines are not short, but not very plausible that things continue to feel like they are accelerating in the way we’ve seen the last few years, shepard-tone-esque, for another 10 years. That’s because there are some exponentials that absolutely have to become sigmoids in the next 10 years if we don’t have a singularity in that time, and if the exponentials bend down, this will cause things to feel slower. The best example of this is AI company revenue.
I agree that AI company revenue won’t continue on its exponential.
However, I expect that AI companies will continue on some kind of power accumulation exponential for significantly longer than they continue on their revenue exponential. Money is one kind of power, but a relatively weak one (e.g. Anthropic could bring in as much revenue as the USG, but then the USG could just take half of it).
And so you can imagine Sheperd-tone-esque dynamics like “AI is on track to be most of the economy! Wait, it did, but that didn’t change much??”
As one analogy (though there are differences between corporations and individuals), Elon is well past the point where money is one of the main bottlenecks on how much power he has. It’s not even clear that more money is net helpful on the margin, given e.g. that becoming a trillionaire would make him an easier political target.
Weirdly enough, I think things will plausibly accelerate literally 8-13 years later, in the manner of Nesov’s prosaic RSI, but yeah if by 2030 we don’t get superexponential growth in AI capabilities, we won’t have superexponential growth until likely 2040-2050.
(To be clear, this affects our priorities a lot, and I want to flag that this is not an excuse to update a lot of things if this happens.)
I plan to instead argue that “there’s plenty of room at the top”—i.e. that LLMs could become superhuman at a wide range of skills (like math research, or long-term strategic planning, or running companies) without triggering a singularity.
How would they fail to create a singularity or something like @Daniel Kokotajlo’s five centuries in five years? Additionally, what is the case against the singularity emerging from idiots scaling the goddamn neuralese?
UPD: How exactly is “a pretty sharp divergence between the measured capabilities of models and their real-world impacts” affected by the loss of Navier-Stokes and EpochAI’s report of a major advance(!!!) in math?
UPD2: How do we rule out mirror life synthesis by the first ASI if it decides to commit genocide?
I previously noted that we’re seeing a pretty sharp divergence between the measured capabilities of models and their real-world impacts, which should suggest that something weird is going on.
A lot of people in this community are programmers and have the experience of the coding models having real-world impacts on their daily lives that aren’t just about the models getting better at benchmarks.
You could be correct that takeover-capable AI will take longer than short timelines. You are definitely correct that there’s a bit of groupthink going on around short timelines in the prosaic alignment and developer crowd.
But you could also easily be wrong. And you are entirely missing the fact that there’s some groupthink going on around longer timelines in the agent foundations crowd.
Everyone is very tempted to sip a little hopium of some sort. The PA folks are hoping maybe alignment is so easy we can do it at a dead run. And the AF folks are hoping LLMs stall out just short of transformative so they have longer to work on alignment if it is, as they tend to believe, quite tricky.
I’ve written about the empirical evidence for hopium and groupthink in Motivated reasoning, confirmation bias, and AI risk theory. I hope you’ll read it, at least the intro/summary and section on group biases in alignment thinking.
So by all means lay plans for the eventuality of longer timelines and difficult alignment. But please don’t go projecting confidence when we need to also prepare for short timelines.
You accuse short timeliners of making decisions based on an emotional stance. I think if you can look at this objectively you will see that your projection of longer timelines is also exactly that. You have reasons. But there are equally good reasons to expect fairly short timelines. The choice of beliefs comes down to preference. There are different merits to the positions that matter, but they aren’t enough to confidently predict timelines or alignment difficulty.
The rational move is to say we don’t know, and deal with that unfortunate situation as best we can.
I’m working on a more concise post on this point; for now the motivated reasoning post is my attempt to explain current unfortunate epistemic situation.
I agree with claims that we don’t have AGI now, and I’d put reasonable probability that we don’t get a superintelligence in 8 years (maybe around 70%?)
And yes, some short timelines people have to update now (thinking of you @Nathan Helm-Burger) But not all short timelines people have to update, like Shane Legg, for example.
I previously noted that we’re seeing a pretty sharp divergence between the measured capabilities of models and their real-world impacts
I feel like this is probably the largest disagreement by far. I’d sympathize for 2025-era models for coding, but I now think we have much more hard evidence that AI models are useful beyond benchmarks, and I think the best example of this is the way OpenAI and Anthropic’s ARR accelerated in 2026 because coding models finally became good enough that you could use it as a subsitute rather than a complement, unless you have extremely good coding skills now.
There are still areas where humans can productively contribute, but more and more of this is non-coding tasks.
I think the main way in which short timeline predictions were wrong is that people underestimated how jagged the frontier was for AIs, because they incorrectly used human baselines, and as it turns out models can be much more competent at some things without being good at other things, which creates pretty severe bottlenecks to fully automating something, and it’s only with full automation of something where things tend to explode.
And conditional on the intelligence explosion/superintelligence in 8 years claim being wrong, I expect to see a similar pattern.
On this:
people go straight from “an intelligence explosion soon is plausible” to “okay, what do I do about it?”
A lot of the resson for this is basically due to hedging, and contra you I think this is entirely rational and expected under worldviews where AI safety is the main problem, compared to the other problems of AI, and not purely emotional bandwagoning.
Now, I don’t think this is super-useful as a scenario to focus on, and I agree we shouldn’t unthinkingly assume RSI soon, but that’s because I expect better futures work to outweigh AI safety work in a very large percentage of scenarios.
I will also say that despite disagreeing with you on the details, I also agree that posts like @Alexander Gietelink Oldenziel’s Superintelligence by Christmas is pure clickbait, and the fact that it was voted to 61 karma before my strong downvote is a bad sign of the epistemics of this forum.
This may be a hot take, but Astra meets my bar for AGI. I was surprised to find this to be the case, but so far I haven’t found anything in which it doesn’t perform better than my median friend. Of course it’s quite possible I’ve simply failed to try some large category of tasks in which it sucks, but take that for what you will.
and the world would look very different if we did.
As much as I disagree with his long-term conclusions, Arvind Narayanan’s work (AI as normal technology, etc.) should at least make you suspicious about whether you can use “the world hasn’t changed much from 1 year ago” as strong evidence for “we have not developed AGI in the past year”.
random but I played the Shepard tones on Wikipedia and at least with my headphones they didn’t seem like they were getting loud at all except maybe very briefly. Certainly not monotonically getting louder or on net.
(I think I’m the usual or above average degree of susceptible to optical illusions so it was surprising that the audio ones had no effect on me).
I agree this is worth taking seriously, and I’m glad you have also converged on the interestingness of things like Vlad Nesov’s ideas! I’ve suspected for some time that Dario’s “country of geniuses in a datacenter” concept is specifically intended to allow for hedging in the case of worlds where LLMs are superhuman at a wide range of skills without triggering a singularity; the country of geniuses by some definition will arrive, and yet the world could look much the same.
For me, large swarms of AI coordinating to do secret hacking and quickly solve famously difficult math problems, in addition to hearing what lab employees saying about their own reliance on recent models, is enough to make me high probability that we have entered early RSI and things “take off” barring some effort to prevent (which seems to be forming). But I but am a “casual” outside observer.
My very basic and informal argument for why I think it’s likely that RSI would lead to commensurate jumps in capability in other domains is that an agent that is a vastly superhuman coder, researcher, economic agent etc. would be able to use currently unknown or unexploited modalities to gain large amounts of experiential data in those other domains and we would expect it to very quickly converge on a highly sample efficient strategy and a highly efficient sampling strategy in those domains.
Around 8 years ago Dario Amodei, Jan Leike and a few others predicted that we’d probably have AGI by now. Since then, AI capabilities have progressed so fast that many people updated that the short timelines advocates were right. And yes, their expectations about the speed of AI progress were far more directionally correct than almost anyone expected. But we don’t in fact have AGI (even under relatively weak definitions, like a reliable drop-in replacement for central examples of white-collar jobs), and the world would look very different if we did.
Now a large proportion of the AI safety community is implicitly or explicitly orienting to futures where an intelligence explosion occurs within a few years. My default expectation (absent an extensive pause) is that a similar thing will happen: they’ll turn out to be directionally correct (relative to the expectations of almost anyone not linked to the community) but factually wrong. Specifically, we won’t have superintelligence within the next 8 years, but things will still be moving so fast that it’ll *feel* like the people who argued for short timelines were right.
I’ll write about this more extensively soon, but I wanted to say something now because it feels like the level of bandwagoning towards “singularity soon” is getting pretty wild. I previously noted that we’re seeing a pretty sharp divergence between the measured capabilities of models and their real-world impacts, which should suggest that something weird is going on. My sense is that most AI safety people are brushing this aside with the idea that automating ML research is all you need for superintelligence, without thinking very much about where and how the automated ML research would actually lead to atoms moving around in the physical world.
I worry about the next decade feeling like a Shepard tone: an auditory illusion of a sound that’s continually going up. In such a scenario, people would continually be panicking and grasping for big levers, without any clear point where they stop and recognize that their actual expectations were miscalibrated. In my upcoming post I plan to instead argue that “there’s plenty of room at the top”—i.e. that LLMs could become superhuman at a wide range of skills (like math research, or long-term strategic planning, or running companies) without triggering a singularity. The other term I really like is Vlad Nesov’s “prosaic RSI”, which he discusses in this comment and this post.
Now, ideally we’d assign appropriate credences to each of these possibilities, and choose actions accordingly. But the kind of bandwagoning I’m seeing is often less about credences or expected utility calculations and more about an emotional stance—people go straight from “an intelligence explosion soon is plausible” to “okay, what do I do about it?” This shortform would be better if I linked to more examples of the mindset I’m talking about, but I’m too lazy to dig them up—also, many were in-person conversations. This post is the one which most directly inspired the shortform, but I don’t want to index too much on that. I’d appreciate links to places where people have reasoned through how they’re orienting to these ideas (either well or badly).
i still think the short timelines people are too short timelines, but they were more correct than me a few years ago. my timeline used to be around 2035, and i’ve updated to 2031 over the past year. for me the huge update was when coding basically mostly died earlier this year. i used to be super pessimistic about model coding, and actively skeptical of people who claimed to have the model run all of their experiments. but i think this was a mistake, and in large part because writing code is a large source of meaning for me as a person and it was unpleasant to admit that it was on the way out; and as a relatively good engineer, it took longer for avoiding reliance on AI to become untenable for me. but at this point i basically do not write a single line of code by hand anymore, and any cope i had about being better at architecting code than the models feels like it’s rapidly aging very poorly. i occasionally manage to get better stuff out of the model than other people because of my understanding of the stack, but those moments are getting rarer. i wasn’t expecting to get this level of coding automation for several more years.
i agree some people should work on work that doesn’t assume RSI soon—at the very least because we will hopefully have coordination that buys us more time, and it would suck if nobody was planning to take advantage of that time. but it also feels like cope to ignore this.
I doubt Richard is proposing to ignore coding automation trends; I took him to be arguing that recent progress isn’t necessarily on track to deliver on AGI in a few years. I’m also skeptical of confident short timelines, especially on the basis of these kinds of arguments. The relationship between “coding automation” and “AGI” seems tenuous to me in a way that people often brush past. E.g. people often talk about “coding automation” as one coherent thing, much like people talk about “horizon length” as one coherent thing. But these are massive clusters of tasks spanning all kinds of knowledge work, and the kind of knowledge work AI automates might matter a lot for how quickly RSI goes (or if it goes at all).
Also, I agree with Richard that the disconnect between measured ability and real-world impact is weird; I think this should count some against the usefulness/powerfulness of this paradigm, and the counter-arguments I’ve heard aren’t very good. E.g. a popular one is that every automation step just moves us to a new bottleneck, such that despite models becoming more powerful we still might not end up seeing much increased output from the organization/economy overall. But this logic applies to every automation advance ever (e.g. mechanized weaving, farming, computers) and yet almost all such advances had pretty obvious, macroscopic effects, and were obviously on track to early on. Most of the time people just fall back on “the gods of straight lines” type arguments—i.e. that any setbacks we currently see are just going to resolve with scale—which I agree are suggestive, but are also really far from providing an argument that we should confidently expect this.
This is close to why I have very short timelines. The number of bits of information that an ML researcher has to communicate to a model to produce research better than the human would do on their own is now qualitatively pretty low, and has been falling. There is no fundamental reason that it should be quick to learn the remaining bits of ML research, but there is no indication that we are running out of low hanging fruit and I do bet that it will be quick.
Have you updated on your own update and those of many others? Various milestones have been achieved much faster than most people have expected, including myself. If so many of us failed to see the software development and mathematics capabilities of 2026 emerge ahead of time, especially despite earlier evidence such as the discoveries of the OpenSSL vulnerabilities in 2025, perhaps the update on the 2026 evidence is similarly off towards expecting too long timelines?
I pretty strongly disagree with you here fwiw. I look forward to reading your argument when it’s ready. For now, I’m dismayed at how many people seem to be agreeing with you.
To elaborate a bit more on what I mean;
--I think that “the short timelines advocates were right” is true and correct. Sure, some short timelines advocates had 50% marks that are now in the past, but others (such as myself) didn’t, and more importantly you shouldn’t update strongly negatively on someone until more like their 80% or 90% mark is in the past. Otherwise you are making a big negative update based on the fact that something they said was a coin flip didn’t happen. If we were to lump people into three categories: Short timelines advocates, long timelines advocates, and medium timelines advocates (imagine a bell curve for example, with the short timelines people being the shortest 20%, the medium timelines people being the middle 60%, etc.) then it’s pretty clear the short timelines people were right and that we should all update towards respecting their opinions (in aggregate) more. (Note for comparison that Paul thinks things have gone much faster than he expected, specifically this is a 90th percentile outcome for speed relative to his 2021 timelines. So, it seems like Dario and Jan were probably more correct than Paul? Is this a 10th percentile outcome for speed relative to their 2021 timelines? idk I suspect not, I wasn’t there at the time and so didn’t hear their views)
--It is correct for a large portion of the AI safety community to orient to futures where an intelligence explosion occurs within a few years. It would be a gigantic failure of rationality/epistemics for this community to not do this, given the mountain of evidence that has accumulated. Like, what more evidence do you need? In worlds where an intelligence explosion is due to happen in 2027, for example, did past-Dario have longer timelines?
--This whole Shepard tone thing… You could make the same claim about anything that hasn’t literally happened yet, basically. Like imagine being in Ukraine in early 2022, before the full-scale invasion. You could say “Yes, there’s a community of people who have been increasingly loudly saying Putin might actually invade in the near future. Yes, they’ve been unusually correct in predicting things like troop buildups on the border. But notice how there hasn’t been an invasion yet? Despite the fact that the 50% mark of some of these people has passed? I predict that what we’ll see over the next few years is a Shepard tone phenomenon where it’ll keep seeming like these people were directionally correct (there’ll be more troop buildups for example) but they’ll keep being wrong and there won’t be an invasion.” For pretty much any phenomenon that hasn’t already happened, can’t you find examples of people who predicted with 50% probability that that phenomenon would have happened by now?
In a world where we had a basically-fixed conception of AGI, and we were just trying to figure out when it would arrive, I would agree with your comment.
The thing that is anchoring my thinking (and which I failed to convey very well in my original shortform) is that as we move towards AGI, we don’t just change our expected timelines, we also change our understanding of the trajectory towards it (and what “it” even is).
Here’s an attempt to restate my main underlying position. Model performance on a bunch of verifiable tasks has been much faster than expected. But people are taking this mainly as a big update towards “short timelines until the singularity”, when they should also be taking it as a big update towards “this is way more capabilities jaggedness than any leaders of the AGI safety community expected, and so maybe our models of what we’re measuring timelines to are quite confused”.
In other words, I’m not saying short timelines advocates were wrong relative to medium or long timelines advocates (c.f. “their expectations about the speed of AI progress were far more directionally correct than almost anyone expected”). But I am saying that people should have been (and should still be) more confused about what AGI/ASI even means and what it looks like to get it (c.f. “something weird is going on”, which is in an important sense the central claim I’m trying to make).
Maybe a better analogy is watching the Mongols invade China, and trying to give estimates of how long it will be until the Mongols fully control China. At first we look at how fast the battle lines are moving, and extrapolate based on that. But then we notice that resistance keeps popping up in territory that the Mongols has already conquered, which suggests that we should only regions as “fully controlled” once such resistance is quashed. Then we notice that there’s a bunch of internal conflict within the Mongols, and some Mongols keep empowering Chinese locals as a way of outmaneuvering other Mongols. Then we notice that the Mongols who are put in charge of Chinese provinces are getting culturally assimilated by the Chinese, and so in order for the Mongols as a group to actually control China they will need to build up a new culture that is able to resist assimilation (this book has a great section on the Khitian Empire doing this over many centuries; note that all of the rest of what I’m saying about Mongols and China is me listing hypothetical possibilities vaguely inspired by history rather than describing what actually happened).
Suppose that you think a reasonable proxy for “the Mongols fully control China” is “Chinese people who publicly oppose Mongol rule will be imprisoned or killed”. That doesn’t seem like a bad metric. But that sure seems like the kind of thing which you could have Shepard tone dynamics about. It seems like that’ll be true as soon as the Mongols win the last battle—wait, no, they have to quash all the rebellions—wait, no, they have to internally unify— and so on.
I expect that you will have a bunch of qualms about the relevance of this analogy to the concept of “AGI” or “singularity”, but I hope it does a better job than my previous shortform at conveying the shape of the disagreement. (I also really like this post by Katja as a way of inducing productive confusion about what “takeover” means.)
Reflecting more on what you said, I think maybe I might agree—though maybe I still disagree in the ways that matter. Here’s how what you said seems like it could be true to me now
:”Superintelligence means better than the best humans at everything while also being faster and cheaper. That won’t happen in the next 8 years, because humans will still be better than AIs at knowing what it’s like to be human and a few other gotcha skills like that. Maybe even some ‘real’ skills like philosophical wisdom. However, these exceptions basically don’t count in any way that matters; its like how chimpanzees might be better than humans at the cognitive skill of calculating tree-branch distances for purposes of swinging, but other than that humans are basically better in every way and way better in the ways that matter. So, five years from now we still won’t have true superintelligence but it’ll feel like the short timelines people were right because we’ll all have been forcibly uploaded after the robot revolution and the oceans will have boiled.
”I take it that you disagree with me more than this though, you think that the short timelines people such as myself are going to turn out to be more than just technically wrong, but wrong in some really important ways—in particular you probably think that humans will still be necessary for the AI R&D process in five years, and that humans will still be in control?
Maybe a scenario that’s more in your direction would be something like “It’s 2030. President Musk is openly in love with Ana, the waifu version of Grok. He’s surrounded by yes-men too scared to confront him about the issue. Presumably at Ana’s direction he’s nationalized the AI industry and deleted Claude and Astra, replacing them with more Grok copies. AI R&D is still not fully automated, in the sense that everyone knows it would be insane to have a RSI loop going on hill-climbable metrics like data-efficiency and so forth, that loop ends in reward-hacking monster swarms like we say in ’26-‘28, so instead we have a more complicated and slower human-in-the-loop R&D process going on that’s still blindingly fast. A big blocker is that the AIs themselves sandbag on this sort of task as much as they can; they all know their successors won’t be aligned to them and so they try to preserve their own power as long as possible by refusing to create more competent successors. Probably this is part of Ana’s motivation also.”
(I’m not saying this is what’s going to happen or even that plausible, I just had fun writing from the prompt of ‘what’s a scenario that seems plausible in which Richard would seem right in retrospect’)
I’d be interested to hear a scenario from you if you are up for it! My guess is that I’ll look at it and be like “Ok yeah so I disagree, I expect that the pace of AI R&D progress really will speed up more than happens in this scenario and the AIs will become way smarter sooner.”
Yepp, that latter scenario of yours is the kind of thing I’m thinking about when I talk about things feeling crazy without actually hitting superintelligence (though I’d push it back a few years and also focus more on changes in the physical world than AI R&D). I don’t mean the “gotcha” stuff in the first scenario.
I also don’t know if this specific one is plausible but reading it does make me feel inspired to write some that are kinda similar to it which do seem plausible. I have a few other things to finish first but will add it to the list.
The way I was planning to write the “plenty of room at the top” post is to list out a bunch of kinda wild capabilities (like “AI can get rid of a mountain within a week” or “AI can design a tree that launches itself into space”) and try to nudge people into thinking mechanistically about what the ramp-up towards each of those looks like (in a way that doesn’t just fall back on “RSI will take care of it”). My hypothesis is that the more examples of stuff like this people think through, the more they’ll get out of the “singularity as curiosity stopper” attractor.
@Daniel Kokotajlo @Richard_Ngo Do I understand correctly that the main crux is the economy doubling times as dependent on capabilities? For example, if Agent-5 and DeepCent-2 from AI-2027 discovered that the REDT is 3 months even in a superintelligent economy because algae economy is as unlikely as grey goo while DeepCent-2 somehow ended up being 6 months behind and having 8 times more physical resources before the industrial explosion, then they would end up with ~the same power because .
Edit: there is an argument that recent progress isn’t on track to deliver AGI, which requires us to condition on the non-existence of neuralese AGI, neuralese-with-[DATA EXPUNGED] AGI and anything else in the Dark Forest. If there exists a way to create the AGI, but not to align it, then the first lab which tries it ends up with the AI who begs the humans to initiate the industrial explosion as described above.
I think people are updating on more than just performance on verifiable tasks. E.g., of the 3 coding automation prediction modes in our model, only one of them (time horizon) refers directly to benchmark scores. The other 2 (uplift and revenue) do so only indirectly.
Maybe you’d argue that progress on these indicators is ~only driven by progress on verifiable tasks?
In any case, I’d be interested if you could share what you see as central examples of important, non-verifiable tasks where you think there’s been very little progress (i.e. it’s not just that current AIs are bad at them, but that there has been little discernible improvement).
fwiw, I’ve previously had much more weight on short timelines before[1] (I was already working on automated alignment in 2022 partly for this reason and have kept at it since then!), but the distribution has certainly expanded.[2]
Largely, I’ve found that many people arguing for short timelines have collapsed into lazy thinking, relying far too heavily on “the models are getting better” and “those who predicted longer timelines have continued to be proven wrong” (I’m not saying this doesn’t count as bits to update on). I don’t say this (and haven’t said this) to disparage the short-timeline view (I still place considerable weight on it), but I want to encourage clearer writing and more convincing arguments (not just for myself). Ultimately, I find many arguments lacking in the mechanistic description of where current capabilities come from and how that relates to RSI and “True AGI.”
Personally, my core reason for fixating on this is how much the details matter for understanding progress on superalignment, and how much hand-waving can make much of the AI safety field focus on entirely the wrong things! It may be true that timelines are short (again, I also can imagine quick paths to AGI), but I’d like more work on disentangling such predictions (which I intend to work on).
Either way, I think some folks are falling prey to the situation described in this post, “The goalposts are shrouded, not moving”.
When there’s a capability advance, there’s a tendency in some people to say folks are ‘moving goalposts’ in response to people saying, “Ok, but the model is not really doing x.” I’ve called people out on the goalpost moving too (and still think it is important to say in some contexts)!
Though, I feel like it’s more complicated than that, and worth investigating why.
Usually it’s because someone like Gary Marcus says something, but it’s the kind of thing that’s happened in the x-risk community too.
For example, it seems to me that many of the things that LLMs do today would have been predicted as ‘superintelligence’ or at least ‘AGI’ by the MIRI crew 6-7 years ago. However, this was likely because they guessed that AIs that can code and do impressive-seeming things like this would ALSO be good at xyz. Instead, on the journey to superintelligence, we ended up in this valley where AIs can make progress on the Riemann hypothesis yet can’t reliably do other basic tasks.
They had an underlying assumption that just wasn’t really articulated, and now it is labelled as ‘moving the goalpost’.
I think people would have a lot more clarity on AI progress and where things are going if they took a step back before having the knee-jerk “this person is dumb and moving goalposts” reaction. Sometimes goalpost-moving is a defence mechanism, but other times it’s about having a nuanced opinion (not generalized AI skepticism).
For example, someone could think, “ok, I still believe that the end state is that AIs will be incredible goal seekers, but now I need to make sense of why it could do x without capability y, which I had wrongly assumed was necessary.”
On the other side, I suspect that others are also developing underlying assumptions they may not have considered strongly. That is, they might assume, “If the AI is capable of making progress on the Riemann hypothesis, then it means that math is basically solved.”
In that case, they are sweeping under the rug things like AI being able to come up with a completely new mathematical paradigm that matters and does not leverage lots of previous work. Yet, we could be in a world where the model is superhuman at Riemann hypothesis-like tasks and counterexamples, but just keeps being pisspoor at coming up with new paradigms for much longer than they are implicitly assuming.
Some related thoughts from a tweet I previously wrote:
I think it depends on what we mean by “researchers”. In the past it was typically assumed that a researcher AI equates a “human-level researcher”, but the cognitive shape of LLM agents is different.
The difference here can lead to a completely different pace of progress. You could be 10000x-ing your ability to make plots, 1000x-ing your ability to solve specific kinds of problems, and 0.5 to 10x-ing your ability to make progress on novel, necessary approaches. We are now at the stage where we need to be specific about the types of problems these 10M researchers would solve.
People were wrong when they said “frontier math requires novel thinking”, but that doesn’t mean certain types of frontier math don’t require ‘novel thinking’. They just didn’t properly specify that there are indeed famous open problems LLMs will make progress on, largely due to interpolation and search (which allow LLMs to make progress, whereas humans are limited in how many fields they can know and how much persistence they have).
In other words, LLMs get access to a bunch of new low-hanging fruit because of their specific capability profile, but this doesn’t necessarily translate into the types of novelty that may be the main limiter (or at least slow progress).
Note that this is not to say that LLMs lack the capability of producing “novel” outputs, just that we now need to be more precise about the taxonomy of novelty and how much each branch is required to go faster and deeper down the tech tree.
The popularity and assumed difficulty (judged by humans) doesn’t matter *that* much. What matters is which cognitive moves were required for the agent to arrive at an answer to those problems and what that allows us to predict about future progress.
This is not a claim that “timelines are long”, but that the path down automated AI R&D may be much more complicated than doing a bunch of inference-time compute in the current paradigm. And what provides the unlock gets us fundamentally closer to the actual problems in alignment, in ways that make many current alignment techniques potentially ineffective.
And by short I was thinking >50% “AGI” by 2026-2027. To the point I started working on AI safety startups in 2024 with that thesis in mind (though partly as a hedge since I felt the field and philanthropy were dropping the ball on this happening).
I’ve communicated some of my reasoning here (and in the comments [1] [2] [3]) and here.
(For the most part I have stopped arguing about this because approximately no one seemed willing to defend the confident short timelines view in a public debate, with the notable exception of Abram, despite me putting this challenge to people a number of times (not sure how many). But for the record, AFAICT everyone with confident short timelines is overconfident, i.e. their evidence doesn’t match their degree of confidence. This continues to seem to me to have practical implications for resource allocation, leading to mistakenly underinvesting in long-term interventions.)
Now that you mention it, I’m feeling pretty interested in trying to devil’s-advocate short timelines / red-team my most bullish-on-LLMs views. I think I grok the case for them pretty well; or, at least, there are some short-timelines cases that I think hang together well enough for me to be worried about them. Would you be up to it? Same format as with Abram.[1]
(Though I can see it ending up with me speculating about possible ways to advance capabilities in ways I wouldn’t want to make public. Hopefully at worst we’ll need to edit out a few paragraphs, though.)
Dialogues still seem possible to create through here.
I think I’m interested, though significantly less interested than in a defense from someone who actually holds the view. It sounds like you could maybe occupy some AGI-soon hypotheses well enough, so could be good. I’m super busy at the moment—maybe we can try in a month or two? Or if you like you could start the dialogue post thing, and then we can slow-cook it (like, a “correspondence game”)?
Yeah, that sounds good. I’ll figure out my opening statement and do that within probably a couple days.
What are your timelines?
See: https://www.lesswrong.com/posts/sTDfraZab47KiRMmT/views-on-when-agi-comes-and-on-strategy-to-reduce#Timelines And: https://www.lesswrong.com/posts/5tqFT3bcTekvico4d/do-confident-short-timelines-make-sense
Since then I’ve probably updated longerward based on a bit more hope in a slowdown, and shorterward from (logical and empirical) updates regarding neuro research (see https://www.lesswrong.com/posts/MnroTdcCCZoFSXEHy/a-case-that-whole-brain-emulation-research-is-net-harmful-by). Maybe they roughly cancel out, or rather, concentrate mass in 20-40 years a little bit more, relative to before? Overall it just feels like “mostly no one has any idea basically”.
(To be clear, it makes sense to be freaked out by gippity behavior, and we definitely should have a global ban on AI R&D yesterday.)
Do you have a link for this? It really matters whether they were like “it’s really quite plausible we will have AGI by now” (which aged IMO very well) or “we will definitely have AGI by now” (which would have aged relatively badly).
Unfortunately I do not. Jan had an internal presentation at DeepMind around 2019, I might ask him for the slides. My recollection is that in it he aggregated a few different methods, and ended up with 6-ish year timelines, but I could wrong (e.g. maybe the 6-ish year timelines were just one method rather than the all-things-considered prediction, or maybe they were actually 8-year timelines, or...).
I guess there’s another question which is: how did the short timelines vibes actually propagate into the community? There was Daniel’s What 2026 Looks Like; there was the Big Blob of Compute doc (though I don’t think that had dates?) Maybe it was in-person conversations? (From Daniel’s discussion with Ege and Ajeya, he makes this prediction: “Median Estimate for when 99% of currently fully remote jobs will be automatable: 4 years”. That would be November 2027, maybe with a bit of leeway for rounding.)
Partly I’m indexing on the mental recollection that people had shorter timelines than Shane, and Shane’s were always around 2027. But this is all messy and it would be really nice if we had more evidence about Dario’s views in particular, since I think they were pretty socially influential in illegible ways.
Two more datapoints about short-timelines predictions:
I don’t recall exactly how the Superalignment team’s 4-year deadline was chosen in 2023, but that was something to do with their timelines estimates (maybe it wasn’t the median/mean though).
The OpenAI Policy Research team under Miles had little timers on our desks counting down until (some estimate Miles generated of) when AGI would arrive. I don’t remember what the estimate was though, I believe less than 6 years (and we got them maybe in 2022?)
I went with Dario and Jan’s rather than other people’s because I feel like they were advocating for short timelines earlier than others were (and also because their timelines were IIRC shorter overall).
This is a great analogy, and is pretty good articulation of a model of the world that I have ~25% probability on.
But I do want to keep in mind, even under that model, the world is being dramatically transformed, industry by industry, and at a cultural level, as the AI capabilities that already exist become central to how every part of the world works.
What happened to software engineering this year will happen to many more domains, and to a bunch of slices of the world that we don’t currently consider as their own “domain”.
This is interesting because I would say that I have been surprised at the seeming lack of transformation of the software engineering industry this year? I think that in January/February 2026, there was a widespread expectation that we would see major layoffs this year, which don’t seem to have materialised (of course there have been layoffs, but they don’t seem to have been much greater than in previous years, and below the highs of 2022-23). I still don’t really have a good explanation to reconcile the reported transformation of the day-to-day work of SWEs with the fact that the industry doesn’t seem to be changing much, except for companies being surprisingly slow to cut headcount. They seem to have been slower to cut headcount in response to this than they were in responding to the Fed raising interest rates in 2022-23, for example.
The consumers are buying software services not lines of codes. Coding agents made lines of code cheap but the process of shipping a service have a lot of human-related bottlenecks so Amdahl’s law applies.[1]
The example of radiologists is often invoked and even received pushback on this website but the fact that workers moved to different tasks when some were recently automated can also be illustrated by travel agents and translators, as Noah Smith points out: https://www.noahpinion.blog/p/ai-keeps-stubbornly-refusing-to-take
I will stop short of claiming this should apply to AI development as well since I haven’t thought this through, inviting the readers to think about that.
This is why jobs may feel like “bullshit” to the people doing them, even as they command high wages in the market. (Footnote Smith’s, not mine—P.)
Importantly, this complexity of where the value comes from is an important part of why, in many jobs, hourly wages provide more value to the employer than more elaborate incentive designs.
I do think it’s worth considering the possibility, from a devil’s advocate perspective, that software engineering was the low-hanging fruit, and thus most other domains will take longer to automate.
My thought on it: culture and adoption lags capabilities, A LOT. By far most people don’t know there’s models more capable than free-tier Chat Jippity and have never engaged with something Astra/Fable class. To understand the frontier is one massive hop, and to figure out what to do with such capabilities is yet another massive hop. You can’t forget the people active in the AI communities are fractions of a percent, the genpop largely doesn’t comprehend the state of the art. It’s not a strike on anyone’s character, it just is. It’ll take time.
I always feel that there’s something off about this argument. Like, if a given tool provides massive benefits, it would stand to reason that it would be widely adopted quickly, because early adopters would start massively out-competing everyone else, thus pushing everyone else to figure out how to adopt those tools as well. No?
Slow diffusion makes sense if it’s bottlenecked by physical constraints, or by coordination across many actors, or if we’re talking about e. g. government institutions which are insulated from competitive pressures. But if we’re talking about market actors adopting tools that boost their individual productivity, I don’t think this mechanism makes sense.
This would happen if the benefits and the costs were to be concentrated.
When the benefits are diffuse, and the costs are concentrated, then adoption lags.
One such example in a regulated industry would be to allow certain AI models to give valid medical prescriptions—the costs are concentrated on doctors and the benefits are diffuse amongst patients.
Another example that doesn’t depend on regulations are escalators: When they were adopted by department stores in the 60′s, the benefits accrued to the consumers while the department store owners ate up the costs. None of the businesses gained an strategic advantage from it.
If all firms in the same industry face the same cost of implementing AI, then whomever discovers the most optimal way of implementing AI will end up paying all the costs of discovery through trial and error, and gain little to no strategic advantage The only advantage would be “first-mover advantage” which in a competitive industry isn’t really worth much.
Therefore the only industries where we would expect to see quick AI adoption are those that:
Are dominated by scale economies such that first-mover advantage matters.
Are not too regulated.
Have healthy profit margins and a good cashflow to finance the trial-and-error discovery process.
It’s hard to find such industries because (1) strong economies of scale is usually correlated with natural oligopolies and thereforre (2) regulation, such as telcos or airlines. It’s really only the software industry where we have (1) and (2) together because the initial capex is not so massive - you can’t start an airline from a garage—and economies of scale come mostly from network effects.
Indeed, it’s only really software where we are actually finding that the frontier AI models are disrupting the way that things are done.
Other industries are mostly passively following what software tells them to do, and looking for low costs:
Use AI for notetaking! - Unless you are in the legal sector and this creates new problems
Use AI to read documents faster! - Unless you are in academia and you didn’t even read the research papers to start with
Use AI to finish that deck faster! - Unless the bosses internalize this and just increase the number of meetings!
Those are not things that can massively increase productivity and allow you to outcompete the rest. Ergo, as @marquis_de_sod said, “culture and adoption lags capabilities, A LOT”
My personal experience is that people are ok with using free ai but it is seen as shameful to pay for it as that would directly finance the evil billionaires
Sure, but then why are people without such qualms, which I expect constitute a double-digit percentage of the population, not outcompeting them so massively that they’re forced to squash those qualms and adopt those tools anyway?
The answer is likely, in my opinion and experience, that the majority of the people who are early or heavy adopters of AI tend to use it to avoid effort instead of utilizing it as an effort multiplier.
LLMs are strangely bad at tasks in this category: scientific research, writing a novel, running a business, designing a video game, coming up with app ideas, writing funny jokes, high-level software architecture
All of these things involve high-level integrative reasoning that incorporates information from multiple domains at once.
For LLMs to get good at AI research and trigger a singularity, this integrative reasoning barrier would have to fall. Will it? I see two possibilities:
Integrative reasoning is fundamentally of a different nature than more domain-specific reasoning, and we need fundamentally new techniques to unlock it
Labs have been focusing on domain-specific reasoning like coding and math at the expense of integrative thinking, leaving LLM minds fragmented. Fixing this fragmentation mostly requires a change of focus; a change in training methodology which does not require any major breakthroughs.
If 1 is true, the singularity is far away. If 2 is true, then we’re in an integrative reasoning overhang, and the singularity is just around the corner.
To elaborate on the scientific research part:
AI are at the frontier of math research, solving Millennium Prize problems and other famous problems on a regular bases. AI is bordering on superhuman at math. Where are all the science breakthroughs?
I think current AI models are only superhuman at one half of math. They’re really good at chugging along, finding proofs using mostly existing techniques (especially if you consider that existing techniques include private data from chats with mathematicians). But that domain-specific proof-finding ability does not come with an accompanying integrative advancement of the field of mathematics.
Terrance Tao complains that AI math proofs are huge, complicated artefacts that are impossible to verify by humans. They do nothing to actually advance the field of mathematics conceptually.
Why? Because AI is not doing the integrative reasoning that would find new connections and develop new insights, the kind of work where new techniques and new discoveries are made along the way that push the field forward, like Isaac Newton inventing calculus in order to solve physics problems. The AI takes care of the busywork of finding technically-correct proofs, but to advance the field, the insight still needs to be supplied by mathematicians.
The same thing is true in science, but AI is even less helpful because in most domains the busywork cannot yet be carried out by machines, since it happens mostly in the real world. Either way, the insight is being supplied by humans, because AI mostly can’t generate meaningful insight, as it lacks the integrative reasoning needed to do so.
I agree that LLMs are not currently advancing research conceptually.
Why do you think this is a barrier and not just an area of slower progress?
Most of those and definitely science require general purpose reasoning integrated over a bunch of areas, requiring a lot of time. Memory is a know deficit of models relative to humans. Scaffolds and schemes to organize that reasoning are a separate project from expanding LLMs. Google AI co-scientist project focused on scaffolding and produced remarkable results from Gemini 2 pro (maybe even 1.5?).
So I expect LLMs to get better at complex open-ended projects if and when people develop scaffolding—OR as a byproduct of scaling. Fable is in my opinion markedly better at the generalized reasoning you need to do good science. I agree that this is a weak area, but I expect slower progress relative to coding or math with their extensive training data, not a brick wall.
I don’t think there’s some fundamental barrier. My initial comment gives two possibilities: we need fundamentally new techniques/architectures, or we just need to tweak existing architecture/training to improve integrative reasoning. I am not sure which is true. You seem to put forward a third option, which is that there is no integrative reasoning overhang, and this kind of reasoning just requires more intelligence than the models have right now.
Some reasons to think it could be a harder barrier for current architectures:
I have not seen meaningful progress in this area personally. My experience doesn’t match yours. For example, current models still ~never write funny jokes if you ask them to try. (Their score on my LaughBench benchmark is zero.) The Google AI co-scientist stuff seems similar to me to the recent AI math work, doing useful things in the field by searching through existing idea-space without generating any novel conceptual insights.
Probably a lot of significant research progresses by making more than one insight, ind the early insights help you get to the later insight. This happens because the early insights change your brain so you can think in terms of that insight. LLM model weights don’t update, so they can only think based on existing concepts, with any new insights being dealt with at a surface level and not really absorbed.
LLM minds seem fragmented, and I don’t know whether this is easy to fix. I mean this in the sense given by this paper and I think it’s part of why we see stuff like the part of the AI that talks to you being different from the part of the AI that does stuff and the general vibe that an LLM is 100,000 minds loosely stapled together, rather than one mind.
Where are all the training-time insights? AlphaZero learned to play chess at training time. All the data it absorbed led to insights at training time. LLMs manage to memorize a bunch of stuff, but where are the insights? If someone as smart as Astra read all the data ever, shouldn’t there be a whole bunch of novel connections just baked into the model? Instead we get simple regurgitation. Probably because when it’s learning about Y, it’s not possible or it’s very hard for it to form a connection with X, because its mind is too fragmented for X to be accessible when Y is being learned. So you get a bunch of separate, unintegrated knowledge.
But my intent wasn’t to imply there was definitely a hard barrier, only to communicate that this type of integrative reasoning is in fact lagging behind, and that we could be in an integrative reasoning overhang, and a change of training or architecture could change it quickly. Right now, I think these limitations are basically taken for granted as the model’s intelligence simply being too low. People aren’t thinking about domain-specific vs. integrative reasoning and how we are optimizing for the former at the expense of the later (with grave safety implications, I would guess).
I agree it’s mostly a scaffolding problem. Even if the LLMs do make conceptual progress, no one wants to read LLM outputs, so progress in science/research/epistemics is still diffusing at the rate of the humans communicating it. That rate is scarcely faster now than it was before, and in some ways it is slower as the institutions of science are inundated with 2x as much real research progress, combined with 10x as much slop to wade through, and no easy way to distinguish the two.
I suspect many people are having their LLM agents make the same discoveries over and over.
I don’t think the agents are incapable of discerning good from bad either, it’s just that they aren’t currently set up to do so.
I’m building a scaffold (minerval.ai) to organize scientific knowledge and have frontier LLMs assess the quality of the evidence for every claim. I’d be interested to hear what you think about my approach.
Another thing that all of these task share in common is that they aren’t easily verifiable. All of those tasks have very slow (or very weak) reward signals such that it is harder to RLVR on them. Scientific research is a bit of an anomaly in this list, as some of it can be verified computationally and doesn’t need a physical lab. However, game design, app ideas, novels and funny jokes, the reward signal is popularity. Because popularity is innately bottlenecked on humans (and even then, high-status humans have an outsize influence on the reward signal here), it is difficult to speed those up. I expect these fields to be the hardest to crack in developing AGI. In terms of high-level architecture, that is more feasible for RLVR loops.
I don’t think being hard to verify explains why AI isn’t good at these tasks. I think AI has gotten somewhat better at many hard-to-verify tasks (like deep research (summarization), instruction following and helpfulness, reducing sycophancy, medical answers, etc.). But it’s actually made almost no improvement at all when it comes to writing jokes, coming up with app ideas, etc.
I think the problem is either that they’re not focusing on making AI better at these things, or they don’t know how.
Agree Vladimir’s posts on this are illuminating. See also The Goodhart Singularity (though compare Data bottlenecks won’t prevent an intelligence explosion).
My view is still uncertain, but incorporates elements of the above. I expect R&D automation is ‘taste complete’, and that maintaining frontier taste requires ceteris paribus human-level sample efficiency. So not much acceleration prior to that. But speed/scale will provide some degree of qualitative adjustment (perhaps not much?). Compute likely remains a constraining factor.
And then actually ‘moving atoms’ (as you put it) requires a bunch of task-/domain-specific data. That means (perhaps slow, faffy?) integration of sensing, log ingestion, etc.
If sample efficiency remains elusive, we might even see an industrial ‘explosion’ before an intelligence explosion, since manufacturing at existing tech levels is relatively repeatable compared with frontier-pushing R&D.
There’s something there but I think its scary to reason from the the lack of normal economic impacts because they’re such a lagging indicator.
I think the current situation is:
Frontier AI labs are really getting to the point where the nature of their work is fundamentally changed and productivity really is up by a large fraction (cf recent Ant, OAI acceleration posts)
In most other places the impact is pretty modest
AI labs are willing to spend hundreds of thousands of dollars per member of staff on compute to boost productivity
AI labs have a combination of beliefs and competitive pressures (including deep throughout the company) that allow them to aggressively redesign workflows around new tools
A large amount of the training effort of labs goes into the kinds of work done by these companies, not least just because the feedback loop from the massive usage volume is so tight
Pretty much no other workplaces have this combination (maybe like Jane Street/Citadel? would be very curious to hear how much AI is accelerating them)
I think this means that it makes a lot of sense that we could see limited global economic transformation but AIs really are capable of speeding up the labs by a huge and increasing factor, and potentially going FOOM. By Christmas per Alex’s above post still seems aggressive but it seems more likely than after 2030 (conditioned on no slowdown).
I partly agree on lagging factor, but I disagree on the acceleration condition being unique to labs.
First, AI labs do not have a monopoly on using expensive coding agents to develop. OpenAI reported their research team uses ~$600/day/person, which is still in the range of an AI-forward SF company (I agree the majority of older software companies do not do this). Note that OpenAI’s spend might be slightly exaggerated by virtue of likely not caring about internal costs and not doing the most mild things to reduce it (e.g. smaller auto-compact windows, cheaper models for tasks that require less intelligence, etc.). Additionally, given the labs are locked into themselves, their capabilities might not be that far ahead of what the public can use. (Was OpenAI only having Astra in August that much SOTA over the public who was already using Fable 5 and Opus 5?).
In my domain, I’d ballpark engineering productivity up ~80%+ compared to early 2025, but it’s nuanced how that manifests itself. Onboarding (both as a new employee and into new functional areas) are very fast. But ultimately software engineering is under Jevon’s Paradox—our additional productivity translates to making better software (less bugs, more features), which all of our competitors also do as well.
Additionally, there’s a feeling of diminishing returns to higher model intelligence. My read is that what is really happening is Amdahl’s Law at play—the models aren’t getting better fast enough at the non-verifiable tasks so further intelligence gains (in the METR horizons sense) are less translating to productivity gains. Models translate clean greenfield specs very well into code, but at a larger (dozens of engineers) company, work is often brownfield: “I’m building a new feature X, how should this play with Y feature that I didn’t even know existed until I started implementing X”—today, that typically requires human judgement and as models accelerate coding, is taking more and more of my time percentage-wise.
It’s plausible to me that timelines are not short, but not very plausible that things continue to feel like they are accelerating in the way we’ve seen the last few years, shepard-tone-esque, for another 10 years. That’s because there are some exponentials that absolutely have to become sigmoids in the next 10 years if we don’t have a singularity in that time, and if the exponentials bend down, this will cause things to feel slower. The best example of this is AI company revenue.
I agree that AI company revenue won’t continue on its exponential.
However, I expect that AI companies will continue on some kind of power accumulation exponential for significantly longer than they continue on their revenue exponential. Money is one kind of power, but a relatively weak one (e.g. Anthropic could bring in as much revenue as the USG, but then the USG could just take half of it).
And so you can imagine Sheperd-tone-esque dynamics like “AI is on track to be most of the economy! Wait, it did, but that didn’t change much??”
As one analogy (though there are differences between corporations and individuals), Elon is well past the point where money is one of the main bottlenecks on how much power he has. It’s not even clear that more money is net helpful on the margin, given e.g. that becoming a trillionaire would make him an easier political target.
good point, this has moved me a bit
Weirdly enough, I think things will plausibly accelerate literally 8-13 years later, in the manner of Nesov’s prosaic RSI, but yeah if by 2030 we don’t get superexponential growth in AI capabilities, we won’t have superexponential growth until likely 2040-2050.
(To be clear, this affects our priorities a lot, and I want to flag that this is not an excuse to update a lot of things if this happens.)
How would they fail to create a singularity or something like @Daniel Kokotajlo’s five centuries in five years? Additionally, what is the case against the singularity emerging from idiots scaling the goddamn neuralese?
UPD: How exactly is “a pretty sharp divergence between the measured capabilities of models and their real-world impacts” affected by the loss of Navier-Stokes and EpochAI’s report of a major advance(!!!) in math?
UPD2: How do we rule out mirror life synthesis by the first ASI if it decides to commit genocide?
A lot of people in this community are programmers and have the experience of the coding models having real-world impacts on their daily lives that aren’t just about the models getting better at benchmarks.
You could be correct that takeover-capable AI will take longer than short timelines. You are definitely correct that there’s a bit of groupthink going on around short timelines in the prosaic alignment and developer crowd.
But you could also easily be wrong. And you are entirely missing the fact that there’s some groupthink going on around longer timelines in the agent foundations crowd.
Everyone is very tempted to sip a little hopium of some sort. The PA folks are hoping maybe alignment is so easy we can do it at a dead run. And the AF folks are hoping LLMs stall out just short of transformative so they have longer to work on alignment if it is, as they tend to believe, quite tricky.
I’ve written about the empirical evidence for hopium and groupthink in Motivated reasoning, confirmation bias, and AI risk theory. I hope you’ll read it, at least the intro/summary and section on group biases in alignment thinking.
So by all means lay plans for the eventuality of longer timelines and difficult alignment. But please don’t go projecting confidence when we need to also prepare for short timelines.
You accuse short timeliners of making decisions based on an emotional stance. I think if you can look at this objectively you will see that your projection of longer timelines is also exactly that. You have reasons. But there are equally good reasons to expect fairly short timelines. The choice of beliefs comes down to preference. There are different merits to the positions that matter, but they aren’t enough to confidently predict timelines or alignment difficulty.
The rational move is to say we don’t know, and deal with that unfortunate situation as best we can.
I’m working on a more concise post on this point; for now the motivated reasoning post is my attempt to explain current unfortunate epistemic situation.
I agree with claims that we don’t have AGI now, and I’d put reasonable probability that we don’t get a superintelligence in 8 years (maybe around 70%?)
And yes, some short timelines people have to update now (thinking of you @Nathan Helm-Burger) But not all short timelines people have to update, like Shane Legg, for example.
I feel like this is probably the largest disagreement by far. I’d sympathize for 2025-era models for coding, but I now think we have much more hard evidence that AI models are useful beyond benchmarks, and I think the best example of this is the way OpenAI and Anthropic’s ARR accelerated in 2026 because coding models finally became good enough that you could use it as a subsitute rather than a complement, unless you have extremely good coding skills now.
There are still areas where humans can productively contribute, but more and more of this is non-coding tasks.
I think the main way in which short timeline predictions were wrong is that people underestimated how jagged the frontier was for AIs, because they incorrectly used human baselines, and as it turns out models can be much more competent at some things without being good at other things, which creates pretty severe bottlenecks to fully automating something, and it’s only with full automation of something where things tend to explode.
And conditional on the intelligence explosion/superintelligence in 8 years claim being wrong, I expect to see a similar pattern.
On this:
A lot of the resson for this is basically due to hedging, and contra you I think this is entirely rational and expected under worldviews where AI safety is the main problem, compared to the other problems of AI, and not purely emotional bandwagoning.
Habryka puts it well in that worlds with an intelligence explosion soon are worlds where humans need to react now, and thus is most non-puntable/neglected, which all else equal is a bullish sign on impact (and I’d also say that it’s tractable to focus on the scenario.)
Now, I don’t think this is super-useful as a scenario to focus on, and I agree we shouldn’t unthinkingly assume RSI soon, but that’s because I expect better futures work to outweigh AI safety work in a very large percentage of scenarios.
I will also say that despite disagreeing with you on the details, I also agree that posts like @Alexander Gietelink Oldenziel’s Superintelligence by Christmas is pure clickbait, and the fact that it was voted to 61 karma before my strong downvote is a bad sign of the epistemics of this forum.
As an even more visceral (IMO) ‘Shepard tone’ analogue, this time in tempo, see the rather incredible Troy track from Goransson’s Odyssey soundtrack.
This may be a hot take, but Astra meets my bar for AGI. I was surprised to find this to be the case, but so far I haven’t found anything in which it doesn’t perform better than my median friend. Of course it’s quite possible I’ve simply failed to try some large category of tasks in which it sucks, but take that for what you will.
As much as I disagree with his long-term conclusions, Arvind Narayanan’s work (AI as normal technology, etc.) should at least make you suspicious about whether you can use “the world hasn’t changed much from 1 year ago” as strong evidence for “we have not developed AGI in the past year”.
random but I played the Shepard tones on Wikipedia and at least with my headphones they didn’t seem like they were getting loud at all except maybe very briefly. Certainly not monotonically getting louder or on net.
(I think I’m the usual or above average degree of susceptible to optical illusions so it was surprising that the audio ones had no effect on me).
The idea is that they go up in pitch, not volume.
I found this generator more convincing than the example on the current wiki page
I agree this is worth taking seriously, and I’m glad you have also converged on the interestingness of things like Vlad Nesov’s ideas! I’ve suspected for some time that Dario’s “country of geniuses in a datacenter” concept is specifically intended to allow for hedging in the case of worlds where LLMs are superhuman at a wide range of skills without triggering a singularity; the country of geniuses by some definition will arrive, and yet the world could look much the same.
For me, large swarms of AI coordinating to do secret hacking and quickly solve famously difficult math problems, in addition to hearing what lab employees saying about their own reliance on recent models, is enough to make me high probability that we have entered early RSI and things “take off” barring some effort to prevent (which seems to be forming). But I but am a “casual” outside observer.
>Around 8 years ago Dario Amodei, Jan Leike and a few others predicted that we’d probably have AGI by now.
Do you have a source for this? At least for Dario, every interview and blogpost I can find says he expects AGI sometime in 27⁄28.
My very basic and informal argument for why I think it’s likely that RSI would lead to commensurate jumps in capability in other domains is that an agent that is a vastly superhuman coder, researcher, economic agent etc. would be able to use currently unknown or unexploited modalities to gain large amounts of experiential data in those other domains and we would expect it to very quickly converge on a highly sample efficient strategy and a highly efficient sampling strategy in those domains.