I also talked about how “mundane” misalignment and underelicitation could doom us. I say more about this here.
(I think there are other important threat models, like AIs ending up with (shared) long-run preferences and deciding to fake alignment based on these preferences. For reference, see here and here.)
Some of people at Redwood wrote up some notes to help me prep:
My median for full automation of AI R&D is around late 2030/early 2031.[1] But my “modal”/best guess prediction for this milestone would be significantly earlier (mid 2029).
Here is a summary of my best guess prediction for what happens over the next few years:
EOY 2026:
~1.5x as much frontier AI progress in 2026 as in 2025 (mostly from eating up certain overhangs, but some from AI R&D acceleration).
AIs accelerate AI R&D labor at Anthropic by ~2.5x (as in, as useful as making all researchers/engineers think/work 2.5x faster).
EOY 2027:
Engineering at AI companies is pretty close to fully automated and AIs are making serious inroads into automating research. AI R&D labor acceleration: ~8.5x.
Some people claim AI R&D is fully automated in 2027. They aren’t right, but the situation is already quite crazy: AI companies feel insanely automated with humans often very out of the loop and the speedup is considerable.
~1.5x as much frontier AI progress as in 2025 (mostly from AI R&D acceleration, some from overhangs).
2028:
Automated coder (AC) around April. (AIs that can basically fully automate research engineering / SWE.)
Rough parity with human AI R&D researchers is reached late 2028, though humans still add significant value for a while (takes, taste, pointing out blind spots/errors).
In the second half of the year, AI progress runs ~1.6x the 2025 rate: 6 months of calendar time yields ~0.8 years of AI progress.
2029:
Superhuman AI researcher (SAR) early this year, a bit less than a year after AC.
Progress is picking up with ~1.3 years of AI progress in the first half of the year (2.6x rate).
By EOY, significantly past top-expert-dominating AI (TEDAI), with ~2.5 years of AI progress in the second half of the year (5x rate). AIs are now very superhuman in many domains (though this varies).
2030 (??):
Mid: AIs are somewhere between TEDAI and wildly superhuman AIs (ASI). Crazy shit. Compute is maybe doubling every ~4 months (downstream of robots).
EOY: Singularity™. We’ve had a bunch of economic doublings. Compute is doubling every ~2 months (???).
2031 (??????):
Mid: doubling time is more like ~2 weeks. Truly insane new technology is coming online.
Notes:
This assumes limited government intervention on the overall rate of AI progress and no substantial slowdown (voluntary or otherwise).
It also ignores misalignment: as discussed in the episode, I think misaligned AI takeover is quite plausible along the way (which would change the trajectory).
Milestones (AC, SAR, TEDAI) are roughly as defined in the AI Futures Model.
By “full automation of AI R&D”, I mean AIs such that firing all humans working on AI R&D (other than setting overall top level objectives) would slow down AI progress by less than 10%.
Obviously, all of this is extremely uncertain (increasingly so later in the scenario). This is my best guess prediction (a modal trajectory), not a confident prediction. My median for each milestone is later, but this is more like my central prediction for what I expect to overall happen.
Human (and animal) brains, by their existence, prove that much more efficient ways of getting high intelligence from limited data exist. Naturally we don’t know exactly what’s the “secret sauce” yet, but I find Dwarkesh’s stance that “Our current best models haven’t made much progress, therefore we can never get RSI” a bit weird.
I could also envision that what makes humans so good at learning is a kind of weird, evolved hack that isn’t easy to formalize, understand, or reason about, even though it could be replicated relatively easily given a few more pointers. In that case, too, “figuring it out” would yield large gains in short time as the algorithmic overhang falls, and we don’t know what the slope would then be.
Unfortunately, I don’t think that what makes human learning so efficient is very mysterious, difficult to understand, or difficult to guess. I say that after half a career of studying exactly that question, among other aspects of computational brain function.
I am currently finding this pretty concerning in relation to Ryan’s already fast timelines, which roughly match my own best guesses but are better thought out.
So I’m not going to repeat the hypothesis here, but I think we should be prepared for the AI industry to basically figure it out. As you say, that would accelerate progress from a new direction.
Thanks! I’ll have to take your word for it (since it’s probably unwise to ask for what exactly would make machine learning more efficient), but it does sound concerning, in favor of faster takeoff speeds.
Fantastic. I’m listening to it now, because this feels important. Dwarkesh is a huge voice and his skepticism about RSI and fast timelines may make a difference for attitudes in general, and therefore caution.
At 14:29 he says he gives his main crux: skepticism that intelligence is that intelligence is that important for AI progress. He’s arguing that compute is dominant over researcher intelligence.
This seems pretty strange to me. Porque no los dos? Which is pretty much your answer. Obviously compute is super important. And it seems equally important that AI research depends on intelligence. The idea that there are no more good ideas and it’s all just scaling from here seems so wildly counterintuitive to me that I have trouble understanding the mindset.
This is probably at root based on my perspective on human brain function. Scale is important. To a large extent, human intelligence is just a result of scaling up primate intelligence. But even primate intelligence involves a lot more moving parts than the general learning mechanisms of LLMs. At the very least there are a bunch of motivational tricks that direct our continual learning very efficiently. And it’s pretty inarguable that episodic memory and continual learning play a big role in human intelligence, on top of core learning mechanisms that are (much more arguably) fairly similar to the combination of predictive and RL learning used to train LLMs.
Anyway that’s just one intuition pump for the intuition that we haven’t likely discovered nearly all of the tricks that will make scaffolded LLMs effectively smarter/more competent.
Actually, it seems contradictory if I’m right that Dwarkesh believes ASI will need continual learning (or mountains of new RLVR data). Why doesn’t he think intelligence won’t help create better CL mechanisms? I guess he’s assuming the RLVR method beats the CL method. I assume the opposite. I think well-selected fine-tuning will work just fine for CL and it’s the likely breakthrough.
The idea that there are no more good ideas and it’s all just scaling from here
Fwiw, this is not how I understand his take. I think he’s saying AI progress will be bottlenecked by compute (and human expert data), which I interpret to mean that the elasticity of substitution between compute and intelligence isn’t high enough. In fact, I think this his view includes compute to run experiments to do AI research.
Right but it seems that increased intelligence will reduce that bottleneck by making the limited amount of compute for experiments go farther. How much farther is an open question.
I briefly discussed with Dwarkesh why I’m skeptical AI progress is heavily driven by scaling up spending on human experts labeling/making data. My main argument is that spending on researchers and experiment compute seem much higher. But I didn’t say very much in the podcast.
More precisely, my view is that if spending on having human experts label/make individual data points were fixed at ~$100 million / year (per company), then AI progress would be <25% slower.
We didn’t have time to get into everything in this podcast (and some content about this recorded at an earlier point was cut), so I’ll spell out my view in a bit more detail here:
Spending directly on data (rather than on R&D about data) isn’t growing that fast and isn’t that high (relative to spending on researchers and especially compared to spending on experiment compute).
It’s important to make a distinction between spending directly on making data and science about better processes for making data. E.g., better data mixes like FineWeb count as R&D (and the person doing the R&D needs almost no understanding of individual sequences).
AI automation seems differentially good at accelerating data generation, such that I think improvements in RL environments have mostly been driven by improved AI (and R&D into making better environments) rather than spending on humans making/labeling data, and I expect this to continue. Like data stuff seems particularly amenable to acceleration from weaker AIs.
Structurally, most of what human data labeling does (though not all!) depends on having generally decent judgment rather than on having more expertise than the AIs being trained.
Transfer without domain-specific labels looks decent in practice. E.g., it doesn’t seem like Anthropic is hiring a ton of mathematicians and cyber experts to do data labeling, and the AIs are still good and getting better at these domains. Maybe this depends on having labels in some domain, but so long as AIs can label in domains that transfer well enough and/or can make RL envs that don’t require much labeling, that would be fine.
I recently recorded a podcast with Dwarkesh about the potential for (very) fast AI progress and how misaligned AI takeover might happen.
Our conversation focused a lot on threat models from “reward-seeking” AIs. If you’re interested in reading more about this, Alex (who works with me at Redwood) has written in a lot more detail about this threat model here
I also talked about how “mundane” misalignment and underelicitation could doom us. I say more about this here.
(I think there are other important threat models, like AIs ending up with (shared) long-run preferences and deciding to fake alignment based on these preferences. For reference, see here and here.)
Some of people at Redwood wrote up some notes to help me prep:
Lukas Finnveden on why human checks-and-balances would fail for AIs
Lukas on disanalogies between (misaligned) AIs and human labor
Alex Mallen on whether reward-seeking is too “unambitious” to cause takeover
Alex on what sloppy AIs might look like
My median for full automation of AI R&D is around late 2030/early 2031. [1] But my “modal”/best guess prediction for this milestone would be significantly earlier (mid 2029).
Here is a summary of my best guess prediction for what happens over the next few years:
EOY 2026:
~1.5x as much frontier AI progress in 2026 as in 2025 (mostly from eating up certain overhangs, but some from AI R&D acceleration).
AIs accelerate AI R&D labor at Anthropic by ~2.5x (as in, as useful as making all researchers/engineers think/work 2.5x faster).
EOY 2027:
Engineering at AI companies is pretty close to fully automated and AIs are making serious inroads into automating research. AI R&D labor acceleration: ~8.5x.
Some people claim AI R&D is fully automated in 2027. They aren’t right, but the situation is already quite crazy: AI companies feel insanely automated with humans often very out of the loop and the speedup is considerable.
~1.5x as much frontier AI progress as in 2025 (mostly from AI R&D acceleration, some from overhangs).
2028:
Automated coder (AC) around April. (AIs that can basically fully automate research engineering / SWE.)
Rough parity with human AI R&D researchers is reached late 2028, though humans still add significant value for a while (takes, taste, pointing out blind spots/errors).
In the second half of the year, AI progress runs ~1.6x the 2025 rate: 6 months of calendar time yields ~0.8 years of AI progress.
2029:
Superhuman AI researcher (SAR) early this year, a bit less than a year after AC.
Progress is picking up with ~1.3 years of AI progress in the first half of the year (2.6x rate).
By EOY, significantly past top-expert-dominating AI (TEDAI), with ~2.5 years of AI progress in the second half of the year (5x rate). AIs are now very superhuman in many domains (though this varies).
2030 (??):
Mid: AIs are somewhere between TEDAI and wildly superhuman AIs (ASI). Crazy shit. Compute is maybe doubling every ~4 months (downstream of robots).
EOY: Singularity™. We’ve had a bunch of economic doublings. Compute is doubling every ~2 months (???).
2031 (??????):
Mid: doubling time is more like ~2 weeks. Truly insane new technology is coming online.
Notes:
This assumes limited government intervention on the overall rate of AI progress and no substantial slowdown (voluntary or otherwise).
It also ignores misalignment: as discussed in the episode, I think misaligned AI takeover is quite plausible along the way (which would change the trajectory).
Milestones (AC, SAR, TEDAI) are roughly as defined in the AI Futures Model.
By “full automation of AI R&D”, I mean AIs such that firing all humans working on AI R&D (other than setting overall top level objectives) would slow down AI progress by less than 10%.
Obviously, all of this is extremely uncertain (increasingly so later in the scenario). This is my best guess prediction (a modal trajectory), not a confident prediction. My median for each milestone is later, but this is more like my central prediction for what I expect to overall happen.
See also here for some more predictions.
Pretty cool, just listened to it.
Human (and animal) brains, by their existence, prove that much more efficient ways of getting high intelligence from limited data exist. Naturally we don’t know exactly what’s the “secret sauce” yet, but I find Dwarkesh’s stance that “Our current best models haven’t made much progress, therefore we can never get RSI” a bit weird.
I could also envision that what makes humans so good at learning is a kind of weird, evolved hack that isn’t easy to formalize, understand, or reason about, even though it could be replicated relatively easily given a few more pointers. In that case, too, “figuring it out” would yield large gains in short time as the algorithmic overhang falls, and we don’t know what the slope would then be.
Unfortunately, I don’t think that what makes human learning so efficient is very mysterious, difficult to understand, or difficult to guess. I say that after half a career of studying exactly that question, among other aspects of computational brain function. I am currently finding this pretty concerning in relation to Ryan’s already fast timelines, which roughly match my own best guesses but are better thought out.
So I’m not going to repeat the hypothesis here, but I think we should be prepared for the AI industry to basically figure it out. As you say, that would accelerate progress from a new direction.
Thanks! I’ll have to take your word for it (since it’s probably unwise to ask for what exactly would make machine learning more efficient), but it does sound concerning, in favor of faster takeoff speeds.
Fantastic. I’m listening to it now, because this feels important. Dwarkesh is a huge voice and his skepticism about RSI and fast timelines may make a difference for attitudes in general, and therefore caution.
At 14:29 he says he gives his main crux: skepticism that intelligence is that intelligence is that important for AI progress. He’s arguing that compute is dominant over researcher intelligence.
This seems pretty strange to me. Porque no los dos? Which is pretty much your answer. Obviously compute is super important. And it seems equally important that AI research depends on intelligence. The idea that there are no more good ideas and it’s all just scaling from here seems so wildly counterintuitive to me that I have trouble understanding the mindset.
This is probably at root based on my perspective on human brain function. Scale is important. To a large extent, human intelligence is just a result of scaling up primate intelligence. But even primate intelligence involves a lot more moving parts than the general learning mechanisms of LLMs. At the very least there are a bunch of motivational tricks that direct our continual learning very efficiently. And it’s pretty inarguable that episodic memory and continual learning play a big role in human intelligence, on top of core learning mechanisms that are (much more arguably) fairly similar to the combination of predictive and RL learning used to train LLMs.
Anyway that’s just one intuition pump for the intuition that we haven’t likely discovered nearly all of the tricks that will make scaffolded LLMs effectively smarter/more competent.
Actually, it seems contradictory if I’m right that Dwarkesh believes ASI will need continual learning (or mountains of new RLVR data). Why doesn’t he think intelligence won’t help create better CL mechanisms? I guess he’s assuming the RLVR method beats the CL method. I assume the opposite. I think well-selected fine-tuning will work just fine for CL and it’s the likely breakthrough.
Fwiw, this is not how I understand his take. I think he’s saying AI progress will be bottlenecked by compute (and human expert data), which I interpret to mean that the elasticity of substitution between compute and intelligence isn’t high enough. In fact, I think this his view includes compute to run experiments to do AI research.
Right but it seems that increased intelligence will reduce that bottleneck by making the limited amount of compute for experiments go farther. How much farther is an open question.
I agree with that!
I briefly discussed with Dwarkesh why I’m skeptical AI progress is heavily driven by scaling up spending on human experts labeling/making data. My main argument is that spending on researchers and experiment compute seem much higher. But I didn’t say very much in the podcast.
More precisely, my view is that if spending on having human experts label/make individual data points were fixed at ~$100 million / year (per company), then AI progress would be <25% slower.
We didn’t have time to get into everything in this podcast (and some content about this recorded at an earlier point was cut), so I’ll spell out my view in a bit more detail here:
Spending directly on data (rather than on R&D about data)
isn’t growing that fast andisn’t that high (relative to spending on researchers and especially compared to spending on experiment compute).It’s important to make a distinction between spending directly on making data and science about better processes for making data. E.g., better data mixes like FineWeb count as R&D (and the person doing the R&D needs almost no understanding of individual sequences).
AI automation seems differentially good at accelerating data generation, such that I think improvements in RL environments have mostly been driven by improved AI (and R&D into making better environments) rather than spending on humans making/labeling data, and I expect this to continue. Like data stuff seems particularly amenable to acceleration from weaker AIs.
Structurally, most of what human data labeling does (though not all!) depends on having generally decent judgment rather than on having more expertise than the AIs being trained.
Transfer without domain-specific labels looks decent in practice. E.g., it doesn’t seem like Anthropic is hiring a ton of mathematicians and cyber experts to do data labeling, and the AIs are still good and getting better at these domains. Maybe this depends on having labels in some domain, but so long as AIs can label in domains that transfer well enough and/or can make RL envs that don’t require much labeling, that would be fine.