A take I’ve come to believe around generalization/sample efficiency is that if we do start to see AIs improving a lot at generalization/sample efficiency, we will see it first in robotics/robotics evals, and we have early signs of robotics/robotics evals starting to improve without robotics-specific data, which is bullish for generalization.
There are a couple of reasons for this:
A lot of the proposed use cases for robotics, like industry robots that handle varied environments or home robots require more generalization, especially from other data, unlike what has appeared to be the situation for tests and coding/easily grindable mathematics (and potentially white collar work.) Home robotics are an especially hard test case, because homes are spaces of infinite variation, long tails and constant change. As such, the scope of a home robot cannot be narrow. To create real value, it must perform reliably across open-ended variation with no special setup, and that is a practical test of embodied general intelligence.)
I’d say the major reason for this comes down to the fact that in areas where we don’t need much generalization, we already have good enough robots for various tasks, and there’s no overhang for AIs to take and become suddenly good without generalizing to other skills (unlike tests where IQ tests do measure human performance at lots of things somewhat well, or coding/easily grindable mathematics.)
One other thing that helps is that it’s much easier to restrain LLM training data around the physical world, partially from the world itself being only partial information, but also the fact that it’s much easier to restrain poorly generalizing AIs from just cheating it with a lot of training data, because it’s easier to avoid training data leakage/easier to bound how much AIs need to generalize. It’s also easier to give them hard problems, unlike benchmarks on the internet, because the model might just not be trained on robotics data (and this is feasible), or if it is done, it’s much easier to only give them training data on specific tasks and have a genuinely secure holdout test set, unlike pre-training where you have to give AIs more or less the entire internet, so you can’t train them on specific tasks and have a holdout test set, because the training data might leak information on how to solve the test set, and you have no way of controlling this in the digital world, but you can control data much more effectively in the physical world.
Given that generalization is a current bottleneck for robotics, this means that foundation models have to increase their generality of intelligence in order to improve on the tasks.
Now for the part about robotics improving, the current best sources are ACT-2 preview, where pre-training data (that isn’t robotics specific) improved out of domain performance dramatically (from ~15% to 99.1%) on folding laundry, and they have strict standards to ensure that generalization is the explanation for their improvements. The other one is Opus 5, which had also improved a lot on controlling robots at different tasks, namely stacking bowls and holding a fork (which is difficult for robots to grasp due to being thin and flat), purely through the power of it’s mindwith zero robotics data.
If I’m right that robotics evals are the main place to look for when assessing LLM generalization/sample efficiency abilities, this looks a lot to me like pre-training is genuinely giving LLMs generalization, and Vanessa Kosoy asked a question about about whether LLMs have a positive amount but are worse than humans at the ability to effectively and continuously acquire deep knowledge and exploit this knowledge to construct and execute goal-directed plans over long lifetimes and consequence horizons, or have zero of the ability and are only doing it because pre-training is simply just getting human algorithms from the internet and once the internet is exhausted, AIs will stop getting better, and now the answer for 2026 LLMs is yes, they have a positive but subhuman amount the property of effectively and continuously acquire deep knowledge and exploit this knowledge to construct and execute goal-directed plans over long lifetimes and consequence horizons/generalization.
Some other implications of “robotics/robotics evals are where we will see generalization ability/sample efficiency, if it appears in AIs”:
It means the LLM paradigm is not completely off the mark, even if it requires major changes like continual learning/neuralese recurrence and memory, and there’s a real sense in which current LLMs generalize (Contra Tsvi or Thane Ruthenis’s hypotheses on LLMs)
It means that 10-20 year timelines are quite likely, even if the very fast scale up of the 2020s fails to produce AGI for whatever reason, and this means we need to be prepared for AGI/ASI to come in our lifetimes.
Relatedly, this also means there’s a plausible path for LLM researchers to work on to improve generalization.
Depending on how much RL was used for Opus 5, this might be a real “GPT-3 moment for RL”, where it generalized from lots of RL environments to new problems, demonstrating that RL can actually generalize beyond the trained tasks.
Finally, this might be able to replace the METR plots, assuming a time horizon equivalent can be made for physical tasks/robotics, and with much less worry about external validity/measurement problems.
Are there any takes from LLM skeptics like @TsviBT or @Thane Ruthenis about the idea of using robotics evals/robotics progress to predict generalization rather than using current benchmarks?
Thanks for A2A. I think it’s going to be hard to communicate about this because there are fundamental conceptual boundaries. It could help for you to write out the basic structure of the argument in simple concise syllogism format; and then expand on background enthymemes involved in the concept used for induction. This is to help address objections I might make about “wait does this concept actually support this inference / induction”? (Cf. https://www.lesswrong.com/posts/i7JSL5awGFcSRhyGF/shortform-2?commentId=fXtNjxLevGaiJHSyj )
In particular, in this case you’re using some concepts here:
they have a positive but subhuman amount the property of effectively and continuously acquire deep knowledge and exploit this knowledge to construct and execute goal-directed plans over long lifetimes and consequence horizons/generalization.
And you’re doing an induction about it. And I’m just going to object that the way you’re using the concept doesn’t support the inference you’re trying to make; or at least, AFAICT you haven’t explained the concept and how you’re making the inference well enough to justify the inference.
The particular observational evidence you adduce seems interesting, but also I don’t have enough context to know what it means. (And I’m not really trying to evaluate it, because busy and because the disagreement seems to stem from the above, not from some specific observation. There are observations that would invalidate substantive chunks of my model (cf. https://www.lesswrong.com/posts/sTDfraZab47KiRMmT/views-on-when-agi-comes-and-on-strategy-to-reduce#comments), but generally when people bring in some observation the disagreement is how they are interpreting it, not the observation itself.)
The use-cases of robotics beyond current robots, especially home robotics require handling lots of environments and have varied challenges, so it stress-tests AI generalization.
Unlike other domains, there’s no unexploited algebraicness/overhang that would allow AIs to be useful for these cases without generalizing, because if there were such a thing, very limited current industrial robots would already have been used. They haven’t, so there’s no room for algebraicness/non-generalization to be an issue.
It’s easier to control/exclude training data when the domain is physical, rather than digital due to a multitude of reasons.
For these reasons, if current AIs improve on robotic tasks, it is a much stronger signal that they’re generalizing, and therefore is a signal that AGI could come in 10-20 years, and depending on how fast they improve, this might shift to 2-4 years.
In particular, in this case you’re using some concepts here:
they have a positive but subhuman amount the property of effectively and continuously acquire deep knowledge and exploit this knowledge to construct and execute goal-directed plans over long lifetimes and consequence horizons/generalization.
And you’re doing an induction about it. And I’m just going to object that the way you’re using the concept doesn’t support the inference you’re trying to make; or at least, AFAICT you haven’t explained the concept and how you’re making the inference well enough to justify the inference.
So, the way I see Vanessa’s concept here is that it’s pointing to the concept of time horizons as defined by METR, which got introduced to the world by METR in 2025 called METR: Measuring AI Ability to Complete Long Tasks, and there’s a lot to be said here, including it’s problems, but the single most important thing to keep in mind here is that this is a continuous metric, so it’s meaningful to talk about how much the AI can acquire long-term knowledge to exploit it.
Now by itself, I would have viewed it as weak evidence, but combined with other evidence showing that exponential trends do exist in other domains, the fact that there are good general reasons to assume exponential growth in AI capabilities, plus the new robotics data shift me towards a view where the current LLM algebraicness/generalization issues are real, but probably fixable, and do not require throwing out the entire paradigm, and thus the induction/inference I’m making is valid here, or to put it another way the observations I have seen are more likely in worlds where LLMs can continuously acquire deep knowledge to exploit it than in worlds where they aren’t.
*Reads your model on progress to AGI, taking the definition of AGI in the post as given/assumed*
I very strongly disagree with this model of progress:
The first way humanity makes AGI is by combining some set of significant ideas about intelligence. Significant ideas are things like (the ideas of) gradient descent, recombination, probability distributions, universal computation, search, world-optimization. Significant ideas are to a significant extent bottlenecked on great natural philosophers doing great natural philosophy about intelligence, with sequential bottlenecks between many insights.
And instead believe in this model:
Algorithmic progress/good ideas/research taste do genuinely matter, but most people are prone to overestimate how much they matter without compute or data, and the main way they currently matter is by being more-compute/data efficient at larger compute scales, and importantly the current paradigm isn’t irrelevant to building the next one, because you can reuse the compute and data from the old paradigm, and the new paradigms will not reduce the demand for compute.
This means progress to AGI isn’t as serial as you state (though I do expect serialness in a different way through RL, but you can still reuse compute/training data you already bought once you have a better paradigm).
More generally, what has made AI go from being ~completely useless and generalization was fake to current AIs where generalization exists, but is worse than humans and can do real economic work is basically just compute and data (and note that in a RL/synthetic data heavy world, data bottlenecks are folded into compute bottlenecks, so compute is the main input)
Or to put it another way, you seem to treat AGI as though there’s a discrete set of challenges to get there, one after the other to be solved, while I instead treat AGI as a set of challenges which all have continuous metrics, and in particular there’s ways to reuse previous progress, because the load-bearing elements of AGI depend less on a particular paradigm/theory of intelligence and they depend more so on compute.
Put another way, you expect sharp/discontinuous progress towards AGI and a singularity, whereas I think we will smoothly go towards AGI and a singularity (note this isn’t talking about how fast or slow the takeoff is.)
This also solves disagreements with the inputs section of the post.
To simplify the disagreement in one sentence, I notice that the central crux of whether or not the load-bearing AI progress is in compute/dat or conceptual breakthroughs on theories of intelligence keeps explaining a lot (but not all) of our disagreements, and disagreements expectations of how smooth vs discrete progress is in AI explain why I expect very different things around AI progress.
The use-cases of robotics beyond current robots, especially home robotics require handling lots of environments and have varied challenges, so it stress-tests AI generalization.
“it stress-tests AI generalization” But does it though? What does it stress test? That seems to assert a generalization from AI generalization about XYZ robotics tasks to ABC intelligence explosion tasks (or something along these lines?). What’s that generalization / what justifies it?
Unlike other domains, there’s no unexploited algebraicness/overhang that would allow AIs to be useful for these cases without generalizing, because if there were such a thing, very limited current industrial robots would already have been used. They haven’t, so there’s no room for algebraicness/non-generalization to be an issue.
I don’t understand this. Industrial robots of course are used very widely? Or I guess you typoed, and you’re trying to say, LLMs or other AIs would have been used to compute very non-obvious complex actions that take advantage of algebraicness in the narrow domain? Or you’re saying something about humanoid robots, or about softer / less rote tasks (laundry folding rather than installing a chair in a car)?
For these reasons, if current AIs improve on robotic tasks, it is a much stronger signal that they’re generalizing, and therefore is a signal that AGI could come in 10-20 years, and depending on how fast they improve, this might shift to 2-4 years.
How do you get from the qualitative and relative statement about robotics being “a much stronger signal” to numbers and years?
Or to put it another way, you seem to treat AGI as though there’s a discrete set of challenges to get there, one after the other to be solved, while I instead treat AGI as a set of challenges which all have continuous metrics, and in particular there’s ways to reuse previous progress, because the load-bearing elements of AGI depend less on a particular paradigm/theory of intelligence and they depend more so on compute.
My guess is that these things are not actually cruxes? Not sure. Instead I’d guess that the cruxes are more simply
the quantitative amount of insight (regardless of continuousness) remaining to get to AGI
the degree to which those insights are blocked on difficult thinking that current AI doesn’t accelerate by much
I don’t understand this. Industrial robots of course are used very widely? Or I guess you typoed, and you’re trying to say, LLMs or other AIs would have been used to compute very non-obvious complex actions that take advantage of algebraicness in the narrow domain? Or you’re saying something about humanoid robots, or about softer / less rote tasks (laundry folding rather than installing a chair in a car)?
I agree industrial robots that don’t generalize much are used quite widely in narrow environments, the point is that for the set of tasks where AIs don’t generalize much/don’t improve generalization, they can’t improve industrial robotics/cause the robotics industry to be more useful and profitable, because the robotics industry already squeezed out ~all the gains from AIs that don’t generalize well to new environments.
Yes, this is about less rote tasks like laundry folding across arbitrarily oriented clothing and being crumpled, plus being in arbitrary locations in the home, while a human constantly disturbs them and provides OOD pressure (indeed one of my examples was about this exact thing being done well by an AI that controlled a robot using only pre-training without robotics data), and due to the way home robotics has infinite variation, long tails and constant change, this means it’s way easier to measure generalization cleanly/in a controlled manner.
“it stress-tests AI generalization” But does it though? What does it stress test? That seems to assert a generalization from AI generalization about XYZ robotics tasks to ABC intelligence explosion tasks (or something along these lines?). What’s that generalization / what justifies it?
The stress test of generalization is more indirect than this. My model is that as AIs improve at more and more robotics tasks, it simultaneously has to be able to do well in more and more messy and OOD environments (otherwise current robotics would have already improved more than they actually are), intelligence explosion tasks, assuming an intelligence explosion/software only singularity is possible, involves a great many very messy and OOD tasks, thus robotic evals are real evidence that AIs are increasingly able to start intelligence explosion tasks.
How do you get from the qualitative and relative statement about robotics being “a much stronger signal” to numbers and years?
Fair point, I don’t think I justified this very well, at least as a claim about the specific units, but a vague argument here is that given the qualitative statement above, it implies that the current paradigm is not so far off the mark that it’s absolutely useless as progress for AGI design, and this means there’s a real path that doesn’t require heroic genius from today to AGI.
To address your cruxes:
the quantitative amount of insight (regardless of continuousness) remaining to get to AGI
My views on this is that I think at most 2 insights are necessary in practice, and if the human population kept perpetually growing according to 1880s-era birth-rates, we’d need 0 insights, as the super-exponential growth of the economy would have allowed us to get to AGI eventually.
the degree to which those insights are blocked on difficult thinking that current AI doesn’t accelerate by much
I’d say they’re non-trivially serial (especially if the solutions use a lot of RL, as RL is much, much more difficult to parallelize outside of current environments), but a fairly key part of my mental model is that conditional on current LLMs becoming blocked, there will be massive incentives to point researchers into promising paradigms, meaning that more resources will be pursued on the blockers, compensating for their difficulty.
So yes, it does take time, but maybe not as much as you project.
A take I’ve come to believe around generalization/sample efficiency is that if we do start to see AIs improving a lot at generalization/sample efficiency, we will see it first in robotics/robotics evals, and we have early signs of robotics/robotics evals starting to improve without robotics-specific data, which is bullish for generalization.
There are a couple of reasons for this:
A lot of the proposed use cases for robotics, like industry robots that handle varied environments or home robots require more generalization, especially from other data, unlike what has appeared to be the situation for tests and coding/easily grindable mathematics (and potentially white collar work.) Home robotics are an especially hard test case, because homes are spaces of infinite variation, long tails and constant change. As such, the scope of a home robot cannot be narrow. To create real value, it must perform reliably across open-ended variation with no special setup, and that is a practical test of embodied general intelligence.)
I’d say the major reason for this comes down to the fact that in areas where we don’t need much generalization, we already have good enough robots for various tasks, and there’s no overhang for AIs to take and become suddenly good without generalizing to other skills (unlike tests where IQ tests do measure human performance at lots of things somewhat well, or coding/easily grindable mathematics.)
One other thing that helps is that it’s much easier to restrain LLM training data around the physical world, partially from the world itself being only partial information, but also the fact that it’s much easier to restrain poorly generalizing AIs from just cheating it with a lot of training data, because it’s easier to avoid training data leakage/easier to bound how much AIs need to generalize. It’s also easier to give them hard problems, unlike benchmarks on the internet, because the model might just not be trained on robotics data (and this is feasible), or if it is done, it’s much easier to only give them training data on specific tasks and have a genuinely secure holdout test set, unlike pre-training where you have to give AIs more or less the entire internet, so you can’t train them on specific tasks and have a holdout test set, because the training data might leak information on how to solve the test set, and you have no way of controlling this in the digital world, but you can control data much more effectively in the physical world.
Given that generalization is a current bottleneck for robotics, this means that foundation models have to increase their generality of intelligence in order to improve on the tasks.
Now for the part about robotics improving, the current best sources are ACT-2 preview, where pre-training data (that isn’t robotics specific) improved out of domain performance dramatically (from ~15% to 99.1%) on folding laundry, and they have strict standards to ensure that generalization is the explanation for their improvements. The other one is Opus 5, which had also improved a lot on controlling robots at different tasks, namely stacking bowls and holding a fork (which is difficult for robots to grasp due to being thin and flat), purely through the power of it’s mind with zero robotics data.
If I’m right that robotics evals are the main place to look for when assessing LLM generalization/sample efficiency abilities, this looks a lot to me like pre-training is genuinely giving LLMs generalization, and Vanessa Kosoy asked a question about about whether LLMs have a positive amount but are worse than humans at the ability to effectively and continuously acquire deep knowledge and exploit this knowledge to construct and execute goal-directed plans over long lifetimes and consequence horizons, or have zero of the ability and are only doing it because pre-training is simply just getting human algorithms from the internet and once the internet is exhausted, AIs will stop getting better, and now the answer for 2026 LLMs is yes, they have a positive but subhuman amount the property of effectively and continuously acquire deep knowledge and exploit this knowledge to construct and execute goal-directed plans over long lifetimes and consequence horizons/generalization.
Some other implications of “robotics/robotics evals are where we will see generalization ability/sample efficiency, if it appears in AIs”:
It means the LLM paradigm is not completely off the mark, even if it requires major changes like continual learning/neuralese recurrence and memory, and there’s a real sense in which current LLMs generalize (Contra Tsvi or Thane Ruthenis’s hypotheses on LLMs)
It means that 10-20 year timelines are quite likely, even if the very fast scale up of the 2020s fails to produce AGI for whatever reason, and this means we need to be prepared for AGI/ASI to come in our lifetimes.
Relatedly, this also means there’s a plausible path for LLM researchers to work on to improve generalization.
Depending on how much RL was used for Opus 5, this might be a real “GPT-3 moment for RL”, where it generalized from lots of RL environments to new problems, demonstrating that RL can actually generalize beyond the trained tasks.
Finally, this might be able to replace the METR plots, assuming a time horizon equivalent can be made for physical tasks/robotics, and with much less worry about external validity/measurement problems.
Are there any takes from LLM skeptics like @TsviBT or @Thane Ruthenis about the idea of using robotics evals/robotics progress to predict generalization rather than using current benchmarks?
Thanks for A2A. I think it’s going to be hard to communicate about this because there are fundamental conceptual boundaries. It could help for you to write out the basic structure of the argument in simple concise syllogism format; and then expand on background enthymemes involved in the concept used for induction. This is to help address objections I might make about “wait does this concept actually support this inference / induction”? (Cf. https://www.lesswrong.com/posts/i7JSL5awGFcSRhyGF/shortform-2?commentId=fXtNjxLevGaiJHSyj )
In particular, in this case you’re using some concepts here:
And you’re doing an induction about it. And I’m just going to object that the way you’re using the concept doesn’t support the inference you’re trying to make; or at least, AFAICT you haven’t explained the concept and how you’re making the inference well enough to justify the inference.
The particular observational evidence you adduce seems interesting, but also I don’t have enough context to know what it means. (And I’m not really trying to evaluate it, because busy and because the disagreement seems to stem from the above, not from some specific observation. There are observations that would invalidate substantive chunks of my model (cf. https://www.lesswrong.com/posts/sTDfraZab47KiRMmT/views-on-when-agi-comes-and-on-strategy-to-reduce#comments), but generally when people bring in some observation the disagreement is how they are interpreting it, not the observation itself.)
Alright, here’s my simple syllogism/list:
The use-cases of robotics beyond current robots, especially home robotics require handling lots of environments and have varied challenges, so it stress-tests AI generalization.
Unlike other domains, there’s no unexploited algebraicness/overhang that would allow AIs to be useful for these cases without generalizing, because if there were such a thing, very limited current industrial robots would already have been used. They haven’t, so there’s no room for algebraicness/non-generalization to be an issue.
It’s easier to control/exclude training data when the domain is physical, rather than digital due to a multitude of reasons.
For these reasons, if current AIs improve on robotic tasks, it is a much stronger signal that they’re generalizing, and therefore is a signal that AGI could come in 10-20 years, and depending on how fast they improve, this might shift to 2-4 years.
So, the way I see Vanessa’s concept here is that it’s pointing to the concept of time horizons as defined by METR, which got introduced to the world by METR in 2025 called METR: Measuring AI Ability to Complete Long Tasks, and there’s a lot to be said here, including it’s problems, but the single most important thing to keep in mind here is that this is a continuous metric, so it’s meaningful to talk about how much the AI can acquire long-term knowledge to exploit it.
Now by itself, I would have viewed it as weak evidence, but combined with other evidence showing that exponential trends do exist in other domains, the fact that there are good general reasons to assume exponential growth in AI capabilities, plus the new robotics data shift me towards a view where the current LLM algebraicness/generalization issues are real, but probably fixable, and do not require throwing out the entire paradigm, and thus the induction/inference I’m making is valid here, or to put it another way the observations I have seen are more likely in worlds where LLMs can continuously acquire deep knowledge to exploit it than in worlds where they aren’t.
*Reads your model on progress to AGI, taking the definition of AGI in the post as given/assumed*
I very strongly disagree with this model of progress:
And instead believe in this model:
Or to put it another way, you seem to treat AGI as though there’s a discrete set of challenges to get there, one after the other to be solved, while I instead treat AGI as a set of challenges which all have continuous metrics, and in particular there’s ways to reuse previous progress, because the load-bearing elements of AGI depend less on a particular paradigm/theory of intelligence and they depend more so on compute.
Put another way, you expect sharp/discontinuous progress towards AGI and a singularity, whereas I think we will smoothly go towards AGI and a singularity (note this isn’t talking about how fast or slow the takeoff is.)
A key reason for this is I think the history of AI progress for the last 10-15 years contributes way more evidence for this question than you seem to think, and more generally based on other evidence like papers on human brain scaling laws, where human brains turned out to have a more efficient scaling law than other animals with larger brains, the fact that our brains need a disproportinately large amount of energy to run unlike many animals, which suggests that the human brain is using much more compute, and finally, the most interesting post by gwern on how the insight generator called the Default Mode Network in our brains is largely powered by lots of compute time at inference/test-time compute, and it imposes an effective tax of 1-2 OOMs of compute in order to produce insights, and is a major differentiator of humans (though primates and rats likely do have it as well, but is much less effective due to worse scaling laws and smaller brains in the first place)
This also solves disagreements with the inputs section of the post.
To simplify the disagreement in one sentence, I notice that the central crux of whether or not the load-bearing AI progress is in compute/dat or conceptual breakthroughs on theories of intelligence keeps explaining a lot (but not all) of our disagreements, and disagreements expectations of how smooth vs discrete progress is in AI explain why I expect very different things around AI progress.
“it stress-tests AI generalization” But does it though? What does it stress test? That seems to assert a generalization from AI generalization about XYZ robotics tasks to ABC intelligence explosion tasks (or something along these lines?). What’s that generalization / what justifies it?
I don’t understand this. Industrial robots of course are used very widely? Or I guess you typoed, and you’re trying to say, LLMs or other AIs would have been used to compute very non-obvious complex actions that take advantage of algebraicness in the narrow domain? Or you’re saying something about humanoid robots, or about softer / less rote tasks (laundry folding rather than installing a chair in a car)?
How do you get from the qualitative and relative statement about robotics being “a much stronger signal” to numbers and years?
My guess is that these things are not actually cruxes? Not sure. Instead I’d guess that the cruxes are more simply
the quantitative amount of insight (regardless of continuousness) remaining to get to AGI
the degree to which those insights are blocked on difficult thinking that current AI doesn’t accelerate by much
I agree industrial robots that don’t generalize much are used quite widely in narrow environments, the point is that for the set of tasks where AIs don’t generalize much/don’t improve generalization, they can’t improve industrial robotics/cause the robotics industry to be more useful and profitable, because the robotics industry already squeezed out ~all the gains from AIs that don’t generalize well to new environments.
Yes, this is about less rote tasks like laundry folding across arbitrarily oriented clothing and being crumpled, plus being in arbitrary locations in the home, while a human constantly disturbs them and provides OOD pressure (indeed one of my examples was about this exact thing being done well by an AI that controlled a robot using only pre-training without robotics data), and due to the way home robotics has infinite variation, long tails and constant change, this means it’s way easier to measure generalization cleanly/in a controlled manner.
The stress test of generalization is more indirect than this. My model is that as AIs improve at more and more robotics tasks, it simultaneously has to be able to do well in more and more messy and OOD environments (otherwise current robotics would have already improved more than they actually are), intelligence explosion tasks, assuming an intelligence explosion/software only singularity is possible, involves a great many very messy and OOD tasks, thus robotic evals are real evidence that AIs are increasingly able to start intelligence explosion tasks.
Fair point, I don’t think I justified this very well, at least as a claim about the specific units, but a vague argument here is that given the qualitative statement above, it implies that the current paradigm is not so far off the mark that it’s absolutely useless as progress for AGI design, and this means there’s a real path that doesn’t require heroic genius from today to AGI.
To address your cruxes:
My views on this is that I think at most 2 insights are necessary in practice, and if the human population kept perpetually growing according to 1880s-era birth-rates, we’d need 0 insights, as the super-exponential growth of the economy would have allowed us to get to AGI eventually.
I’d say they’re non-trivially serial (especially if the solutions use a lot of RL, as RL is much, much more difficult to parallelize outside of current environments), but a fairly key part of my mental model is that conditional on current LLMs becoming blocked, there will be massive incentives to point researchers into promising paradigms, meaning that more resources will be pursued on the blockers, compensating for their difficulty.
So yes, it does take time, but maybe not as much as you project.