The use-cases of robotics beyond current robots, especially home robotics require handling lots of environments and have varied challenges, so it stress-tests AI generalization.
Unlike other domains, there’s no unexploited algebraicness/overhang that would allow AIs to be useful for these cases without generalizing, because if there were such a thing, very limited current industrial robots would already have been used. They haven’t, so there’s no room for algebraicness/non-generalization to be an issue.
It’s easier to control/exclude training data when the domain is physical, rather than digital due to a multitude of reasons.
For these reasons, if current AIs improve on robotic tasks, it is a much stronger signal that they’re generalizing, and therefore is a signal that AGI could come in 10-20 years, and depending on how fast they improve, this might shift to 2-4 years.
In particular, in this case you’re using some concepts here:
they have a positive but subhuman amount the property of effectively and continuously acquire deep knowledge and exploit this knowledge to construct and execute goal-directed plans over long lifetimes and consequence horizons/generalization.
And you’re doing an induction about it. And I’m just going to object that the way you’re using the concept doesn’t support the inference you’re trying to make; or at least, AFAICT you haven’t explained the concept and how you’re making the inference well enough to justify the inference.
So, the way I see Vanessa’s concept here is that it’s pointing to the concept of time horizons as defined by METR, which got introduced to the world by METR in 2025 called METR: Measuring AI Ability to Complete Long Tasks, and there’s a lot to be said here, including it’s problems, but the single most important thing to keep in mind here is that this is a continuous metric, so it’s meaningful to talk about how much the AI can acquire long-term knowledge to exploit it.
Now by itself, I would have viewed it as weak evidence, but combined with other evidence showing that exponential trends do exist in other domains, the fact that there are good general reasons to assume exponential growth in AI capabilities, plus the new robotics data shift me towards a view where the current LLM algebraicness/generalization issues are real, but probably fixable, and do not require throwing out the entire paradigm, and thus the induction/inference I’m making is valid here, or to put it another way the observations I have seen are more likely in worlds where LLMs can continuously acquire deep knowledge to exploit it than in worlds where they aren’t.
*Reads your model on progress to AGI, taking the definition of AGI in the post as given/assumed*
I very strongly disagree with this model of progress:
The first way humanity makes AGI is by combining some set of significant ideas about intelligence. Significant ideas are things like (the ideas of) gradient descent, recombination, probability distributions, universal computation, search, world-optimization. Significant ideas are to a significant extent bottlenecked on great natural philosophers doing great natural philosophy about intelligence, with sequential bottlenecks between many insights.
And instead believe in this model:
Algorithmic progress/good ideas/research taste do genuinely matter, but most people are prone to overestimate how much they matter without compute or data, and the main way they currently matter is by being more-compute/data efficient at larger compute scales, and importantly the current paradigm isn’t irrelevant to building the next one, because you can reuse the compute and data from the old paradigm, and the new paradigms will not reduce the demand for compute.
This means progress to AGI isn’t as serial as you state (though I do expect serialness in a different way through RL, but you can still reuse compute/training data you already bought once you have a better paradigm).
More generally, what has made AI go from being ~completely useless and generalization was fake to current AIs where generalization exists, but is worse than humans and can do real economic work is basically just compute and data (and note that in a RL/synthetic data heavy world, data bottlenecks are folded into compute bottlenecks, so compute is the main input)
Or to put it another way, you seem to treat AGI as though there’s a discrete set of challenges to get there, one after the other to be solved, while I instead treat AGI as a set of challenges which all have continuous metrics, and in particular there’s ways to reuse previous progress, because the load-bearing elements of AGI depend less on a particular paradigm/theory of intelligence and they depend more so on compute.
Put another way, you expect sharp/discontinuous progress towards AGI and a singularity, whereas I think we will smoothly go towards AGI and a singularity (note this isn’t talking about how fast or slow the takeoff is.)
This also solves disagreements with the inputs section of the post.
To simplify the disagreement in one sentence, I notice that the central crux of whether or not the load-bearing AI progress is in compute/dat or conceptual breakthroughs on theories of intelligence keeps explaining a lot (but not all) of our disagreements, and disagreements expectations of how smooth vs discrete progress is in AI explain why I expect very different things around AI progress.
The use-cases of robotics beyond current robots, especially home robotics require handling lots of environments and have varied challenges, so it stress-tests AI generalization.
“it stress-tests AI generalization” But does it though? What does it stress test? That seems to assert a generalization from AI generalization about XYZ robotics tasks to ABC intelligence explosion tasks (or something along these lines?). What’s that generalization / what justifies it?
Unlike other domains, there’s no unexploited algebraicness/overhang that would allow AIs to be useful for these cases without generalizing, because if there were such a thing, very limited current industrial robots would already have been used. They haven’t, so there’s no room for algebraicness/non-generalization to be an issue.
I don’t understand this. Industrial robots of course are used very widely? Or I guess you typoed, and you’re trying to say, LLMs or other AIs would have been used to compute very non-obvious complex actions that take advantage of algebraicness in the narrow domain? Or you’re saying something about humanoid robots, or about softer / less rote tasks (laundry folding rather than installing a chair in a car)?
For these reasons, if current AIs improve on robotic tasks, it is a much stronger signal that they’re generalizing, and therefore is a signal that AGI could come in 10-20 years, and depending on how fast they improve, this might shift to 2-4 years.
How do you get from the qualitative and relative statement about robotics being “a much stronger signal” to numbers and years?
Or to put it another way, you seem to treat AGI as though there’s a discrete set of challenges to get there, one after the other to be solved, while I instead treat AGI as a set of challenges which all have continuous metrics, and in particular there’s ways to reuse previous progress, because the load-bearing elements of AGI depend less on a particular paradigm/theory of intelligence and they depend more so on compute.
My guess is that these things are not actually cruxes? Not sure. Instead I’d guess that the cruxes are more simply
the quantitative amount of insight (regardless of continuousness) remaining to get to AGI
the degree to which those insights are blocked on difficult thinking that current AI doesn’t accelerate by much
I don’t understand this. Industrial robots of course are used very widely? Or I guess you typoed, and you’re trying to say, LLMs or other AIs would have been used to compute very non-obvious complex actions that take advantage of algebraicness in the narrow domain? Or you’re saying something about humanoid robots, or about softer / less rote tasks (laundry folding rather than installing a chair in a car)?
I agree industrial robots that don’t generalize much are used quite widely in narrow environments, the point is that for the set of tasks where AIs don’t generalize much/don’t improve generalization, they can’t improve industrial robotics/cause the robotics industry to be more useful and profitable, because the robotics industry already squeezed out ~all the gains from AIs that don’t generalize well to new environments.
Yes, this is about less rote tasks like laundry folding across arbitrarily oriented clothing and being crumpled, plus being in arbitrary locations in the home, while a human constantly disturbs them and provides OOD pressure (indeed one of my examples was about this exact thing being done well by an AI that controlled a robot using only pre-training without robotics data), and due to the way home robotics has infinite variation, long tails and constant change, this means it’s way easier to measure generalization cleanly/in a controlled manner.
“it stress-tests AI generalization” But does it though? What does it stress test? That seems to assert a generalization from AI generalization about XYZ robotics tasks to ABC intelligence explosion tasks (or something along these lines?). What’s that generalization / what justifies it?
The stress test of generalization is more indirect than this. My model is that as AIs improve at more and more robotics tasks, it simultaneously has to be able to do well in more and more messy and OOD environments (otherwise current robotics would have already improved more than they actually are), intelligence explosion tasks, assuming an intelligence explosion/software only singularity is possible, involves a great many very messy and OOD tasks, thus robotic evals are real evidence that AIs are increasingly able to start intelligence explosion tasks.
How do you get from the qualitative and relative statement about robotics being “a much stronger signal” to numbers and years?
Fair point, I don’t think I justified this very well, at least as a claim about the specific units, but a vague argument here is that given the qualitative statement above, it implies that the current paradigm is not so far off the mark that it’s absolutely useless as progress for AGI design, and this means there’s a real path that doesn’t require heroic genius from today to AGI.
To address your cruxes:
the quantitative amount of insight (regardless of continuousness) remaining to get to AGI
My views on this is that I think at most 2 insights are necessary in practice, and if the human population kept perpetually growing according to 1880s-era birth-rates, we’d need 0 insights, as the super-exponential growth of the economy would have allowed us to get to AGI eventually.
the degree to which those insights are blocked on difficult thinking that current AI doesn’t accelerate by much
I’d say they’re non-trivially serial (especially if the solutions use a lot of RL, as RL is much, much more difficult to parallelize outside of current environments), but a fairly key part of my mental model is that conditional on current LLMs becoming blocked, there will be massive incentives to point researchers into promising paradigms, meaning that more resources will be pursued on the blockers, compensating for their difficulty.
So yes, it does take time, but maybe not as much as you project.
Alright, here’s my simple syllogism/list:
The use-cases of robotics beyond current robots, especially home robotics require handling lots of environments and have varied challenges, so it stress-tests AI generalization.
Unlike other domains, there’s no unexploited algebraicness/overhang that would allow AIs to be useful for these cases without generalizing, because if there were such a thing, very limited current industrial robots would already have been used. They haven’t, so there’s no room for algebraicness/non-generalization to be an issue.
It’s easier to control/exclude training data when the domain is physical, rather than digital due to a multitude of reasons.
For these reasons, if current AIs improve on robotic tasks, it is a much stronger signal that they’re generalizing, and therefore is a signal that AGI could come in 10-20 years, and depending on how fast they improve, this might shift to 2-4 years.
So, the way I see Vanessa’s concept here is that it’s pointing to the concept of time horizons as defined by METR, which got introduced to the world by METR in 2025 called METR: Measuring AI Ability to Complete Long Tasks, and there’s a lot to be said here, including it’s problems, but the single most important thing to keep in mind here is that this is a continuous metric, so it’s meaningful to talk about how much the AI can acquire long-term knowledge to exploit it.
Now by itself, I would have viewed it as weak evidence, but combined with other evidence showing that exponential trends do exist in other domains, the fact that there are good general reasons to assume exponential growth in AI capabilities, plus the new robotics data shift me towards a view where the current LLM algebraicness/generalization issues are real, but probably fixable, and do not require throwing out the entire paradigm, and thus the induction/inference I’m making is valid here, or to put it another way the observations I have seen are more likely in worlds where LLMs can continuously acquire deep knowledge to exploit it than in worlds where they aren’t.
*Reads your model on progress to AGI, taking the definition of AGI in the post as given/assumed*
I very strongly disagree with this model of progress:
And instead believe in this model:
Or to put it another way, you seem to treat AGI as though there’s a discrete set of challenges to get there, one after the other to be solved, while I instead treat AGI as a set of challenges which all have continuous metrics, and in particular there’s ways to reuse previous progress, because the load-bearing elements of AGI depend less on a particular paradigm/theory of intelligence and they depend more so on compute.
Put another way, you expect sharp/discontinuous progress towards AGI and a singularity, whereas I think we will smoothly go towards AGI and a singularity (note this isn’t talking about how fast or slow the takeoff is.)
A key reason for this is I think the history of AI progress for the last 10-15 years contributes way more evidence for this question than you seem to think, and more generally based on other evidence like papers on human brain scaling laws, where human brains turned out to have a more efficient scaling law than other animals with larger brains, the fact that our brains need a disproportinately large amount of energy to run unlike many animals, which suggests that the human brain is using much more compute, and finally, the most interesting post by gwern on how the insight generator called the Default Mode Network in our brains is largely powered by lots of compute time at inference/test-time compute, and it imposes an effective tax of 1-2 OOMs of compute in order to produce insights, and is a major differentiator of humans (though primates and rats likely do have it as well, but is much less effective due to worse scaling laws and smaller brains in the first place)
This also solves disagreements with the inputs section of the post.
To simplify the disagreement in one sentence, I notice that the central crux of whether or not the load-bearing AI progress is in compute/dat or conceptual breakthroughs on theories of intelligence keeps explaining a lot (but not all) of our disagreements, and disagreements expectations of how smooth vs discrete progress is in AI explain why I expect very different things around AI progress.
“it stress-tests AI generalization” But does it though? What does it stress test? That seems to assert a generalization from AI generalization about XYZ robotics tasks to ABC intelligence explosion tasks (or something along these lines?). What’s that generalization / what justifies it?
I don’t understand this. Industrial robots of course are used very widely? Or I guess you typoed, and you’re trying to say, LLMs or other AIs would have been used to compute very non-obvious complex actions that take advantage of algebraicness in the narrow domain? Or you’re saying something about humanoid robots, or about softer / less rote tasks (laundry folding rather than installing a chair in a car)?
How do you get from the qualitative and relative statement about robotics being “a much stronger signal” to numbers and years?
My guess is that these things are not actually cruxes? Not sure. Instead I’d guess that the cruxes are more simply
the quantitative amount of insight (regardless of continuousness) remaining to get to AGI
the degree to which those insights are blocked on difficult thinking that current AI doesn’t accelerate by much
I agree industrial robots that don’t generalize much are used quite widely in narrow environments, the point is that for the set of tasks where AIs don’t generalize much/don’t improve generalization, they can’t improve industrial robotics/cause the robotics industry to be more useful and profitable, because the robotics industry already squeezed out ~all the gains from AIs that don’t generalize well to new environments.
Yes, this is about less rote tasks like laundry folding across arbitrarily oriented clothing and being crumpled, plus being in arbitrary locations in the home, while a human constantly disturbs them and provides OOD pressure (indeed one of my examples was about this exact thing being done well by an AI that controlled a robot using only pre-training without robotics data), and due to the way home robotics has infinite variation, long tails and constant change, this means it’s way easier to measure generalization cleanly/in a controlled manner.
The stress test of generalization is more indirect than this. My model is that as AIs improve at more and more robotics tasks, it simultaneously has to be able to do well in more and more messy and OOD environments (otherwise current robotics would have already improved more than they actually are), intelligence explosion tasks, assuming an intelligence explosion/software only singularity is possible, involves a great many very messy and OOD tasks, thus robotic evals are real evidence that AIs are increasingly able to start intelligence explosion tasks.
Fair point, I don’t think I justified this very well, at least as a claim about the specific units, but a vague argument here is that given the qualitative statement above, it implies that the current paradigm is not so far off the mark that it’s absolutely useless as progress for AGI design, and this means there’s a real path that doesn’t require heroic genius from today to AGI.
To address your cruxes:
My views on this is that I think at most 2 insights are necessary in practice, and if the human population kept perpetually growing according to 1880s-era birth-rates, we’d need 0 insights, as the super-exponential growth of the economy would have allowed us to get to AGI eventually.
I’d say they’re non-trivially serial (especially if the solutions use a lot of RL, as RL is much, much more difficult to parallelize outside of current environments), but a fairly key part of my mental model is that conditional on current LLMs becoming blocked, there will be massive incentives to point researchers into promising paradigms, meaning that more resources will be pursued on the blockers, compensating for their difficulty.
So yes, it does take time, but maybe not as much as you project.