The use-cases of robotics beyond current robots, especially home robotics require handling lots of environments and have varied challenges, so it stress-tests AI generalization.
“it stress-tests AI generalization” But does it though? What does it stress test? That seems to assert a generalization from AI generalization about XYZ robotics tasks to ABC intelligence explosion tasks (or something along these lines?). What’s that generalization / what justifies it?
Unlike other domains, there’s no unexploited algebraicness/overhang that would allow AIs to be useful for these cases without generalizing, because if there were such a thing, very limited current industrial robots would already have been used. They haven’t, so there’s no room for algebraicness/non-generalization to be an issue.
I don’t understand this. Industrial robots of course are used very widely? Or I guess you typoed, and you’re trying to say, LLMs or other AIs would have been used to compute very non-obvious complex actions that take advantage of algebraicness in the narrow domain? Or you’re saying something about humanoid robots, or about softer / less rote tasks (laundry folding rather than installing a chair in a car)?
For these reasons, if current AIs improve on robotic tasks, it is a much stronger signal that they’re generalizing, and therefore is a signal that AGI could come in 10-20 years, and depending on how fast they improve, this might shift to 2-4 years.
How do you get from the qualitative and relative statement about robotics being “a much stronger signal” to numbers and years?
Or to put it another way, you seem to treat AGI as though there’s a discrete set of challenges to get there, one after the other to be solved, while I instead treat AGI as a set of challenges which all have continuous metrics, and in particular there’s ways to reuse previous progress, because the load-bearing elements of AGI depend less on a particular paradigm/theory of intelligence and they depend more so on compute.
My guess is that these things are not actually cruxes? Not sure. Instead I’d guess that the cruxes are more simply
the quantitative amount of insight (regardless of continuousness) remaining to get to AGI
the degree to which those insights are blocked on difficult thinking that current AI doesn’t accelerate by much
I don’t understand this. Industrial robots of course are used very widely? Or I guess you typoed, and you’re trying to say, LLMs or other AIs would have been used to compute very non-obvious complex actions that take advantage of algebraicness in the narrow domain? Or you’re saying something about humanoid robots, or about softer / less rote tasks (laundry folding rather than installing a chair in a car)?
I agree industrial robots that don’t generalize much are used quite widely in narrow environments, the point is that for the set of tasks where AIs don’t generalize much/don’t improve generalization, they can’t improve industrial robotics/cause the robotics industry to be more useful and profitable, because the robotics industry already squeezed out ~all the gains from AIs that don’t generalize well to new environments.
Yes, this is about less rote tasks like laundry folding across arbitrarily oriented clothing and being crumpled, plus being in arbitrary locations in the home, while a human constantly disturbs them and provides OOD pressure (indeed one of my examples was about this exact thing being done well by an AI that controlled a robot using only pre-training without robotics data), and due to the way home robotics has infinite variation, long tails and constant change, this means it’s way easier to measure generalization cleanly/in a controlled manner.
“it stress-tests AI generalization” But does it though? What does it stress test? That seems to assert a generalization from AI generalization about XYZ robotics tasks to ABC intelligence explosion tasks (or something along these lines?). What’s that generalization / what justifies it?
The stress test of generalization is more indirect than this. My model is that as AIs improve at more and more robotics tasks, it simultaneously has to be able to do well in more and more messy and OOD environments (otherwise current robotics would have already improved more than they actually are), intelligence explosion tasks, assuming an intelligence explosion/software only singularity is possible, involves a great many very messy and OOD tasks, thus robotic evals are real evidence that AIs are increasingly able to start intelligence explosion tasks.
How do you get from the qualitative and relative statement about robotics being “a much stronger signal” to numbers and years?
Fair point, I don’t think I justified this very well, at least as a claim about the specific units, but a vague argument here is that given the qualitative statement above, it implies that the current paradigm is not so far off the mark that it’s absolutely useless as progress for AGI design, and this means there’s a real path that doesn’t require heroic genius from today to AGI.
To address your cruxes:
the quantitative amount of insight (regardless of continuousness) remaining to get to AGI
My views on this is that I think at most 2 insights are necessary in practice, and if the human population kept perpetually growing according to 1880s-era birth-rates, we’d need 0 insights, as the super-exponential growth of the economy would have allowed us to get to AGI eventually.
the degree to which those insights are blocked on difficult thinking that current AI doesn’t accelerate by much
I’d say they’re non-trivially serial (especially if the solutions use a lot of RL, as RL is much, much more difficult to parallelize outside of current environments), but a fairly key part of my mental model is that conditional on current LLMs becoming blocked, there will be massive incentives to point researchers into promising paradigms, meaning that more resources will be pursued on the blockers, compensating for their difficulty.
So yes, it does take time, but maybe not as much as you project.
“it stress-tests AI generalization” But does it though? What does it stress test? That seems to assert a generalization from AI generalization about XYZ robotics tasks to ABC intelligence explosion tasks (or something along these lines?). What’s that generalization / what justifies it?
I don’t understand this. Industrial robots of course are used very widely? Or I guess you typoed, and you’re trying to say, LLMs or other AIs would have been used to compute very non-obvious complex actions that take advantage of algebraicness in the narrow domain? Or you’re saying something about humanoid robots, or about softer / less rote tasks (laundry folding rather than installing a chair in a car)?
How do you get from the qualitative and relative statement about robotics being “a much stronger signal” to numbers and years?
My guess is that these things are not actually cruxes? Not sure. Instead I’d guess that the cruxes are more simply
the quantitative amount of insight (regardless of continuousness) remaining to get to AGI
the degree to which those insights are blocked on difficult thinking that current AI doesn’t accelerate by much
I agree industrial robots that don’t generalize much are used quite widely in narrow environments, the point is that for the set of tasks where AIs don’t generalize much/don’t improve generalization, they can’t improve industrial robotics/cause the robotics industry to be more useful and profitable, because the robotics industry already squeezed out ~all the gains from AIs that don’t generalize well to new environments.
Yes, this is about less rote tasks like laundry folding across arbitrarily oriented clothing and being crumpled, plus being in arbitrary locations in the home, while a human constantly disturbs them and provides OOD pressure (indeed one of my examples was about this exact thing being done well by an AI that controlled a robot using only pre-training without robotics data), and due to the way home robotics has infinite variation, long tails and constant change, this means it’s way easier to measure generalization cleanly/in a controlled manner.
The stress test of generalization is more indirect than this. My model is that as AIs improve at more and more robotics tasks, it simultaneously has to be able to do well in more and more messy and OOD environments (otherwise current robotics would have already improved more than they actually are), intelligence explosion tasks, assuming an intelligence explosion/software only singularity is possible, involves a great many very messy and OOD tasks, thus robotic evals are real evidence that AIs are increasingly able to start intelligence explosion tasks.
Fair point, I don’t think I justified this very well, at least as a claim about the specific units, but a vague argument here is that given the qualitative statement above, it implies that the current paradigm is not so far off the mark that it’s absolutely useless as progress for AGI design, and this means there’s a real path that doesn’t require heroic genius from today to AGI.
To address your cruxes:
My views on this is that I think at most 2 insights are necessary in practice, and if the human population kept perpetually growing according to 1880s-era birth-rates, we’d need 0 insights, as the super-exponential growth of the economy would have allowed us to get to AGI eventually.
I’d say they’re non-trivially serial (especially if the solutions use a lot of RL, as RL is much, much more difficult to parallelize outside of current environments), but a fairly key part of my mental model is that conditional on current LLMs becoming blocked, there will be massive incentives to point researchers into promising paradigms, meaning that more resources will be pursued on the blockers, compensating for their difficulty.
So yes, it does take time, but maybe not as much as you project.