For me, a big part of the reason is that it went off topic immediately in an area that’s still charged by politics, so I didn’t want the discussion to centralize around something that would be predictably useless but very dramatic/distracting.
In essence, I’m trying to cut short a likely demon thread around COVID.
I agree industrial robots that don’t generalize much are used quite widely in narrow environments, the point is that for the set of tasks where AIs don’t generalize much/don’t improve generalization, they can’t improve industrial robotics/cause the robotics industry to be more useful and profitable, because the robotics industry already squeezed out ~all the gains from AIs that don’t generalize well to new environments.
Yes, this is about less rote tasks like laundry folding across arbitrarily oriented clothing and being crumpled, plus being in arbitrary locations in the home, while a human constantly disturbs them and provides OOD pressure (indeed one of my examples was about this exact thing being done well by an AI that controlled a robot using only pre-training without robotics data), and due to the way home robotics has infinite variation, long tails and constant change, this means it’s way easier to measure generalization cleanly/in a controlled manner.
The stress test of generalization is more indirect than this. My model is that as AIs improve at more and more robotics tasks, it simultaneously has to be able to do well in more and more messy and OOD environments (otherwise current robotics would have already improved more than they actually are), intelligence explosion tasks, assuming an intelligence explosion/software only singularity is possible, involves a great many very messy and OOD tasks, thus robotic evals are real evidence that AIs are increasingly able to start intelligence explosion tasks.
Fair point, I don’t think I justified this very well, at least as a claim about the specific units, but a vague argument here is that given the qualitative statement above, it implies that the current paradigm is not so far off the mark that it’s absolutely useless as progress for AGI design, and this means there’s a real path that doesn’t require heroic genius from today to AGI.
To address your cruxes:
My views on this is that I think at most 2 insights are necessary in practice, and if the human population kept perpetually growing according to 1880s-era birth-rates, we’d need 0 insights, as the super-exponential growth of the economy would have allowed us to get to AGI eventually.
I’d say they’re non-trivially serial (especially if the solutions use a lot of RL, as RL is much, much more difficult to parallelize outside of current environments), but a fairly key part of my mental model is that conditional on current LLMs becoming blocked, there will be massive incentives to point researchers into promising paradigms, meaning that more resources will be pursued on the blockers, compensating for their difficulty.
So yes, it does take time, but maybe not as much as you project.