Maybe it sounds absurd for me to compare what one RLVR training can do, versus the whole edifice of ideas painstakingly built by the mathematics community over the course of 200 years. But it’s not absurd: AlphaZero really did blow past the whole edifice of ideas painstakingly built over the course of centuries by the chess and go communities in its 72-hour training runs. So the idea of building real new knowledge at a massive scale through RL is not absurd on its face.
I should flag here that Go has a number of properties that aren’t replicable in many other domains, and thus we shouldn’t anchor our estimates of takeoff in a given domain or in general to Go:
It’s very parallelizable, like easily parallelizable mathematics/SWE work, which is something a lot of domains lack. In particular, non-trivial serial feedback loops of the order of weeks, months or even years/decades in the most complicated, taste-driven tasks is much more common than people intuitively expect. This is related to failures of the factored cognition hypothesis in practice, and Dwarkesh talks about this in this podcast episode.
Go has a ~perfect simulator, and it’s not a coincidence that RLVR makes the greatest strides/adds new capabilities to LLMs precisely when they either have a very good simulator or a very good way to verify if you correctly did the thing or not. Again, this is lacking for many essential tasks for AIs to transform the economy. (This is discussed in AI safety/capability discourse as the difference between clean tasks vs messy tasks, where messy tasks lack both a good simulator and a good verifier, whereas clean tasks only need one of them to hold,)
And that’s more or less the entire issue here.
Notably, this does not mean that AIs can’t get good at the necessary messy tasks, or that the takeoff if directed by AGIs won’t lead to intuitively insane outcomes (Damon Binder’s series on robotics, AGI, and growth rates, especially parts 1, 4and 5are enough here), but it does mean that we should initially expect much slower takeoff speeds than the Go anchor/AlphaZero suggests.
I was making a narrow point about what I mean by “LLMs are (still) mostly powered by imitative learning, not RL”.
You seem to be saying: “Nobody ever would have expected that RLVR to build a whole giant new edifice of deep interconnected knowledge into an LLM, the way AlphaZero’s MCTS+RL training did. Duh, that’s obvious and overdetermined. So why are we even talking about that?”
But to the extent that that’s true, it would make my point stronger, not weaker. Because imitative learning can (and does) build a whole giant new edifice of deep interconnected knowledge into an LLM.
It seems like you’re trying to argue about takeoff speeds instead? If so, that seems off-topic here, I think. My take on that is at Foom & Doom 1: “Brain in a box in a basement”, including the part starting at “To be clear, the resulting ASI after those 0–2 years would not be an AI that already knows everything about everything…”
I should flag here that Go has a number of properties that aren’t replicable in many other domains, and thus we shouldn’t anchor our estimates of takeoff in a given domain or in general to Go:
It’s very parallelizable, like easily parallelizable mathematics/SWE work, which is something a lot of domains lack. In particular, non-trivial serial feedback loops of the order of weeks, months or even years/decades in the most complicated, taste-driven tasks is much more common than people intuitively expect. This is related to failures of the factored cognition hypothesis in practice, and Dwarkesh talks about this in this podcast episode.
Go has a ~perfect simulator, and it’s not a coincidence that RLVR makes the greatest strides/adds new capabilities to LLMs precisely when they either have a very good simulator or a very good way to verify if you correctly did the thing or not. Again, this is lacking for many essential tasks for AIs to transform the economy. (This is discussed in AI safety/capability discourse as the difference between clean tasks vs messy tasks, where messy tasks lack both a good simulator and a good verifier, whereas clean tasks only need one of them to hold,)
And that’s more or less the entire issue here.
Notably, this does not mean that AIs can’t get good at the necessary messy tasks, or that the takeoff if directed by AGIs won’t lead to intuitively insane outcomes (Damon Binder’s series on robotics, AGI, and growth rates, especially parts 1, 4 and 5 are enough here), but it does mean that we should initially expect much slower takeoff speeds than the Go anchor/AlphaZero suggests.
I was making a narrow point about what I mean by “LLMs are (still) mostly powered by imitative learning, not RL”.
You seem to be saying: “Nobody ever would have expected that RLVR to build a whole giant new edifice of deep interconnected knowledge into an LLM, the way AlphaZero’s MCTS+RL training did. Duh, that’s obvious and overdetermined. So why are we even talking about that?”
But to the extent that that’s true, it would make my point stronger, not weaker. Because imitative learning can (and does) build a whole giant new edifice of deep interconnected knowledge into an LLM.
It seems like you’re trying to argue about takeoff speeds instead? If so, that seems off-topic here, I think. My take on that is at Foom & Doom 1: “Brain in a box in a basement”, including the part starting at “To be clear, the resulting ASI after those 0–2 years would not be an AI that already knows everything about everything…”