CoT that’s legible isn’t good evidence that LLMs are not wielding (or could not wield) novel latent abstractions/reasoning primitives/control heuristics.
I feel like you’re saying: LLMs can wield more than zero novel latent abstractions/reasoning primitives/control heuristics while CoT remains generally legible.
Whereas what I’m saying is: LLMs cannot be totally transformed by RLVR while CoT remains generally legible.
These aren’t contradictory. I think both are true.
As an example, think about humans learning things, like a teen going from her first number theory class as a teen at time 0, to deeply understanding very advanced math (e.g. the Langlands program) as an adult at time T = many years later. It’s an arduous and time-consuming process. And her notes at time T would be deeply, deeply inscrutable from the perspective of her former teen self at time 0—even the notes that are in the form of words rather than symbols.
This suggests that the delta between pre-RLVR vs post-RLVR LLMs is much less of a wrenching change than the delta between the teen at time 0 vs the now-adult mathematician at time T.
Novelty builds heavily on existing frameworks and often involves a subtle reframe, reconfiguration, or saliencing of a known/slightly modified concept in a new setting. Insight is about discovering relevance, but what’s available to be relevant is often familiar primitives.
Mathematical work is a great example here, since mathematical discoveries very often have a subtle kernel of insight, a small new idea, that reconfigures existing concepts around a problem in a significant way to illuminate something unknown. The reconfiguration is almost all in terms of familiar stuff—the moving pieces don’t change much. In the case of LLMs, that’s the content imitative learning provides.
I think instead of “novelty” here you should have said “a sufficiently small increment of novelty”.
If we instead consider mathematics as a collective human enterprise, it went from “number theory doesn’t exist at all” to the Langlands program, over the course of 200 years.
Maybe it sounds absurd for me to compare what one RLVR training can do, versus the whole edifice of ideas painstakingly built by the mathematics community over the course of 200 years. But it’s not absurd: AlphaZero really did blow past the whole edifice of ideas painstakingly built over the course of centuries by the chess and go communities in its 72-hour training runs. So the idea of building real new knowledge at a massive scale through RL is not absurd on its face. I’m just saying: RLVR-on-LLMs is not doing anything like that, at least not today. Rather, the LLM approach is to use imitative learning to suck in the whole edifice of ideas painstakingly built by humans, and then tweak it a bit at the end via RLVR.
Relatedly, when mathematicians study the recent LLM-generated math results, they’re not “learning something new” in a way that’s analogous to that teen spending years poring over her math textbooks. Rather they’re “learning something new” in a way that’s analogous to some guy telling me what his name is. If the mathematicians already have all the right background knowledge, they can quickly understand the solution within their existing conceptual frameworks. Or if they don’t already have all the right background knowledge, they can read human-created textbooks to get it. (I’m mainly thinking of the unit distance conjecture; in other examples that I looked into like the Jacobian conjecture, IIUC, the CoTs weren’t released, so mathematicians are still be a bit puzzled about how the LLMs came upon the answer.)
…LLMs too, seem to perform much of their computation through nonlinguistic representations and compress the actionable results into language, mostly because of the structural thing that language is the medium through which they maintain serial state and communicate.
This is an argument against
if you create AI capabilities via imitative learning, you get models that follow the human distribution of outputs
I don’t think it’s an argument against that, or sorry if I’m misunderstanding. You’re saying that LLMs may be computing their outputs in a different way from humans. Fine. But their outputs are still following the human distribution of outputs. “Following the human distribution of outputs” is just another way to say “low perplexity”, right?
(I will concede that the phrase “the human distribution of outputs” is a bit misleading on various other grounds, e.g. LLMs may generalize OOD in a different way from humans; not all imitative learning training tokens are created by humans; etc.)
I feel like you’re saying: LLMs can wield more than zero novel latent abstractions/reasoning primitives/control heuristics while CoT remains generally legible.
Whereas what I’m saying is: LLMs cannot be totally transformed by RLVR while CoT remains generally legible.
These aren’t contradictory. I think both are true.
As an example, think about humans learning things, like a teen going from her first number theory class as a teen at time 0, to deeply understanding very advanced math (e.g. the Langlands program) as an adult at time T = many years later. It’s an arduous and time-consuming process. And her notes at time T would be deeply, deeply inscrutable from the perspective of her former teen self at time 0—even the notes that are in the form of words rather than symbols.
This suggests that the delta between pre-RLVR vs post-RLVR LLMs is much less of a wrenching change than the delta between the teen at time 0 vs the now-adult mathematician at time T.
I think instead of “novelty” here you should have said “a sufficiently small increment of novelty”.
If we instead consider mathematics as a collective human enterprise, it went from “number theory doesn’t exist at all” to the Langlands program, over the course of 200 years.
Maybe it sounds absurd for me to compare what one RLVR training can do, versus the whole edifice of ideas painstakingly built by the mathematics community over the course of 200 years. But it’s not absurd: AlphaZero really did blow past the whole edifice of ideas painstakingly built over the course of centuries by the chess and go communities in its 72-hour training runs. So the idea of building real new knowledge at a massive scale through RL is not absurd on its face. I’m just saying: RLVR-on-LLMs is not doing anything like that, at least not today. Rather, the LLM approach is to use imitative learning to suck in the whole edifice of ideas painstakingly built by humans, and then tweak it a bit at the end via RLVR.
Relatedly, when mathematicians study the recent LLM-generated math results, they’re not “learning something new” in a way that’s analogous to that teen spending years poring over her math textbooks. Rather they’re “learning something new” in a way that’s analogous to some guy telling me what his name is. If the mathematicians already have all the right background knowledge, they can quickly understand the solution within their existing conceptual frameworks. Or if they don’t already have all the right background knowledge, they can read human-created textbooks to get it. (I’m mainly thinking of the unit distance conjecture; in other examples that I looked into like the Jacobian conjecture, IIUC, the CoTs weren’t released, so mathematicians are still be a bit puzzled about how the LLMs came upon the answer.)
I don’t think it’s an argument against that, or sorry if I’m misunderstanding. You’re saying that LLMs may be computing their outputs in a different way from humans. Fine. But their outputs are still following the human distribution of outputs. “Following the human distribution of outputs” is just another way to say “low perplexity”, right?
(I will concede that the phrase “the human distribution of outputs” is a bit misleading on various other grounds, e.g. LLMs may generalize OOD in a different way from humans; not all imitative learning training tokens are created by humans; etc.)