Caveat: Just as in the previous section, this paper is old (by RLVR standards), and not based on bleeding-edge frontier LLMs with bleeding-edge RLVR best practices. I guess in this case I’m slightly reassured by the fact that one of the coauthors, @Neel Nanda, works with near-SOTA LLMs at his day job at DeepMind. But only slightly.
My guess is that the paper was true of early RLVR models especially distills, remains a useful mental model, but probably is explaining less and less of the picture over time
I think RL gives way fewer bits of information per token, but does give some, and if people are eg spending comparable amounts of compute on RL as on pretraining that’s quite a lot
My guess is that the paper was true of early RLVR models especially distills, remains a useful mental model, but probably is explaining less and less of the picture over time
I think RL gives way fewer bits of information per token, but does give some, and if people are eg spending comparable amounts of compute on RL as on pretraining that’s quite a lot