Caveat: Just as in the previous section, this paper is old (by RLVR standards), and not based on bleeding-edge frontier LLMs with bleeding-edge RLVR best practices. I guess in this case I’m slightly reassured by the fact that one of the coauthors, @Neel Nanda, works with near-SOTA LLMs at his day job at DeepMind. But only slightly.
My guess is that the paper was true of early RLVR models especially distills, remains a useful mental model, but probably is explaining less and less of the picture over time
I think RL gives way fewer bits of information per token, but does give some, and if people are eg spending comparable amounts of compute on RL as on pretraining that’s quite a lot
I like the premise of this post! Can you share your full strategic deception evals? I’m concerned that the examples given in the post seem way too transparently fake, a capable model ought to be able to tell that its fake, and may infer that either the user wants it to say yes, or that its being invited to roleplay. But it’s hard to assess because you’re comparing pretty dumb models to the highly capable and likely distilled ones. Further I suspect intelligent models that are not distilled will still respond similarly to these prompts, and would probably also respond to similar prompts about being Gemini or GPT