I took Linch’s “I do not have a clear understanding of why this occurs.” sentence to refer to the phenomenon in general, not to the headline claim of unreleased models being substantially more reward-hacky/locally deceptive/etc than released models.
Huh. It turns out the same seemingly simple sentence can have 3 different interpretations, without my intending it.
In this case I meant what you said partially, but part of the implicit model in my head for why something needed to be explained isn’t just the phenomenon of reward-hacking in general against an undefined baseline but the (probable-but-uncertain) hypothesis that the degree of reward-hacking has grown over time. This is why I reached for RLVR and new model sizes as the most likely salient potential explanations, rather than eg. Goodhart’s Law or pretraining data stuff.
Huh. It turns out the same seemingly simple sentence can have 3 different interpretations, without my intending it.
In this case I meant what you said partially, but part of the implicit model in my head for why something needed to be explained isn’t just the phenomenon of reward-hacking in general against an undefined baseline but the (probable-but-uncertain) hypothesis that the degree of reward-hacking has grown over time. This is why I reached for RLVR and new model sizes as the most likely salient potential explanations, rather than eg. Goodhart’s Law or pretraining data stuff.