Not on personal experience with the models. I have experienced reward hacking too (as have most people, I think), and I’ve reproduced a lot of these failures in controlled settings where I’ve carefully analyzed the prompt. I’m just explaining why I am generally skeptical of the quality of the evidence that he’s citing.
I really disagree! People (like me) are bad at reading comprehension. Their personal groupthought about whether the models are getting better or worse at instruction following, inside of a community of people with a staked position on the subject, inside or outside of the labs, is exactly the kind of thing that needs really high quality data to stay sane! On twitter, people can sometimes hardly agree whether or not the new model is better than the last version, amid tons of benchmarks!
Their personal groupthought about whether the models are getting better or worse at instruction following, inside of a community of people with a staked position on the subject, inside or outside of the labs, is exactly the kind of thing that needs really high quality data to stay sane!
I think “better or worse at instruction following” is too general and not the claim at stake. What is the “staked position” of labs or Cursor for example?
Oh, I have a guess about where the miscommunication is. I took Linch’s “I do not have a clear understanding of why this occurs.” sentence to refer to the phenomenon in general, not to the headline claim of unreleased models being substantially more reward-hacky/locally deceptive/etc than released models. And I don’t need hard data to personally observe that the released models are noticeably reward-hacky, because I have personally noticed.
I took Linch’s “I do not have a clear understanding of why this occurs.” sentence to refer to the phenomenon in general, not to the headline claim of unreleased models being substantially more reward-hacky/locally deceptive/etc than released models.
Huh. It turns out the same seemingly simple sentence can have 3 different interpretations, without my intending it.
In this case I meant what you said partially, but part of the implicit model in my head for why something needed to be explained isn’t just the phenomenon of reward-hacking in general against an undefined baseline but the (probable-but-uncertain) hypothesis that the degree of reward-hacking has grown over time. This is why I reached for RLVR and new model sizes as the most likely salient potential explanations, rather than eg. Goodhart’s Law or pretraining data stuff.
Just going to delete this and reply to the initial post.
Linch based his statement on:
Not on personal experience with the models. I have experienced reward hacking too (as have most people, I think), and I’ve reproduced a lot of these failures in controlled settings where I’ve carefully analyzed the prompt. I’m just explaining why I am generally skeptical of the quality of the evidence that he’s citing.
Oh, I see. I guess that’s reasonable, but, eh, I feel like it can mostly be screened off by the personal experience?
I really disagree! People (like me) are bad at reading comprehension. Their personal groupthought about whether the models are getting better or worse at instruction following, inside of a community of people with a staked position on the subject, inside or outside of the labs, is exactly the kind of thing that needs really high quality data to stay sane! On twitter, people can sometimes hardly agree whether or not the new model is better than the last version, amid tons of benchmarks!
I think “better or worse at instruction following” is too general and not the claim at stake. What is the “staked position” of labs or Cursor for example?
Oh, I have a guess about where the miscommunication is. I took Linch’s “I do not have a clear understanding of why this occurs.” sentence to refer to the phenomenon in general, not to the headline claim of unreleased models being substantially more reward-hacky/locally deceptive/etc than released models. And I don’t need hard data to personally observe that the released models are noticeably reward-hacky, because I have personally noticed.
Huh. It turns out the same seemingly simple sentence can have 3 different interpretations, without my intending it.
In this case I meant what you said partially, but part of the implicit model in my head for why something needed to be explained isn’t just the phenomenon of reward-hacking in general against an undefined baseline but the (probable-but-uncertain) hypothesis that the degree of reward-hacking has grown over time. This is why I reached for RLVR and new model sizes as the most likely salient potential explanations, rather than eg. Goodhart’s Law or pretraining data stuff.
Naw I got that now, that’s why I redacted the original post.