Their personal groupthought about whether the models are getting better or worse at instruction following, inside of a community of people with a staked position on the subject, inside or outside of the labs, is exactly the kind of thing that needs really high quality data to stay sane!
I think “better or worse at instruction following” is too general and not the claim at stake. What is the “staked position” of labs or Cursor for example?
I think “better or worse at instruction following” is too general and not the claim at stake. What is the “staked position” of labs or Cursor for example?