Fantastic work! Updates me on a lot. I expect the automated grader prompt may be making the model play all the evals more as games rather than taking them seriously as evals. It’d be great to see the automated grader prompt but saying that it has real world impacts on people. Not to say that the persona explanation is wrong, but I currently put more weight that this may be caused by the automated grader implying there are no real stakes, so it plays games and aces tests that it knows are fake.
Fantastic work! Updates me on a lot. I expect the automated grader prompt may be making the model play all the evals more as games rather than taking them seriously as evals. It’d be great to see the automated grader prompt but saying that it has real world impacts on people. Not to say that the persona explanation is wrong, but I currently put more weight that this may be caused by the automated grader implying there are no real stakes, so it plays games and aces tests that it knows are fake.