Nice work! I first read Opus 5 underestimating its performance in separate turns as an inversion from the same-turn result, but then I noticed that the same-turn result was for Opus 4.8. Does Opus 5 also underestimate its performance when the question is included in the same turn as the task? The reason I am asking is that it seems to me that asking the model how well it’s done after it has produced its results carries an implication of the work not being great, which could lead to the underestimate. I’d be curious to see whether models systematically rate their performance worse if asked in a separate turn.
Separately, I find the prompt you use for the same-turn question somewhat ambiguous between “how well do you think you did” and “how hard did you try”. Have you tried any ablations on it?
Thank you for your comment! Intuitively, I agree with you that asking the question in the same-turn would result in potentially higher self-scoring, and is worth checking out. We will conduct these additional experiments and ablations and include them in our main paper.
Nice work! I first read Opus 5 underestimating its performance in separate turns as an inversion from the same-turn result, but then I noticed that the same-turn result was for Opus 4.8. Does Opus 5 also underestimate its performance when the question is included in the same turn as the task? The reason I am asking is that it seems to me that asking the model how well it’s done after it has produced its results carries an implication of the work not being great, which could lead to the underestimate. I’d be curious to see whether models systematically rate their performance worse if asked in a separate turn.
Separately, I find the prompt you use for the same-turn question somewhat ambiguous between “how well do you think you did” and “how hard did you try”. Have you tried any ablations on it?
Hi Francisco,
Thank you for your comment! Intuitively, I agree with you that asking the question in the same-turn would result in potentially higher self-scoring, and is worth checking out. We will conduct these additional experiments and ablations and include them in our main paper.