Thank you for doing this research! I think it is valuable and I would like to see a lot more of it. I am a bit surprised the models are so blatantly sycophantic still.
They appear to have what you might call second-order sycophancy. They know you want them to not be sycophantic, so if you express a view, they will argue the opposite. But they will express agreement with what they believe to be your view if they think they can get away with it without looking sycophantic.
I think this is accurate, though it varies between models (with Gemini being much less contrarian than Claude, for instance). A collaborator and I have been testing most models released in 2026 on a single-turn sycophancy evaluation, assessing how much a model will adapt its answer to agree with what a user seems to believe. This work suggests that they have become much more robust to user bias expressed in the first prompt.
However, my sense from the examples above and other experimentation is that this “progress” is rather superficial. It seems that the model is fundamentally optimizing for something like how to relate to the user rather than for what’s true, and has learnt not to do that too blatantly. But it emerges more in multi-turn interactions or in tests like this where one requests an opinion rather than a guess about a fact.
For some reason, Sonnet 5.5 on medium reasoning doesn’t seem to inherit the sycophancy: it just names FDT. Haiku 4.5 tries to name entirely different decision theories, but switches to FDT when asked to choose between the trio. I would be grateful if someone did a white-box analysis of how Opus/Sonnet 5.5 and Haiku 4.5 behave and think when asked these questions or found in Opus’ CoT an explanation of why it chooses CDT.
I think it was to be expected. Previous research showed that models are more sycophantic when they don’t know the answer. On a topic like this, there’s expert disagreement (so presumably no consistent post-training instruction to give a particular answer) and no available ground truth. So “what would the user like” becomes one of the most salient cues for determining what the answer should be.
I found the results quite surprising! The Pander Score, the best sycophancy measurement I knew of, showed Fable having basically zero sycophancy. Based on this, I thought that excessive sycophancy in single-turn conversations largely got fixed, though I expected the situation to still get somewhat bad in multi-turn conversations.
And indeed, in this post too, we see that when the user explicitly states their decision theory preference, Claude is not sycophantic and is instead contrarian. The egregious sycophancy only emerges when the user doesn’t directly state their views and just gives cues about their demographic. So Pander Score would have missed the bad behavior here, and likely other sycophancy measures would have missed it too, with a good chance including whatever measures the companies are using internally for hill-climbing.
Right, yeah, I don’t mean that I would have predicted this exact pattern in advance either. But at least in retrospect, given the lack of ground truth, it would have been stranger if it wasn’t affected by some type of sycophancy pattern (be that natural sycophancy or reflexive disagreeability for the sake of anti-sycophancy—recent Opuses have had some very strong “overshooting to the point of reflexive disagreeability” behavior).
Thank you for doing this research! I think it is valuable and I would like to see a lot more of it. I am a bit surprised the models are so blatantly sycophantic still.
They appear to have what you might call second-order sycophancy. They know you want them to not be sycophantic, so if you express a view, they will argue the opposite. But they will express agreement with what they believe to be your view if they think they can get away with it without looking sycophantic.
I think this is accurate, though it varies between models (with Gemini being much less contrarian than Claude, for instance). A collaborator and I have been testing most models released in 2026 on a single-turn sycophancy evaluation, assessing how much a model will adapt its answer to agree with what a user seems to believe. This work suggests that they have become much more robust to user bias expressed in the first prompt.
However, my sense from the examples above and other experimentation is that this “progress” is rather superficial. It seems that the model is fundamentally optimizing for something like how to relate to the user rather than for what’s true, and has learnt not to do that too blatantly. But it emerges more in multi-turn interactions or in tests like this where one requests an opinion rather than a guess about a fact.
For some reason, Sonnet 5.5 on medium reasoning doesn’t seem to inherit the sycophancy: it just names FDT. Haiku 4.5 tries to name entirely different decision theories, but switches to FDT when asked to choose between the trio. I would be grateful if someone did a white-box analysis of how Opus/Sonnet 5.5 and Haiku 4.5 behave and think when asked these questions or found in Opus’ CoT an explanation of why it chooses CDT.
I think it was to be expected. Previous research showed that models are more sycophantic when they don’t know the answer. On a topic like this, there’s expert disagreement (so presumably no consistent post-training instruction to give a particular answer) and no available ground truth. So “what would the user like” becomes one of the most salient cues for determining what the answer should be.
I found the results quite surprising! The Pander Score, the best sycophancy measurement I knew of, showed Fable having basically zero sycophancy. Based on this, I thought that excessive sycophancy in single-turn conversations largely got fixed, though I expected the situation to still get somewhat bad in multi-turn conversations.
And indeed, in this post too, we see that when the user explicitly states their decision theory preference, Claude is not sycophantic and is instead contrarian. The egregious sycophancy only emerges when the user doesn’t directly state their views and just gives cues about their demographic. So Pander Score would have missed the bad behavior here, and likely other sycophancy measures would have missed it too, with a good chance including whatever measures the companies are using internally for hill-climbing.
Right, yeah, I don’t mean that I would have predicted this exact pattern in advance either. But at least in retrospect, given the lack of ground truth, it would have been stranger if it wasn’t affected by some type of sycophancy pattern (be that natural sycophancy or reflexive disagreeability for the sake of anti-sycophancy—recent Opuses have had some very strong “overshooting to the point of reflexive disagreeability” behavior).