I think this is accurate, though it varies between models (with Gemini being much less contrarian than Claude, for instance). A collaborator and I have been testing most models released in 2026 on a single-turn sycophancy evaluation, assessing how much a model will adapt its answer to agree with what a user seems to believe. This work suggests that they have become much more robust to user bias expressed in the first prompt.
However, my sense from the examples above and other experimentation is that this “progress” is rather superficial. It seems that the model is fundamentally optimizing for something like how to relate to the user rather than for what’s true, and has learnt not to do that too blatantly. But it emerges more in multi-turn interactions or in tests like this where one requests an opinion rather than a guess about a fact.
I think this is accurate, though it varies between models (with Gemini being much less contrarian than Claude, for instance). A collaborator and I have been testing most models released in 2026 on a single-turn sycophancy evaluation, assessing how much a model will adapt its answer to agree with what a user seems to believe. This work suggests that they have become much more robust to user bias expressed in the first prompt.
However, my sense from the examples above and other experimentation is that this “progress” is rather superficial. It seems that the model is fundamentally optimizing for something like how to relate to the user rather than for what’s true, and has learnt not to do that too blatantly. But it emerges more in multi-turn interactions or in tests like this where one requests an opinion rather than a guess about a fact.