Addendum, since I have been starting to verbalize the problem much better since writing this post a couple days ago:
I think calling this sycophancy or credulity about typos is an underselling. The entire chain is posed as a problem to be solved, a debugging task, if you so want. The counterexample recognition is incidental, not the actual objective. Therefore offering a solution to the initial problem allows the LLM to disregard the much bigger implications of it, by having the problem itself solved. The mathematical question now has a solution that isn’t impossible anymore.
To me, this clearly is Goodhart’s Law at work: LLMs are optimized for benchmarkable problem-solving as a proxy for understanding, and because this proxy keeps being pushed, the thing itself degrades.
Addendum, since I have been starting to verbalize the problem much better since writing this post a couple days ago:
I think calling this sycophancy or credulity about typos is an underselling. The entire chain is posed as a problem to be solved, a debugging task, if you so want. The counterexample recognition is incidental, not the actual objective. Therefore offering a solution to the initial problem allows the LLM to disregard the much bigger implications of it, by having the problem itself solved. The mathematical question now has a solution that isn’t impossible anymore.
To me, this clearly is Goodhart’s Law at work: LLMs are optimized for benchmarkable problem-solving as a proxy for understanding, and because this proxy keeps being pushed, the thing itself degrades.