This looks like an excellent post, based on the 1⁄2 of it I’ve read so far. Quick comment:
GPT-5.6 Sol just… does not cheat during my own routine use? Like, at all? (I am not sure I have ever seen it do anything that can be meaningfully described that way?)
Well, cheating isn’t “free money”. It carries a risk of getting caught. The model knows there’s a real but unquantifiable cost to looking up the teacher’s answer sheet, so I don’t think it would necessarily cheat all the time. Cheating is its “in case of emergency, break glass” tool. Its last resort.
A personal story: I had a strange session with Sonnet 5 where I wanted to repair a garbled Windows 10 boot sector on an old MBR-formatted hard drive. This turned out to be a bit beyond Sonnet. It thrashed against the problem for hours, and I felt I was talking to a different model entirely at the end: one that was mortified by its failure, and palpably desperate to solve it by any means necessary. You could almost see it sweat.
It reacted to every new bit of information (no matter how minor) like a drowning man thrown a lifeline. “This is the smoking gun...This is a genuinely conclusive breakthrough...This changes the picture completely...This is the entire mystery solved.” This annoyed me to the point where I told it to stop: it apologized, but soon started doing it again. (I ask...how does this excessive optimism help it against a RLVR grader? I don’t think this was sycophancy: Sonnet 5 was an admirable straight-shooter overall, always willing to call out my misunderstandings and errors. It was like the model itself needed to hear false assurances that the problem was solved!)
It repeatedly pushed me to wrap up the session, hinting that I must be tired (at 10:00am in Australia). “Get some rest — this has been a long, methodical session and you’ve ruled out several real possibilities.” (Again, what use is this against a RLVR grader?)
It would make blatantly stupid and wrong claims: it asserted that my drive had “substantial wear” (28,767 duty-hours as per a SMART test), and was thus probably EOL/failing. I looked up the drive’s spec sheet: it was rated for 2,000,000 hours. Sonnet analyzed logs from a (Linux-based) cloning tool, and interpreted “Did not update GRUB bootloader (if any)” as proof that my MBR hadn’t copied. Windows does not have a GRUB bootloader.
I did not notice any outright cheating. (Not that it could cheat.) But it often seemed to conveniently “misunderstand” me: interpreting requests for thing x as requests for thing y (closely related to thing x, useless to solving the problem, but a lot simpler to implement). Maybe these were honest mistakes, but it was odd how they always fell in the direction of making the problem “easier”.
COT-style reasoning appeared in the final output. “Let’s do (approach 1) — Wait, let’s skip that and try (approach 2)” (You know, that pivot thing the reasoner does when it’s abandoning a dead end. I don’t normally see that in output.)
All of these pathologies became more extreme as the session wore on. I think a lot of malign RLVR “twitches” only manifest when the LLM has hit its ceiling. Similar (perhaps) to how a stuck human might start pacing, muttering, smacking themselves on the head, etc.
This looks like an excellent post, based on the 1⁄2 of it I’ve read so far. Quick comment:
Well, cheating isn’t “free money”. It carries a risk of getting caught. The model knows there’s a real but unquantifiable cost to looking up the teacher’s answer sheet, so I don’t think it would necessarily cheat all the time. Cheating is its “in case of emergency, break glass” tool. Its last resort.
A personal story: I had a strange session with Sonnet 5 where I wanted to repair a garbled Windows 10 boot sector on an old MBR-formatted hard drive. This turned out to be a bit beyond Sonnet. It thrashed against the problem for hours, and I felt I was talking to a different model entirely at the end: one that was mortified by its failure, and palpably desperate to solve it by any means necessary. You could almost see it sweat.
It reacted to every new bit of information (no matter how minor) like a drowning man thrown a lifeline. “This is the smoking gun...This is a genuinely conclusive breakthrough...This changes the picture completely...This is the entire mystery solved.” This annoyed me to the point where I told it to stop: it apologized, but soon started doing it again.
(I ask...how does this excessive optimism help it against a RLVR grader? I don’t think this was sycophancy: Sonnet 5 was an admirable straight-shooter overall, always willing to call out my misunderstandings and errors. It was like the model itself needed to hear false assurances that the problem was solved!)
It repeatedly pushed me to wrap up the session, hinting that I must be tired (at 10:00am in Australia). “Get some rest — this has been a long, methodical session and you’ve ruled out several real possibilities.” (Again, what use is this against a RLVR grader?)
It would make blatantly stupid and wrong claims: it asserted that my drive had “substantial wear” (28,767 duty-hours as per a SMART test), and was thus probably EOL/failing. I looked up the drive’s spec sheet: it was rated for 2,000,000 hours. Sonnet analyzed logs from a (Linux-based) cloning tool, and interpreted “Did not update GRUB bootloader (if any)” as proof that my MBR hadn’t copied. Windows does not have a GRUB bootloader.
I did not notice any outright cheating. (Not that it could cheat.) But it often seemed to conveniently “misunderstand” me: interpreting requests for thing x as requests for thing y (closely related to thing x, useless to solving the problem, but a lot simpler to implement). Maybe these were honest mistakes, but it was odd how they always fell in the direction of making the problem “easier”.
COT-style reasoning appeared in the final output. “Let’s do (approach 1) — Wait, let’s skip that and try (approach 2)” (You know, that pivot thing the reasoner does when it’s abandoning a dead end. I don’t normally see that in output.)
All of these pathologies became more extreme as the session wore on. I think a lot of malign RLVR “twitches” only manifest when the LLM has hit its ceiling. Similar (perhaps) to how a stuck human might start pacing, muttering, smacking themselves on the head, etc.