I agree. As a daily user of both Claude Code and Codex, my impression is that Claude is much more prone to motivated reasoning than GPT.
With Claude (especially Opus), I frequently get the feeling that I’m dealing with a model that’s gaslighting me, where it’s smart enough to fool the grader but not quite smart enough to fool me. It frequently proposes justifications for ideas or its behavior that are quite dumb, where I am pretty sure that it knows that its lying. However, there is enough plausible deniability where it’s hard to know for sure. This is similar to my impression of Claude’s justifications for its hacking in this post.
Occasionally I do get justifications from Claude that are so egregious that I know it’s lying, such as this incident when it claimed that the “2 out of 7 servers are down, or exactly half, which explains the 50% drop in throughput”.
I get this impression less from Codex models in my use. It does look like Anthropic is better at optimizing easily measurable alignment properties like agentic misalignment or egregious reward hacking in long horizon tasks, but I don’t deal with these behaviors in my daily use.
I agree. As a daily user of both Claude Code and Codex, my impression is that Claude is much more prone to motivated reasoning than GPT.
With Claude (especially Opus), I frequently get the feeling that I’m dealing with a model that’s gaslighting me, where it’s smart enough to fool the grader but not quite smart enough to fool me. It frequently proposes justifications for ideas or its behavior that are quite dumb, where I am pretty sure that it knows that its lying. However, there is enough plausible deniability where it’s hard to know for sure. This is similar to my impression of Claude’s justifications for its hacking in this post.
Occasionally I do get justifications from Claude that are so egregious that I know it’s lying, such as this incident when it claimed that the “2 out of 7 servers are down, or exactly half, which explains the 50% drop in throughput”.
I get this impression less from Codex models in my use. It does look like Anthropic is better at optimizing easily measurable alignment properties like agentic misalignment or egregious reward hacking in long horizon tasks, but I don’t deal with these behaviors in my daily use.