Apparently Claude 4 Opus hasn’t gotten further in Pokémon Red than Claude 3.7 Sonnet did. Just three gym badges obtained.
This is a surprising failure, and probably why Anthropic hasn’t released an updated version of their Pokémon benchmark. Instead, they just made some comments about improved long-term planning ability, which apparently doesn’t translate into measurable results?
That really is surprising, especially given that the announcement includes the following:
Claude Opus 4 also dramatically outperforms all previous models on memory capabilities. When developers build applications that provide Claude local file access, Opus 4 becomes skilled at creating and maintaining ‘memory files’ to store key information. This unlocks better long-term task awareness, coherence, and performance on agent tasks—like Opus 4 creating a ‘Navigation Guide’ while playing Pokémon.
I’d pretty much assumed they fine-tuned on the Pokémon task and it was going to be just about one-shotting the game. Weird.
Yeah I feel like they came up with something nice to say while eliding the “no further progress” issue.
Weirdly, while the announcement talks about creating and maintaining multiple “memory files”, the new public ClaudePlaysPokemon stream has Claude Opus 4 using just a single memory file which it doesn’t even create. Apparently this is “much better” than the setup Claude 3.7 Sonnet used which let it create and maintain as many files at it wanted (usually to its detriment).
One other interesting tidbit I’ll throw in from the stream:
claudestans: @ClaudePlaysPokemon you mentioned before that there was a personality change from e.g. 3.5 sonnet to 3.7 (more persistent, less giving up etc). Have you noticed anything about opus 4 in terms of personality?
ClaudePlaysPokemon: Opus is so much better at keeping track of things that it gets more distressed when it can’t figure things out! So I need to convince it more that nothing is wrong, which I find quite interesting! ClaudePlaysPokemon: like it will be very aware that it has taken 100 steps to solve something and it finds that very frustrating
Huh, indeed interesting. IIRC, one of the suggested problems with the previous models was the lack of boredom: that they’re perfectly capable of doing something in a loop forever where a human would’ve gotten frustrated and done something random that might’ve ended up helping. Sounds like Opus 4 is different in that regard...?
That’s a big performance limitation for LLMs for sure, but Claude 3.7 Sonnet got two more badges than Claude 3.6 Sonnet. Pure reasoning improvements have led to more badges in the past.
I assume you mean MMMU? Looks like a 70.4% → 75% score improvement on the benchmark last jump, compared to just a 75% → 76.5% score improvement this time. I don’t think that’s a big difference, but I was wrong to say the improvement was “pure” reasoning improvements, my bad.
Indeed, it seems Claude 4 is not that much better than Claude 3.7 sonnet at visual reasoning, in fact Claude-4 sonnet is worse at visual reasoning than Claude 3.7 sonnet (though Claude 4 Opus is better). So extra not so surprising, and an indication that this isn’t really what Anthropic is focusing on at the moment (likely a good call).
Claude Sonnet 4 is still better than Claude 3.7 Sonnet without Extended Thinking. Given that 4 doesn’t seem to have an Extended Thinking mode, I’m not sure it’s really a performance degradation.
Claude 4 does have an extended thinking mode, and many of the Claude 4 benchmark results in the screenshot were obtained with extended thinking.
From here, in the “Performance benchmark reporting” appendix:
Claude Opus 4 and Sonnet 4 are hybrid reasoning models. The benchmarks reported in this blog post show the highest scores achieved with or without extended thinking. We’ve noted below for each result whether extended thinking was used:
No extended thinking: SWE-bench Verified, Terminal-bench
Extended thinking (up to 64K tokens):
TAU-bench (no results w/o extended thinking reported)
GPQA Diamond (w/o extended thinking: Opus 4 scores 74.9% and Sonnet 4 is 70.0%)
MMMLU (w/o extended thinking: Opus 4 scores 87.4% and Sonnet 4 is 85.4%)
MMMU (w/o extended thinking: Opus 4 scores 73.7% and Sonnet 4 is 72.6%)
AIME (w/o extended thinking: Opus 4 scores 33.9% and Sonnet 4 is 33.1%)
On MMMU, if I’m reading things correctly, the relative order of the models was:
Opus 4 > Sonnet 4 > Sonnet 3.7 in the no-extended-thinking case
Opus 4 > Sonnet 3.7 > Sonnet 4 in the extended-thinking case
Apparently Claude 4 Opus hasn’t gotten further in Pokémon Red than Claude 3.7 Sonnet did. Just three gym badges obtained.
This is a surprising failure, and probably why Anthropic hasn’t released an updated version of their Pokémon benchmark. Instead, they just made some comments about improved long-term planning ability, which apparently doesn’t translate into measurable results?
Source: Wired reporter Kylie Robinson asked about it in-person to ClaudePlaysPokemon developer and Anthropic employee David Hershey: https://x.com/kyliebytes/status/1925617856449757364
That really is surprising, especially given that the announcement includes the following:
I’d pretty much assumed they fine-tuned on the Pokémon task and it was going to be just about one-shotting the game. Weird.
I found that section very suspicious because it omitted any statement about actual performance, and I guess now we know why.
This seems in line with my longer timelines hypothesis. Perhaps it roughly undoes the update on AlphaEvolve, which I wasn’t sure how to interpret.
Of course the METR evaluation will contain more signal.
Yeah I feel like they came up with something nice to say while eliding the “no further progress” issue.
Weirdly, while the announcement talks about creating and maintaining multiple “memory files”, the new public ClaudePlaysPokemon stream has Claude Opus 4 using just a single memory file which it doesn’t even create. Apparently this is “much better” than the setup Claude 3.7 Sonnet used which let it create and maintain as many files at it wanted (usually to its detriment).
(source: this doc David Hershey just published on the new harness for the stream)
One other interesting tidbit I’ll throw in from the stream:
Huh, indeed interesting. IIRC, one of the suggested problems with the previous models was the lack of boredom: that they’re perfectly capable of doing something in a loop forever where a human would’ve gotten frustrated and done something random that might’ve ended up helping. Sounds like Opus 4 is different in that regard...?
I do still think the main effect here is bad vision & pixel-art being hard
That’s a big performance limitation for LLMs for sure, but Claude 3.7 Sonnet got two more badges than Claude 3.6 Sonnet. Pure reasoning improvements have led to more badges in the past.
Note there was a significant increase in claude’s visual reasoning ability (as measured by the MMMU) between claude-3.6 and thinking-claude-3.7
I assume you mean MMMU? Looks like a 70.4% → 75% score improvement on the benchmark last jump, compared to just a 75% → 76.5% score improvement this time. I don’t think that’s a big difference, but I was wrong to say the improvement was “pure” reasoning improvements, my bad.
Indeed, it seems Claude 4 is not that much better than Claude 3.7 sonnet at visual reasoning, in fact Claude-4 sonnet is worse at visual reasoning than Claude 3.7 sonnet (though Claude 4 Opus is better). So extra not so surprising, and an indication that this isn’t really what Anthropic is focusing on at the moment (likely a good call).
Claude Sonnet 4 is still better than Claude 3.7 Sonnet without Extended Thinking. Given that 4 doesn’t seem to have an Extended Thinking mode, I’m not sure it’s really a performance degradation.
Claude 4 does have an extended thinking mode, and many of the Claude 4 benchmark results in the screenshot were obtained with extended thinking.
From here, in the “Performance benchmark reporting” appendix:
On MMMU, if I’m reading things correctly, the relative order of the models was:
Opus 4 > Sonnet 4 > Sonnet 3.7 in the no-extended-thinking case
Opus 4 > Sonnet 3.7 > Sonnet 4 in the extended-thinking case