Claude 4 does have an extended thinking mode, and many of the Claude 4 benchmark results in the screenshot were obtained with extended thinking.
From here, in the “Performance benchmark reporting” appendix:
Claude Opus 4 and Sonnet 4 are hybrid reasoning models. The benchmarks reported in this blog post show the highest scores achieved with or without extended thinking. We’ve noted below for each result whether extended thinking was used:
No extended thinking: SWE-bench Verified, Terminal-bench
Extended thinking (up to 64K tokens):
TAU-bench (no results w/o extended thinking reported)
GPQA Diamond (w/o extended thinking: Opus 4 scores 74.9% and Sonnet 4 is 70.0%)
MMMLU (w/o extended thinking: Opus 4 scores 87.4% and Sonnet 4 is 85.4%)
MMMU (w/o extended thinking: Opus 4 scores 73.7% and Sonnet 4 is 72.6%)
AIME (w/o extended thinking: Opus 4 scores 33.9% and Sonnet 4 is 33.1%)
On MMMU, if I’m reading things correctly, the relative order of the models was:
Opus 4 > Sonnet 4 > Sonnet 3.7 in the no-extended-thinking case
Opus 4 > Sonnet 3.7 > Sonnet 4 in the extended-thinking case
Claude 4 does have an extended thinking mode, and many of the Claude 4 benchmark results in the screenshot were obtained with extended thinking.
From here, in the “Performance benchmark reporting” appendix:
On MMMU, if I’m reading things correctly, the relative order of the models was:
Opus 4 > Sonnet 4 > Sonnet 3.7 in the no-extended-thinking case
Opus 4 > Sonnet 3.7 > Sonnet 4 in the extended-thinking case