Claude’s performance is low on the 2-4 hour range, which mostly consists of cybersecurity tasks, potentially dual-use for safety. In general, training on cybersecurity CTFs and ML code would increase “horizon length” on the METR plot, which only has 14 samples in the relevant (1 − 4hr) range where progress happened in 2025.
If I’m intrepreting these charts correctly, there is a decent amount of progress in 15m-1hr bracket, a small amount of progress in the 1hr bracket, a decent amount of progress in the 2-4 hr bracket, and a large amount of progress in the 4-16hr bracket. It doesn’t look like progress was dominated by the 1-4hr range.
The mean estimate of 50% success horizon length (headline number METR reports) went from ~1 to ~4 hours. The progress within the hour subranges is difficult to draw much information from, given the low number of data points, and distribution biases in topics. This is the precise claim of the new post I made, and linked :)
I have read your link and I understand you seem to have much more knowledge about statistics than I am. Perhaps I’m making a simple mistake somewhere in my reasoning. My thoughts are:
50 success horizon went from 1 to 4 hours. I interpreted your comment and post to posit that the increase is to some large degree due to training on cybersecurity to increase preformance in the 2-4hr range. However, looking at the charts it seems clear that the increase is not dominated by increased preformance in the 2-4hr range.
My point isn’t that there’s enough samples or that it isn’t possible to game the METR test, it’s that the measured improvement for Claude Opus 4.5 doesn’t seem to be primarly in the 2-4 hr bracket, if anything it’s dominated by 4-16hr.
I personally think the stronger argument here is that Claude models are not growing in capability consistent with higher task length = harder. (Grok 4 was similar) if you look at the histograms.
Both Sonnet 4.5 and Opus 4.5 were outperforming in the 8 to 16 hour bracket over the 2 to 4 hour, which is highly inconsistent with the task length difficulty model. The model appears broken at last since 3.5 sonnet given the flatness of the 2-16 hour tasks.
You end up in a case where the 4.5 Sonnet curve has a higher % of the solved tasks under it than 4.5 Opus (note how 4.5 Opus gets 0 tasks right in the 16 hour to 32 hour window even though the distribution implies it should be more like 25%). That is the “gain” this implies is overstated dramatically. [1]
The unfortunate consequence is largely shash42′s point—it’s not clear that modeling “task length horizon” is a valid way to view this data. Raw accuracy seems better correlated with time.
[1] An alternative interpretation is that Sonnet 4.5 was much better than the METR curve then implied.
https://www.lesswrong.com/posts/2RwDgMXo6nh42egoC/how-to-game-the-metr-plot
Claude’s performance is low on the 2-4 hour range, which mostly consists of cybersecurity tasks, potentially dual-use for safety. In general, training on cybersecurity CTFs and ML code would increase “horizon length” on the METR plot, which only has 14 samples in the relevant (1 − 4hr) range where progress happened in 2025.
If I’m intrepreting these charts correctly, there is a decent amount of progress in 15m-1hr bracket, a small amount of progress in the 1hr bracket, a decent amount of progress in the 2-4 hr bracket, and a large amount of progress in the 4-16hr bracket. It doesn’t look like progress was dominated by the 1-4hr range.
The mean estimate of 50% success horizon length (headline number METR reports) went from ~1 to ~4 hours. The progress within the hour subranges is difficult to draw much information from, given the low number of data points, and distribution biases in topics. This is the precise claim of the new post I made, and linked :)
I have read your link and I understand you seem to have much more knowledge about statistics than I am. Perhaps I’m making a simple mistake somewhere in my reasoning. My thoughts are:
50 success horizon went from 1 to 4 hours. I interpreted your comment and post to posit that the increase is to some large degree due to training on cybersecurity to increase preformance in the 2-4hr range. However, looking at the charts it seems clear that the increase is not dominated by increased preformance in the 2-4hr range.
My point isn’t that there’s enough samples or that it isn’t possible to game the METR test, it’s that the measured improvement for Claude Opus 4.5 doesn’t seem to be primarly in the 2-4 hr bracket, if anything it’s dominated by 4-16hr.
I personally think the stronger argument here is that Claude models are not growing in capability consistent with higher task length = harder. (Grok 4 was similar) if you look at the histograms.
Both Sonnet 4.5 and Opus 4.5 were outperforming in the 8 to 16 hour bracket over the 2 to 4 hour, which is highly inconsistent with the task length difficulty model. The model appears broken at last since 3.5 sonnet given the flatness of the 2-16 hour tasks.
You end up in a case where the 4.5 Sonnet curve has a higher % of the solved tasks under it than 4.5 Opus (note how 4.5 Opus gets 0 tasks right in the 16 hour to 32 hour window even though the distribution implies it should be more like 25%). That is the “gain” this implies is overstated dramatically. [1]
The unfortunate consequence is largely shash42′s point—it’s not clear that modeling “task length horizon” is a valid way to view this data. Raw accuracy seems better correlated with time.
[1] An alternative interpretation is that Sonnet 4.5 was much better than the METR curve then implied.