Do you mention Claude Opus 4.6′s 80.8% because your analysis has one free parameter and you set it to fit that 80.8%? How well does your analysis translate other model’s percentage scores to their time horizons?
I mention Opus 4.6 because it is the predecessor model and this allows a comparison between the numbers that pop out of my analysis and the “official” METR values.
My analysis at least recovered the exponential improvement of time horizons with similar doubling times as the METR analysis, but the concrete values depend on modelling assumptions.
If I find the time I might write it up after all, but here is a short sketch:
Two assumptions:
The logistics fitted by METR tend to have quite similar slopes (at least the later models), so I take the average slope for my fit.
The task time completions of SWE-bench verified are log-normally distributed, I derive the concrete distribution from commit timestamps by cleverly trying to correct for pauses. Here different modelling assumptions don’t change the trend but can change the time horizon values.
With the slope and the distribution I can find for each percentage the position of the logistic which gives me the time horizons.
Do you mention Claude Opus 4.6′s 80.8% because your analysis has one free parameter and you set it to fit that 80.8%? How well does your analysis translate other model’s percentage scores to their time horizons?
I mention Opus 4.6 because it is the predecessor model and this allows a comparison between the numbers that pop out of my analysis and the “official” METR values.
My analysis at least recovered the exponential improvement of time horizons with similar doubling times as the METR analysis, but the concrete values depend on modelling assumptions.
If I find the time I might write it up after all, but here is a short sketch:
Two assumptions:
The logistics fitted by METR tend to have quite similar slopes (at least the later models), so I take the average slope for my fit.
The task time completions of SWE-bench verified are log-normally distributed, I derive the concrete distribution from commit timestamps by cleverly trying to correct for pauses. Here different modelling assumptions don’t change the trend but can change the time horizon values.
With the slope and the distribution I can find for each percentage the position of the logistic which gives me the time horizons.