What’s especially impressive is it was solved by a model that’s only been in training for 11 days. (Plus tens of millions of dollars of inference.)
Additionally, I feel like how recent the model is lends at least some credence to the theory that Tristan Buckmaster and Levent Alpoge’s prompts/work made it into the training of the internal model, which I initially thought was unlikely.
tdo
It’s a little hard to buy that 11% of federal lobbyists in DC were AI-related the year GPT-4 released.
We flagged lobbying as related to AI issues if the disclosures mentioned at least one of the following items: 1) artificial intelligence, 2) autonomous vehicles and systems and/or 3) data centers. We also searched for variations of those words (A.I.) as well as a handful of other AI terms including large language model and facial recognition.
Ah. It was just a hit match and everyone started putting the word AI everywhere in 2023.
Note it was forecast for May 2030 exactly one year ago. It’s been fluctuating from 2030-2034 ever since GPT-4 was dropped almost exactly 3 years ago, with a few extended periods closer to both the high and low ends. I think it’s mostly noise.
METR discovered an issue with their task horizon modeling. Fixing the issue reduces some of the recent model scores by 10-20% on the 50% success horizon but also increases 80% horizon a bit. For example, Opus 4.6 dropped from 14.5 hours to 12 hours.
It’s not just your friend, some staff on METR themselves thinks the same. Not even just Opus 4.6, but “Similar deal for other recent models” too.
Likely just the result of noise but extraordinarily funny METR results today.
Opus 4.6 got nearly triple the time horizon of Opus 4.5, immediately followed by GPT-5.3 getting a worse time horizon than GPT-5.2.
.
.
METR has updated their task horizon methodology with more tasks which leads to different task horizon scores.
Claude Sonnet 4.5′s 50% task horizon is 1 hr 53 min, putting it slightly behind GPT-5′s 2 hr 15 min score.
Note that they also claimed Opus 4 worked for over seven hours but only scored 1h20m on the METR task suite. I wouldn’t be surprised to see Sonnet 4.5 get a strong METR score but 30 hours definitely isn’t likely.
Surprised that this hasn’t gotten more discussion. There’s some potentially big implications for the time horizons study, which has become fairly load-bearing in timelines discourse.
METR’s task-horizon score on GPT-5 is 2h17m @ 50% success. For comparison, o3 was 1h32m and Grok 4 (prior SOTA) was 1hr50m. The 80% success score is 25m, prior SOTA was 20m from both o3 and Claude 4 Opus.
https://metr.github.io/autonomy-evals-guide/gpt-5-report/
METR has finally tested Gemini 2.5 Pro (June Preview) and found its 50% success task horizon is only 39 minutes, far worse than o3 or Opus 4 which are at 90 and 80 minutes respectively. Probably shouldn’t be a gigantic update given 2.5 Pro never scored amazingly at SWE-Bench, but still worse than I expected given how good the model is otherwise.
I feel like looking at unreleased models for doubling time mucks things up a bit. For instance I’m assuming the unreleased o3 model from December had a significantly longer time-horizon in math than the released o3, given its much higher benchmarks in FrontierMath, etc.
Worth noting this year’s p3 was really easy, Gemini 2.5 pro even got it some of the time, and Grok 4 Heavy and Gemini Deep Think got problems rated as harder. Still an achievement, though.
From the author of the epoch article:
https://x.com/GregHBurnham/status/1946655635400950211
https://x.com/GregHBurnham/status/1946725960557949227
https://x.com/GregHBurnham/status/1946567312850530522
tdko’s Shortform
METR’s task length horizon analysis for Claude 4 Opus is out. The 50% task success chance is at 80 minutes, slightly worse than o3′s 90 minutes. The 80% task success chance is tied with o3 at 20 minutes.
OpenAI has confirmed to NYT that there is some truth to the rumor.
https://www.nytimes.com/2026/09/10/science/tristan-buckmaster-openai-math-navier-stokes.html