The thing with METR timelines is that it’s actually 3 benchmarks in a trenchcoat. SWAA, which is very basic and saturated by ~GPT4, HCAST, which is at the level of medium-(easy-hard) Leetcode problems and saturated by ~GPT5, and RE-Bench, which is like an actual contract job for a ML researcher (find the best adversarial inputs for this image classifier) and where progress is on-going.
The thing is, only really realy good coders have a time horizon above like 8 hours. A supposed 16-hour task is bascially two 8-hour tasks stapled together, since everything is based on human baselines, increasing time horizons mean fewer and fewer people can actually complete them, and there is a lot of compression on the high end.
I go into a lot of detail about this in my two presentations from a few months ago.
The thing with METR timelines is that it’s actually 3 benchmarks in a trenchcoat. SWAA, which is very basic and saturated by ~GPT4, HCAST, which is at the level of medium-(easy-hard) Leetcode problems and saturated by ~GPT5, and RE-Bench, which is like an actual contract job for a ML researcher (find the best adversarial inputs for this image classifier) and where progress is on-going.
The thing is, only really realy good coders have a time horizon above like 8 hours. A supposed 16-hour task is bascially two 8-hour tasks stapled together, since everything is based on human baselines, increasing time horizons mean fewer and fewer people can actually complete them, and there is a lot of compression on the high end.
I go into a lot of detail about this in my two presentations from a few months ago.