I don’t think that the METR graph had the impact which you describe for the following reasons:
It was released on March 19, 2025 and described the effect of scaling and RLVR. It didn’t affect anything from Claude 3 Opus to Claude 3.7 Sonnet since they were released beforehand.
All known scaling laws are, well, laws connecting capabilities with scaling of parameters like the LLM’s size, the dataset’s size or, in METR’s case, with the amount of compute and discoveries and/or time that went into teaching the model to do tasks. Scaling laws don’t depend on anyone’s will, making them closer to laws of physics than to Moore’s law.
It affected AI-2027 by providing them with the ability to extrapolate the horizons and try to predict when we’ll reach the Potentially Hazardous Horizon where Coding Is Automated.
I don’t think that the METR graph had the impact which you describe for the following reasons:
It was released on March 19, 2025 and described the effect of scaling and RLVR. It didn’t affect anything from Claude 3 Opus to Claude 3.7 Sonnet since they were released beforehand.
All known scaling laws are, well, laws connecting capabilities with scaling of parameters like the LLM’s size, the dataset’s size or, in METR’s case, with the amount of compute and discoveries and/or time that went into teaching the model to do tasks. Scaling laws don’t depend on anyone’s will, making them closer to laws of physics than to Moore’s law.
It affected AI-2027 by providing them with the ability to extrapolate the horizons and try to predict when we’ll reach the Potentially Hazardous Horizon where Coding Is Automated.
Additionally, METR recently released an entire risk report trying to rule out rogue deployments inside or outside the labs and incorporating lots of evals like CoTless math and ARC-AGI-3(!). It also did express opinions like the idea that GPT-5.1′s performance means that “if trends hold, further development would pose low risk for these threat models, based on an aggressive extrapolation of our time horizon metric over the next 6 months” or even claims that GPT-5.6 Sol is such a cheater that METR failed to describe its time horizon and demanded deep access to internal systems.