It appears there were errors with the conversion of footnotes from Google Docs to LessWrong. Corrected Footnotes: [2] The weightings are, roughly, based on the following reasoning (some of these ideas are repeated elsewhere in this post):
In general, we should care more about higher capability levels.
There is a selection effect where people don’t bother making or testing really weak models anymore, so the data omit recent low-capability models, so we down-weight lower capabilities.
There is also a lack of data on early small models. For instance, GPT-4 is the first model to exceed 20 on the Intelligence Index, and it also appears in this data to be the first model to exceed 5, 10, and 15. But it isn’t actually the first model to exceed 5 (and we don’t know for 10 and 15). Earlier models either are not evaluated by Artificial Analysis or do not have Confident/Likely compute estimates (e.g., GPT-3.5 Turbo scores 8.3, but its compute is unknown), so they are not part of the analysis. Therefore, starting the 5, 10, and 15 series at GPT-4 implies misleadingly high compute for early models at these capability levels, and thus misleadingly steep slopes.
Conversely, some of the higher capability levels only have data for a couple months, and only have a couple models on the frontier, so I down-weight these high capability levels, including to zero.
Finally, the 45 threshold has extremely fast progress. Partially this is due to Grok-3 being released about a week before Claude Sonnet 4.7; had Sonnet been released first, Grok-3 would not be on the graph and the slope would be less steep (albeit still steep; removing Grok-3 brings the 45 bucket from a log10 slope of 4.48 to 3.52). I decided to weight this bucket similar to others.
[4] These two models are a good fit for this analysis because:
The Qwen team is a strong model developer and the models they create are often fairly close to the capabilities frontier (sometimes they are the frontier for open weight models) and the compute efficiency frontier (both of these models are on the compute efficiency for some analyses). Importantly, the Qwen2.5-72B-Instruct model was quite good at the time of its release, being the best open-weight model and improving slightly upon the then-recently released Llama-3.1-405B model that was trained with much more compute—Qwen2.5 isn’t an absurdly high compute baseline.
The Qwen team often releases technical reports including lots of details about their training, and they do so for these two models. This allows for reasonably accurate estimation of training compute.
Bothmodels are quite popular, receiving hundreds of thousands of Hugging Face downloads.
I expected them to have vastly different training compute (based on their active parameter counts varying by 22×) while also both being pretty capable, and I expected Qwen3-30B-A3B to be especially compute-efficient due to its use of a MoE architecture.
The Qwen3-8B model is probably the Qwen3 model closest in capabilities to Qwen2.5-72B (AAII score of 28 vs. 29), but is trained with more compute than 30B-A3B due to its dense architecture and is thus not as good of a representation of the mid-2025 frontier of training compute efficiency.
It appears there were errors with the conversion of footnotes from Google Docs to LessWrong. Corrected Footnotes:
[2] The weightings are, roughly, based on the following reasoning (some of these ideas are repeated elsewhere in this post):
In general, we should care more about higher capability levels.
There is a selection effect where people don’t bother making or testing really weak models anymore, so the data omit recent low-capability models, so we down-weight lower capabilities.
There is also a lack of data on early small models. For instance, GPT-4 is the first model to exceed 20 on the Intelligence Index, and it also appears in this data to be the first model to exceed 5, 10, and 15. But it isn’t actually the first model to exceed 5 (and we don’t know for 10 and 15). Earlier models either are not evaluated by Artificial Analysis or do not have Confident/Likely compute estimates (e.g., GPT-3.5 Turbo scores 8.3, but its compute is unknown), so they are not part of the analysis. Therefore, starting the 5, 10, and 15 series at GPT-4 implies misleadingly high compute for early models at these capability levels, and thus misleadingly steep slopes.
Conversely, some of the higher capability levels only have data for a couple months, and only have a couple models on the frontier, so I down-weight these high capability levels, including to zero.
Finally, the 45 threshold has extremely fast progress. Partially this is due to Grok-3 being released about a week before Claude Sonnet 4.7; had Sonnet been released first, Grok-3 would not be on the graph and the slope would be less steep (albeit still steep; removing Grok-3 brings the 45 bucket from a log10 slope of 4.48 to 3.52). I decided to weight this bucket similar to others.
[4] These two models are a good fit for this analysis because:
The Qwen team is a strong model developer and the models they create are often fairly close to the capabilities frontier (sometimes they are the frontier for open weight models) and the compute efficiency frontier (both of these models are on the compute efficiency for some analyses). Importantly, the Qwen2.5-72B-Instruct model was quite good at the time of its release, being the best open-weight model and improving slightly upon the then-recently released Llama-3.1-405B model that was trained with much more compute—Qwen2.5 isn’t an absurdly high compute baseline.
The Qwen team often releases technical reports including lots of details about their training, and they do so for these two models. This allows for reasonably accurate estimation of training compute.
Both models are quite popular, receiving hundreds of thousands of Hugging Face downloads.
I expected them to have vastly different training compute (based on their active parameter counts varying by 22×) while also both being pretty capable, and I expected Qwen3-30B-A3B to be especially compute-efficient due to its use of a MoE architecture.
The Qwen3-8B model is probably the Qwen3 model closest in capabilities to Qwen2.5-72B (AAII score of 28 vs. 29), but is trained with more compute than 30B-A3B due to its dense architecture and is thus not as good of a representation of the mid-2025 frontier of training compute efficiency.