How does one gather novel evidence for this benchmark? METR’s last evaluation of Chinese models was Kimi K2 (Nov. 2025) in the old methodology, ensuring a 58 min 50% TH and a 15 min 80%TH, on par with Claude 3.7 Sonnet. ARC-AGI-2′s last eval was of KimiK2.5 (Jan 2026, outsmarted by Grok 4 (JULY 2025), GPT-5-pro, then Claude Sonnet 4.5; HOW did Kimi end up being TEN months behind according to a data point?) Where did METR evaluate DeepSeek v.3.2? What could one do to incorporate DeepSeek v4 Pro’s CAISI-led eval which had CAISI’s varsion output 46% as opposed to Opus 4.6′s 63% (which was officially evaluated as 68.8% or 69.2%) and GPT-5.5′s 79% by CAISI/85% by the ARC-AGI team?
Incorporating the CAISI-results is not ideal since models are not run in the same setting as the other ARC-AGI2 results. This is a very common pattern, where there are a bunch of benchmarks with interesting results, but it’s rare that we have benchmarks where we have:
very good coverage of both open and closed models
results over a fairly long period of time
all the results are directly comparable
The lack of such benchmarks, we have only a few, is what makes this kind of analysis hard.
New datapoints come in as open models are tested and cross a threshold previously crossed by a closed model. There is of course always a problem that each benchmark does not run all the models, so we have to make a judgement in each case (each threshold/datapoint) if both the closed and the open model was plausibly the first model to have crossed that threshold, or if there are other missing models that likely would change the gap significantly if they were included, if so we reject that datapoint.
As for Kimi k2.5 being 10 months behind in one datapoint, this is because we assign all the gaps closed by an open model to the time of the open model release, that is how we define the backward-looking gap. So even though that particular model was less behind the frontier than 10 months, there was a gap there of around 10 months that needs to be counted in the open vs closed comparison. The reason we choose to define the gap in this way, is that then we can compare the data from today with data from previous times without any bias.
The methodology is simple. You define a set of thresholds, usually just every 5%, and then we start a clock every time a closed model crosses a threshold, and stop the clock when an open model crosses the same threshold. That is the y-value, a measurement of the gap. Then there is the question of which time (on the x-axis) to associate this y-value with, and we go with the backward-looking perspective and associate the gap with the release date of the open model. The forward-looking perspective would associate the gap with the closed model release date.
As for ARC-AGI 2, here we went with smaller intervals between the thresolds, since there were several scores within a small interval. If DeepSeek-V4 would score 40%, it would cross a bunch of thresholds at once. Some of the thresholds may not count because they would be duplicate (if the same model pair shows up in several thresholds in the same benchmark we only count one of them), some of them may not count if we expect another open model would have crossed it sooner if tested. Any remaining thresholds would be datapoints going into the analysis, the y-value would be determined by the gap since a closed model crossed that threshold, and the x-value would be the release date of DeepSeek-V4.
Does it mean that Deepseek V4-Pro (Apr 24, 2026) would be coupled with Gemini 3 Deep Think, Opus 4.5, Gemini 3 Pro (and Grok 4/GPT-5-Pro since no one crossed the 20/25% thresholds without reaching 30+%?), reaching 5 months for this trio (since it was released in Nov 2025) and 8-9 months for Grok 4 or GPT-5-Pro?
Yea, that sounds right. And some of these thresholds may be judgement calls based on if GLM 5.1 was run etc, but probably I would lean towards accepting them, at least the highest ones. I would have expected DSv4 to shorten the gap more than it has done in the results so far, but if we get scores like this coming in it could affect the results.
How does one gather novel evidence for this benchmark? METR’s last evaluation of Chinese models was Kimi K2 (Nov. 2025) in the old methodology, ensuring a 58 min 50% TH and a 15 min 80%TH, on par with Claude 3.7 Sonnet. ARC-AGI-2′s last eval was of KimiK2.5 (Jan 2026, outsmarted by Grok 4 (JULY 2025), GPT-5-pro, then Claude Sonnet 4.5; HOW did Kimi end up being TEN months behind according to a data point?)
Where did METR evaluate DeepSeek v.3.2?What could one do to incorporate DeepSeek v4 Pro’s CAISI-led eval which had CAISI’s varsion output 46% as opposed to Opus 4.6′s 63% (which was officially evaluated as 68.8% or 69.2%) and GPT-5.5′s 79% by CAISI/85% by the ARC-AGI team?Incorporating the CAISI-results is not ideal since models are not run in the same setting as the other ARC-AGI2 results. This is a very common pattern, where there are a bunch of benchmarks with interesting results, but it’s rare that we have benchmarks where we have:
very good coverage of both open and closed models
results over a fairly long period of time
all the results are directly comparable
The lack of such benchmarks, we have only a few, is what makes this kind of analysis hard.
New datapoints come in as open models are tested and cross a threshold previously crossed by a closed model. There is of course always a problem that each benchmark does not run all the models, so we have to make a judgement in each case (each threshold/datapoint) if both the closed and the open model was plausibly the first model to have crossed that threshold, or if there are other missing models that likely would change the gap significantly if they were included, if so we reject that datapoint.
As for Kimi k2.5 being 10 months behind in one datapoint, this is because we assign all the gaps closed by an open model to the time of the open model release, that is how we define the backward-looking gap. So even though that particular model was less behind the frontier than 10 months, there was a gap there of around 10 months that needs to be counted in the open vs closed comparison. The reason we choose to define the gap in this way, is that then we can compare the data from today with data from previous times without any bias.
Could you explaim the method in more detail? What would you do with a counterfactual OFFICIAL ARC-AGI-2 evaluation of DeepSeek v4 Pro as 40% or 52%?
The methodology is simple. You define a set of thresholds, usually just every 5%, and then we start a clock every time a closed model crosses a threshold, and stop the clock when an open model crosses the same threshold. That is the y-value, a measurement of the gap. Then there is the question of which time (on the x-axis) to associate this y-value with, and we go with the backward-looking perspective and associate the gap with the release date of the open model. The forward-looking perspective would associate the gap with the closed model release date.
As for ARC-AGI 2, here we went with smaller intervals between the thresolds, since there were several scores within a small interval. If DeepSeek-V4 would score 40%, it would cross a bunch of thresholds at once. Some of the thresholds may not count because they would be duplicate (if the same model pair shows up in several thresholds in the same benchmark we only count one of them), some of them may not count if we expect another open model would have crossed it sooner if tested. Any remaining thresholds would be datapoints going into the analysis, the y-value would be determined by the gap since a closed model crossed that threshold, and the x-value would be the release date of DeepSeek-V4.
Does it mean that Deepseek V4-Pro (Apr 24, 2026) would be coupled with Gemini 3 Deep Think, Opus 4.5, Gemini 3 Pro (and Grok 4/GPT-5-Pro since no one crossed the 20/25% thresholds without reaching 30+%?), reaching 5 months for this trio (since it was released in Nov 2025) and 8-9 months for Grok 4 or GPT-5-Pro?
Yea, that sounds right. And some of these thresholds may be judgement calls based on if GLM 5.1 was run etc, but probably I would lean towards accepting them, at least the highest ones. I would have expected DSv4 to shorten the gap more than it has done in the results so far, but if we get scores like this coming in it could affect the results.