Thank you for your comments! Just following on from the point about individual dataset influence, we ran a leave-one-out study, sharing those results below. You are correct that SHADE and Vibe-coding (amongst a few others) have a strong impact, but as you can see these don’t materially change the conclusions we landed on around doubling time trends and associated uncertainty. We will add this to the paper, thank you pointing it out.
>The fact there is only single task in the >0.5hr regime looks pretty problematic
I also wanted to add that, as part of the bootstrap, we are incorporating uncertainty in solve times meaning that even if the point-estimate solve times for some questions is <0.5hr, it could still contribute in the bootstrap at >0.5hr. (this is just to say, there are more tasks in the bucket than it might seem!).
Hi Dewi, thanks for the reply. This is interesting.
Sorry I realized I messed up in my top level comment where I incorrectly vibed the subselection of datasets used (I don’t think it was specified in the paper which 32 datasets are included or not, and Claude assumed in a way that looked reasonable enough but was off on a few). Sorry about that. I’m glad to your new figure exists.
Note though, ideally we’d want to ablate SHADE CoT and SHADE Action together (and maybe the two a-level datasets?) as these are very related tasks (I think?). Keeping them separate for like SHADE monitor makes it look like there are two unique tasks in this high end regime.
I attempted to make a new version tweaked version of my plot using the jsons you sent.
Things with large LOO Δ above generally have either have a tall change in this lollipop plot or are important parts of the tails (eg, arithmetic is important for flushing out the low end. Removing drops that whole area closer to the “tower of london” number).
Things in this center area would have to greatly change in worlds of the next measured doubling.
This is not including the possible extra expansion on the ICR due to the variance in the estimates them self. I’m not currently sure how to capture this. Some things like a-level MCQ has no IQR I think because all labeled at 1min.
> but as you can see these don’t materially change the conclusions we landed on around doubling time trends and associated uncertainty
Right, this does seem pretty robust to various slicing of this particular dataset of tasks. I wanted to start the thread with intuition building of what that particular set is and what’s changing. I feel personally now have a bit better sense of this.
The “doesn’t materially change” part should make sure we are truely internalize the associated uncertainty ranging from months to years. Similarly in the MKodama’s comment when looking at the logistic regression plots on the capabilities of a given model make sure we appreciate the error bars.
It still feels like we currently have a pretty confused signal for a lot of this span, especially in the upper half. The logistic fit will be rough. This is also already communicated by the individual estimates on the frontier models spanning 2 OOM from ~20 seconds to ~1 hour, so seems good as long as people internalize that to the extent.
All things to think about as the community thinks about for future work in the area, as well as making sure the community doesn’t over interpret the lowest nuance “373 day doubling time currently at 3min” version. The paper/post conclusion points around models capable of twenty-five minutes of no-CoT reasoning by 2030 are going need different methodology to measure (there’s not task diversity in this regime). Plus thinking on how well time horizon is a way to think about things we care about here.
Thank you for your comments! Just following on from the point about individual dataset influence, we ran a leave-one-out study, sharing those results below. You are correct that SHADE and Vibe-coding (amongst a few others) have a strong impact, but as you can see these don’t materially change the conclusions we landed on around doubling time trends and associated uncertainty. We will add this to the paper, thank you pointing it out.
>The fact there is only single task in the >0.5hr regime looks pretty problematic
I also wanted to add that, as part of the bootstrap, we are incorporating uncertainty in solve times meaning that even if the point-estimate solve times for some questions is <0.5hr, it could still contribute in the bootstrap at >0.5hr. (this is just to say, there are more tasks in the bucket than it might seem!).
Hi Dewi, thanks for the reply. This is interesting.
Sorry I realized I messed up in my top level comment where I incorrectly vibed the subselection of datasets used (I don’t think it was specified in the paper which 32 datasets are included or not, and Claude assumed in a way that looked reasonable enough but was off on a few). Sorry about that. I’m glad to your new figure exists.
Note though, ideally we’d want to ablate SHADE CoT and SHADE Action together (and maybe the two a-level datasets?) as these are very related tasks (I think?). Keeping them separate for like SHADE monitor makes it look like there are two unique tasks in this high end regime.
I attempted to make a new version tweaked version of my plot using the jsons you sent.
Things with large LOO Δ above generally have either have a tall change in this lollipop plot or are important parts of the tails (eg, arithmetic is important for flushing out the low end. Removing drops that whole area closer to the “tower of london” number).
Things in this center area would have to greatly change in worlds of the next measured doubling.
This is not including the possible extra expansion on the ICR due to the variance in the estimates them self. I’m not currently sure how to capture this. Some things like a-level MCQ has no IQR I think because all labeled at 1min.
> but as you can see these don’t materially change the conclusions we landed on around doubling time trends and associated uncertainty
Right, this does seem pretty robust to various slicing of this particular dataset of tasks. I wanted to start the thread with intuition building of what that particular set is and what’s changing. I feel personally now have a bit better sense of this.
The “doesn’t materially change” part should make sure we are truely internalize the associated uncertainty ranging from months to years. Similarly in the MKodama’s comment when looking at the logistic regression plots on the capabilities of a given model make sure we appreciate the error bars.
It still feels like we currently have a pretty confused signal for a lot of this span, especially in the upper half. The logistic fit will be rough. This is also already communicated by the individual estimates on the frontier models spanning 2 OOM from ~20 seconds to ~1 hour, so seems good as long as people internalize that to the extent.
All things to think about as the community thinks about for future work in the area, as well as making sure the community doesn’t over interpret the lowest nuance “373 day doubling time currently at 3min” version. The paper/post conclusion points around models capable of twenty-five minutes of no-CoT reasoning by 2030 are going need different methodology to measure (there’s not task diversity in this regime). Plus thinking on how well time horizon is a way to think about things we care about here.