Hi Dewi, thanks for the reply. This is interesting.
Sorry I realized I messed up in my top level comment where I incorrectly vibed the subselection of datasets used (I don’t think it was specified in the paper which 32 datasets are included or not, and Claude assumed in a way that looked reasonable enough but was off on a few). Sorry about that. I’m glad to your new figure exists.
Note though, ideally we’d want to ablate SHADE CoT and SHADE Action together (and maybe the two a-level datasets?) as these are very related tasks (I think?). Keeping them separate for like SHADE monitor makes it look like there are two unique tasks in this high end regime.
I attempted to make a new version tweaked version of my plot using the jsons you sent.
Things with large LOO Δ above generally have either have a tall change in this lollipop plot or are important parts of the tails (eg, arithmetic is important for flushing out the low end. Removing drops that whole area closer to the “tower of london” number).
Things in this center area would have to greatly change in worlds of the next measured doubling.
This is not including the possible extra expansion on the ICR due to the variance in the estimates them self. I’m not currently sure how to capture this. Some things like a-level MCQ has no IQR I think because all labeled at 1min.
> but as you can see these don’t materially change the conclusions we landed on around doubling time trends and associated uncertainty
Right, this does seem pretty robust to various slicing of this particular dataset of tasks. I wanted to start the thread with intuition building of what that particular set is and what’s changing. I feel personally now have a bit better sense of this.
The “doesn’t materially change” part should make sure we are truely internalize the associated uncertainty ranging from months to years. Similarly in the MKodama’s comment when looking at the logistic regression plots on the capabilities of a given model make sure we appreciate the error bars.
It still feels like we currently have a pretty confused signal for a lot of this span, especially in the upper half. The logistic fit will be rough. This is also already communicated by the individual estimates on the frontier models spanning 2 OOM from ~20 seconds to ~1 hour, so seems good as long as people internalize that to the extent.
All things to think about as the community thinks about for future work in the area, as well as making sure the community doesn’t over interpret the lowest nuance “373 day doubling time currently at 3min” version. The paper/post conclusion points around models capable of twenty-five minutes of no-CoT reasoning by 2030 are going need different methodology to measure (there’s not task diversity in this regime). Plus thinking on how well time horizon is a way to think about things we care about here.
Hi Dewi, thanks for the reply. This is interesting.
Sorry I realized I messed up in my top level comment where I incorrectly vibed the subselection of datasets used (I don’t think it was specified in the paper which 32 datasets are included or not, and Claude assumed in a way that looked reasonable enough but was off on a few). Sorry about that. I’m glad to your new figure exists.
Note though, ideally we’d want to ablate SHADE CoT and SHADE Action together (and maybe the two a-level datasets?) as these are very related tasks (I think?). Keeping them separate for like SHADE monitor makes it look like there are two unique tasks in this high end regime.
I attempted to make a new version tweaked version of my plot using the jsons you sent.
Things with large LOO Δ above generally have either have a tall change in this lollipop plot or are important parts of the tails (eg, arithmetic is important for flushing out the low end. Removing drops that whole area closer to the “tower of london” number).
Things in this center area would have to greatly change in worlds of the next measured doubling.
This is not including the possible extra expansion on the ICR due to the variance in the estimates them self. I’m not currently sure how to capture this. Some things like a-level MCQ has no IQR I think because all labeled at 1min.
> but as you can see these don’t materially change the conclusions we landed on around doubling time trends and associated uncertainty
Right, this does seem pretty robust to various slicing of this particular dataset of tasks. I wanted to start the thread with intuition building of what that particular set is and what’s changing. I feel personally now have a bit better sense of this.
The “doesn’t materially change” part should make sure we are truely internalize the associated uncertainty ranging from months to years. Similarly in the MKodama’s comment when looking at the logistic regression plots on the capabilities of a given model make sure we appreciate the error bars.
It still feels like we currently have a pretty confused signal for a lot of this span, especially in the upper half. The logistic fit will be rough. This is also already communicated by the individual estimates on the frontier models spanning 2 OOM from ~20 seconds to ~1 hour, so seems good as long as people internalize that to the extent.
All things to think about as the community thinks about for future work in the area, as well as making sure the community doesn’t over interpret the lowest nuance “373 day doubling time currently at 3min” version. The paper/post conclusion points around models capable of twenty-five minutes of no-CoT reasoning by 2030 are going need different methodology to measure (there’s not task diversity in this regime). Plus thinking on how well time horizon is a way to think about things we care about here.