Right, bootstrapping here seems like a good approach, and appreciate the thoroughness here. It will communicate an interval, but there is still this wonder about what does the current doubling look like and what might future ones would look like. There’s various taxonomies here (eg, which tasks can gain on knowledge retrieval over reasoning, as you mentioned in the post). I tried this some in the fig I made.
Bootstrapping is a tool that tells us some things, but not others. It will still be grounded in the selected distribution of tasks (eg, in a cartoon case where bootstrapping can mislead us, imagine if we selected 10⁄32 tasks that were at the core about memorizing say birthdays of baseball players. If models suddenly got good at memorizing birthdays, it would look like a large improvement with a tight CI improvement, but wouldn’t capture overall variance of core metric. Again, cartoon extreme). So this is something was curious about after the initial read of the paper.
In some concrete terms here, an individual dataset will have a ~64% chance[1]of being included during bootstrapping. I was just trying to get an impression of this. Can understand that more times than not outlier-ish datasets like “Vibe-code Sabatoge” and “SHADE Arena-Monitor” will be included. I shared some findings in my comment, but as I said vibey and still underexplored for future work.
> Over 0.5hr regime
Right, to be clear I think focusing on a short-answer subset is the cleanest way to get at this non-CoT question. But when this filtering is where it gets somewhat problematic when trying to understand fitting a logistic regression for a given model in the 0 to 4hr range (like in post Fig 2), and when thinking about what will future doublings under the methodology have to look like as push upwards.
As you highlight when including non-short answer tasks for non-headline results, yes there are ~7 rather than 1, so more coverage (sorry if felt like I was ignoring this in a unfair way). I feel like I don’t have much certainty on how CoT-free the other 6 tasks are (which my impression is y’all are also a bit uncertain of, given filter out). As I mentioned in the comment elsewhere with Raemon, I know models, especially of a certain era, were very prone to doing things like writing long in-line comments, and basically packing CoT into the middle of their answers.
I agree that the fact that the regressions with and without these tasks doesn’t change much is reassuring. I still think the questions around dataset influence seem important. I felt like one thing I learned when looking at this was an appreciation for how the tails of logistic regression matter less (counter to my intuitions on linear regression where tail points matter). So then it becomes about looking at what are the most influential datasets near the crossover point as I tried to loosely explore above, where datasets that are near the crossover point will most influence the results. I didn’t make a version of my plot using the non-short-answer sample yet to try to see measures of influence there. Thanks for the out-of-band exchange sharing the data if did later.
> It would be helpful if you could ask more precise questions about the rest :)
I don’t have a quick expansion to give here right now as direct questions (regarding thoughts on time horizon as a metric, how estimateable it is on things we care about, elicitation variations). Possibly might expand later, or happy to chat synchronously sometime.
--
Thanks for engaging with this! I like many aspects of the study, and hopefully discussion is useful.
Thank you for your comments! Just following on from the point about individual dataset influence, we ran a leave-one-out study, sharing those results below. You are correct that SHADE and Vibe-coding (amongst a few others) have a strong impact, but as you can see these don’t materially change the conclusions we landed on around doubling time trends and associated uncertainty. We will add this to the paper, thank you pointing it out.
>The fact there is only single task in the >0.5hr regime looks pretty problematic
I also wanted to add that, as part of the bootstrap, we are incorporating uncertainty in solve times meaning that even if the point-estimate solve times for some questions is <0.5hr, it could still contribute in the bootstrap at >0.5hr. (this is just to say, there are more tasks in the bucket than it might seem!).
Hi Dewi, thanks for the reply. This is interesting.
Sorry I realized I messed up in my top level comment where I incorrectly vibed the subselection of datasets used (I don’t think it was specified in the paper which 32 datasets are included or not, and Claude assumed in a way that looked reasonable enough but was off on a few). Sorry about that. I’m glad to your new figure exists.
Note though, ideally we’d want to ablate SHADE CoT and SHADE Action together (and maybe the two a-level datasets?) as these are very related tasks (I think?). Keeping them separate for like SHADE monitor makes it look like there are two unique tasks in this high end regime.
I attempted to make a new version tweaked version of my plot using the jsons you sent.
Things with large LOO Δ above generally have either have a tall change in this lollipop plot or are important parts of the tails (eg, arithmetic is important for flushing out the low end. Removing drops that whole area closer to the “tower of london” number).
Things in this center area would have to greatly change in worlds of the next measured doubling.
This is not including the possible extra expansion on the ICR due to the variance in the estimates them self. I’m not currently sure how to capture this. Some things like a-level MCQ has no IQR I think because all labeled at 1min.
> but as you can see these don’t materially change the conclusions we landed on around doubling time trends and associated uncertainty
Right, this does seem pretty robust to various slicing of this particular dataset of tasks. I wanted to start the thread with intuition building of what that particular set is and what’s changing. I feel personally now have a bit better sense of this.
The “doesn’t materially change” part should make sure we are truely internalize the associated uncertainty ranging from months to years. Similarly in the MKodama’s comment when looking at the logistic regression plots on the capabilities of a given model make sure we appreciate the error bars.
It still feels like we currently have a pretty confused signal for a lot of this span, especially in the upper half. The logistic fit will be rough. This is also already communicated by the individual estimates on the frontier models spanning 2 OOM from ~20 seconds to ~1 hour, so seems good as long as people internalize that to the extent.
All things to think about as the community thinks about for future work in the area, as well as making sure the community doesn’t over interpret the lowest nuance “373 day doubling time currently at 3min” version. The paper/post conclusion points around models capable of twenty-five minutes of no-CoT reasoning by 2030 are going need different methodology to measure (there’s not task diversity in this regime). Plus thinking on how well time horizon is a way to think about things we care about here.
Thanks for reply!
> which tasks most influence the results
Right, bootstrapping here seems like a good approach, and appreciate the thoroughness here. It will communicate an interval, but there is still this wonder about what does the current doubling look like and what might future ones would look like. There’s various taxonomies here (eg, which tasks can gain on knowledge retrieval over reasoning, as you mentioned in the post). I tried this some in the fig I made.
Bootstrapping is a tool that tells us some things, but not others. It will still be grounded in the selected distribution of tasks (eg, in a cartoon case where bootstrapping can mislead us, imagine if we selected 10⁄32 tasks that were at the core about memorizing say birthdays of baseball players. If models suddenly got good at memorizing birthdays, it would look like a large improvement with a tight CI improvement, but wouldn’t capture overall variance of core metric. Again, cartoon extreme). So this is something was curious about after the initial read of the paper.
In some concrete terms here, an individual dataset will have a ~64% chance[1]of being included during bootstrapping. I was just trying to get an impression of this. Can understand that more times than not outlier-ish datasets like “Vibe-code Sabatoge” and “SHADE Arena-Monitor” will be included. I shared some findings in my comment, but as I said vibey and still underexplored for future work.
> Over 0.5hr regime
Right, to be clear I think focusing on a short-answer subset is the cleanest way to get at this non-CoT question. But when this filtering is where it gets somewhat problematic when trying to understand fitting a logistic regression for a given model in the 0 to 4hr range (like in post Fig 2), and when thinking about what will future doublings under the methodology have to look like as push upwards.
As you highlight when including non-short answer tasks for non-headline results, yes there are ~7 rather than 1, so more coverage (sorry if felt like I was ignoring this in a unfair way). I feel like I don’t have much certainty on how CoT-free the other 6 tasks are (which my impression is y’all are also a bit uncertain of, given filter out). As I mentioned in the comment elsewhere with Raemon, I know models, especially of a certain era, were very prone to doing things like writing long in-line comments, and basically packing CoT into the middle of their answers.
I agree that the fact that the regressions with and without these tasks doesn’t change much is reassuring. I still think the questions around dataset influence seem important. I felt like one thing I learned when looking at this was an appreciation for how the tails of logistic regression matter less (counter to my intuitions on linear regression where tail points matter). So then it becomes about looking at what are the most influential datasets near the crossover point as I tried to loosely explore above, where datasets that are near the crossover point will most influence the results. I didn’t make a version of my plot using the non-short-answer sample yet to try to see measures of influence there. Thanks for the out-of-band exchange sharing the data if did later.
> It would be helpful if you could ask more precise questions about the rest :)
I don’t have a quick expansion to give here right now as direct questions (regarding thoughts on time horizon as a metric, how estimateable it is on things we care about, elicitation variations). Possibly might expand later, or happy to chat synchronously sometime.
--
Thanks for engaging with this! I like many aspects of the study, and hopefully discussion is useful.
Sampling 32 tasks with replacement.
Thank you for your comments! Just following on from the point about individual dataset influence, we ran a leave-one-out study, sharing those results below. You are correct that SHADE and Vibe-coding (amongst a few others) have a strong impact, but as you can see these don’t materially change the conclusions we landed on around doubling time trends and associated uncertainty. We will add this to the paper, thank you pointing it out.
>The fact there is only single task in the >0.5hr regime looks pretty problematic
I also wanted to add that, as part of the bootstrap, we are incorporating uncertainty in solve times meaning that even if the point-estimate solve times for some questions is <0.5hr, it could still contribute in the bootstrap at >0.5hr. (this is just to say, there are more tasks in the bucket than it might seem!).
Hi Dewi, thanks for the reply. This is interesting.
Sorry I realized I messed up in my top level comment where I incorrectly vibed the subselection of datasets used (I don’t think it was specified in the paper which 32 datasets are included or not, and Claude assumed in a way that looked reasonable enough but was off on a few). Sorry about that. I’m glad to your new figure exists.
Note though, ideally we’d want to ablate SHADE CoT and SHADE Action together (and maybe the two a-level datasets?) as these are very related tasks (I think?). Keeping them separate for like SHADE monitor makes it look like there are two unique tasks in this high end regime.
I attempted to make a new version tweaked version of my plot using the jsons you sent.
Things with large LOO Δ above generally have either have a tall change in this lollipop plot or are important parts of the tails (eg, arithmetic is important for flushing out the low end. Removing drops that whole area closer to the “tower of london” number).
Things in this center area would have to greatly change in worlds of the next measured doubling.
This is not including the possible extra expansion on the ICR due to the variance in the estimates them self. I’m not currently sure how to capture this. Some things like a-level MCQ has no IQR I think because all labeled at 1min.
> but as you can see these don’t materially change the conclusions we landed on around doubling time trends and associated uncertainty
Right, this does seem pretty robust to various slicing of this particular dataset of tasks. I wanted to start the thread with intuition building of what that particular set is and what’s changing. I feel personally now have a bit better sense of this.
The “doesn’t materially change” part should make sure we are truely internalize the associated uncertainty ranging from months to years. Similarly in the MKodama’s comment when looking at the logistic regression plots on the capabilities of a given model make sure we appreciate the error bars.
It still feels like we currently have a pretty confused signal for a lot of this span, especially in the upper half. The logistic fit will be rough. This is also already communicated by the individual estimates on the frontier models spanning 2 OOM from ~20 seconds to ~1 hour, so seems good as long as people internalize that to the extent.
All things to think about as the community thinks about for future work in the area, as well as making sure the community doesn’t over interpret the lowest nuance “373 day doubling time currently at 3min” version. The paper/post conclusion points around models capable of twenty-five minutes of no-CoT reasoning by 2030 are going need different methodology to measure (there’s not task diversity in this regime). Plus thinking on how well time horizon is a way to think about things we care about here.