Thank you for your comments! Just following on from the point about individual dataset influence, we ran a leave-one-out study, sharing those results below. You are correct that SHADE and Vibe-coding (amongst a few others) have a strong impact, but as you can see these don’t materially change the conclusions we landed on around doubling time trends and associated uncertainty. We will add this to the paper, thank you pointing it out.
>The fact there is only single task in the >0.5hr regime looks pretty problematic
I also wanted to add that, as part of the bootstrap, we are incorporating uncertainty in solve times meaning that even if the point-estimate solve times for some questions is <0.5hr, it could still contribute in the bootstrap at >0.5hr. (this is just to say, there are more tasks in the bucket than it might seem!).
Thank you for your comments! Just following on from the point about individual dataset influence, we ran a leave-one-out study, sharing those results below. You are correct that SHADE and Vibe-coding (amongst a few others) have a strong impact, but as you can see these don’t materially change the conclusions we landed on around doubling time trends and associated uncertainty. We will add this to the paper, thank you pointing it out.
>The fact there is only single task in the >0.5hr regime looks pretty problematic
I also wanted to add that, as part of the bootstrap, we are incorporating uncertainty in solve times meaning that even if the point-estimate solve times for some questions is <0.5hr, it could still contribute in the bootstrap at >0.5hr. (this is just to say, there are more tasks in the bucket than it might seem!).