Agree that the sigmoid model is clearly not actually right, and also that more benchmarks are needed, and that it’s reasonable to think there exist benchmarks that Astra gets 0% on without CoT. That said, I actually have no idea whether Astra would get literally 0% on benchmarks of 2 hours (which it sounds like is one of the things in your fit), or whether there exist possible benchmarks of >2 hours where Astra would not get 0% on, which would increase the TH estimate. Like, just based on the data you have, it really would not surprise me at all if Astra had a 60% success rate on tasks in the 2-4 hour bucket!
Basically overall, I feel pretty skeptical of the “add some synthetic 0% benchmarks” methodology, and would prefer a takeaway of “Astra doesn’t actually have a well-defined no-CoT TH because success rate is not sigmoidal in task length, but if it did, it would probably be somewhere north of 8 minutes”.
Another way of saying this: the [8m, 1h] CI basically comes from assuming that if you got more >2h benchmarks, Astra would get 0% on literally all of them. I don’t think that’s a reasonable assumption, especially just eye-balling your figure 1 (which would make me think that Astra would plausibly get around 50% on new benchmarks in the 2-4 hour bucket), and from looking at Neel’s post where it seems like Astra has an unusually lopsided no-CoT skill profile and is really great at some types of tasks.