The sigmoid model doesn’t provide an informative fit when all the tasks are saturated either!
As we say in the post, astra does get low performance on some of the shorter benchmarks (e.g., kenken, sudoko). To me it seems pretty reasonable to assume there were a few benchmarks containing very long tasks which astra couldn’t do without CoT.
Let’s assume that we had included just 3 other benchmarks in our suite, and these benchmarks were full of problems that took humans 32, 64, 128 hours respectively, and astra got 0% on these three benchmarks. Then the TH computation would give a median of 60 min [11 min, 5.9 h] CI. This pushes the TH estimate down (both the median, 7.2 hours, and the upper-bound). Even if you assume one 32 hour benchmark the median gets pushed down to 1.6 hours. (The upper-bound is more sensitive to the assumed number of benchmarks.) In the sensitivity analysis for this part, the median TH is always between 13 mins and 1.6 hours, and for N>1 the CIs are always between 6 mins and 6 hours. I basically believe these are much more reasonable bounds over the TH—though it’s not super principled.
I do think this is a big limitation, which is why we say we need more benchmarks with longer problems.
Hi Daniel,
The sigmoid model doesn’t provide an informative fit when all the tasks are saturated either!
As we say in the post, astra does get low performance on some of the shorter benchmarks (e.g., kenken, sudoko). To me it seems pretty reasonable to assume there were a few benchmarks containing very long tasks which astra couldn’t do without CoT.
Let’s assume that we had included just 3 other benchmarks in our suite, and these benchmarks were full of problems that took humans 32, 64, 128 hours respectively, and astra got 0% on these three benchmarks. Then the TH computation would give a median of 60 min [11 min, 5.9 h] CI. This pushes the TH estimate down (both the median, 7.2 hours, and the upper-bound). Even if you assume one 32 hour benchmark the median gets pushed down to 1.6 hours. (The upper-bound is more sensitive to the assumed number of benchmarks.) In the sensitivity analysis for this part, the median TH is always between 13 mins and 1.6 hours, and for N>1 the CIs are always between 6 mins and 6 hours. I basically believe these are much more reasonable bounds over the TH—though it’s not super principled.
I do think this is a big limitation, which is why we say we need more benchmarks with longer problems.