To account for the fact that our task suite has an insufficient number of benchmarks with long time horizons, we assume that there is some number (N) of benchmarks with problems of length x (2-96hours in the main plot). To compute the 50% TH we assume astra gets 0% performance on these benchmarks.
Is this a reasonable assumption? Intuitively, it would be surprising if astra got non-zero no-CoT performance on very long benchmarks (e.g., a no-CoT TH of 32 hours seems a priori implausible).
I don’t get why this is a reasonable assumption. To the degree you believe the sigmoid model, it’s apparently not what that model predicts on your data (given that it changes your TH estimate), so I suppose it must be from other evidence. But what is that other evidence?
The sigmoid model doesn’t provide an informative fit when all the tasks are saturated either!
As we say in the post, astra does get low performance on some of the shorter benchmarks (e.g., kenken, sudoko). To me it seems pretty reasonable to assume there were a few benchmarks containing very long tasks which astra couldn’t do without CoT.
Let’s assume that we had included just 3 other benchmarks in our suite, and these benchmarks were full of problems that took humans 32, 64, 128 hours respectively, and astra got 0% on these three benchmarks. Then the TH computation would give a median of 60 min [11 min, 5.9 h] CI. This pushes the TH estimate down (both the median, 7.2 hours, and the upper-bound). Even if you assume one 32 hour benchmark the median gets pushed down to 1.6 hours. (The upper-bound is more sensitive to the assumed number of benchmarks.) In the sensitivity analysis for this part, the median TH is always between 13 mins and 1.6 hours, and for N>1 the CIs are always between 6 mins and 6 hours. I basically believe these are much more reasonable bounds over the TH—though it’s not super principled.
I do think this is a big limitation, which is why we say we need more benchmarks with longer problems.
I don’t get why this is a reasonable assumption. To the degree you believe the sigmoid model, it’s apparently not what that model predicts on your data (given that it changes your TH estimate), so I suppose it must be from other evidence. But what is that other evidence?
Hi Daniel,
The sigmoid model doesn’t provide an informative fit when all the tasks are saturated either!
As we say in the post, astra does get low performance on some of the shorter benchmarks (e.g., kenken, sudoko). To me it seems pretty reasonable to assume there were a few benchmarks containing very long tasks which astra couldn’t do without CoT.
Let’s assume that we had included just 3 other benchmarks in our suite, and these benchmarks were full of problems that took humans 32, 64, 128 hours respectively, and astra got 0% on these three benchmarks. Then the TH computation would give a median of 60 min [11 min, 5.9 h] CI. This pushes the TH estimate down (both the median, 7.2 hours, and the upper-bound). Even if you assume one 32 hour benchmark the median gets pushed down to 1.6 hours. (The upper-bound is more sensitive to the assumed number of benchmarks.) In the sensitivity analysis for this part, the median TH is always between 13 mins and 1.6 hours, and for N>1 the CIs are always between 6 mins and 6 hours. I basically believe these are much more reasonable bounds over the TH—though it’s not super principled.
I do think this is a big limitation, which is why we say we need more benchmarks with longer problems.