According to this Fable 5 fit (in logodds space, I think linear extrapolation of bounded quantities are cursed), it looks like these benchmark will saturate between 1y and 10y METR time horizon, which seems surprisingly late given how difficult these tasks seem (they look like the sort of tasks which take on the order of hours for human experts?).
> they look like the sort of tasks which take on the order of hours for human experts? If anything, shorter.
For LMCA, the average rating takes me about 15 minutes, and I think there were essentially no data points where I spent more than 30 minutes. (That said, some final ratings are the product of discussions, and then there could be up to a couple of hours of combined human time involved, though that’s a bit different because it’s significantly parallel.)
For DTBench, average time is more like 5 minutes per question. (It took on average 6-7 minutes for me to check questions, which involved solving them, though also involved a lot of looking for and writing comments about ambiguities.) But I would also think the fit is off for DTBench, since Fable is already at 97 on DTBench.
(ACCoRD it’s hard to say—no humans have attempted the task.)
Our sense is that models are relatively bad at tasks like LMCA, and that they’re even worse at more freeform variants (like writing critiques instead of evaluating them).
Our sense is that models are relatively bad at tasks like LMCA
FWIW, my take is that they’re not bad compared to the reasonably smart human distribution! And surprisingly bad at free-form / long-form stuff given their LMCA performance.
(Maybe Em meant bad compared to coding)
(edit: She in fact meant compared coding, physics, Math etc., just in general their other capabilities. Which was actually kinda obvious from her comment, on reflection :P But I still wanted to avoid confusion bc I think ppl sometimes get the wrong impression when we say models are bad at LMCA)
According to this Fable 5 fit (in logodds space, I think linear extrapolation of bounded quantities are cursed), it looks like these benchmark will saturate between 1y and 10y METR time horizon, which seems surprisingly late given how difficult these tasks seem (they look like the sort of tasks which take on the order of hours for human experts?).
Thanks for the graphs!
> they look like the sort of tasks which take on the order of hours for human experts?
If anything, shorter.
For LMCA, the average rating takes me about 15 minutes, and I think there were essentially no data points where I spent more than 30 minutes. (That said, some final ratings are the product of discussions, and then there could be up to a couple of hours of combined human time involved, though that’s a bit different because it’s significantly parallel.)
For DTBench, average time is more like 5 minutes per question. (It took on average 6-7 minutes for me to check questions, which involved solving them, though also involved a lot of looking for and writing comments about ambiguities.) But I would also think the fit is off for DTBench, since Fable is already at 97 on DTBench.
(ACCoRD it’s hard to say—no humans have attempted the task.)
Our sense is that models are relatively bad at tasks like LMCA, and that they’re even worse at more freeform variants (like writing critiques instead of evaluating them).
FWIW, my take is that they’re not bad compared to the reasonably smart human distribution! And surprisingly bad at free-form / long-form stuff given their LMCA performance.
(Maybe Em meant bad compared to coding)
(edit: She in fact meant compared coding, physics, Math etc., just in general their other capabilities. Which was actually kinda obvious from her comment, on reflection :P But I still wanted to avoid confusion bc I think ppl sometimes get the wrong impression when we say models are bad at LMCA)
I would predict they fall significantly faster than this now that they’re targeted (I.e. that the trend will meaningfully diverge).