The most recent Tetlockian forecasting style thing I’ve spent substantial time on is the 2025 and 2026 AI forecasting surveys, in which hundreds of people each year have made predictions a year out on benchmarks, and other indicators such as revenue.
The theory of change is to (a) establish common knowledge about how fast things are going relative to people’s expectations (and we collect data on people’s overall views on when AGI will be reached so we can sort of see if we’re “on track” for that), and (b) identify which people seem to be making the most accurate predictions. Importantly, it is not to elicit predictions that are directly useful for important decisions.
I’ve observed some evidence of this working, e.g. re: (a) establishing common knowedge, Anson of Epoch wrote an analysis that I’ve seen referenced a few times. I’m glad to have a data point against the common refrain of “people underpredict benchmark scores and overpredict real-world impact” from revenue outpacing people’s predictions (though it is a narrow and single data point).
Re: (b) identifying who is making the most accurate predictions, I found it informative that in Anson’s analysis (footnote 1), forecasters with pre and post-2030 timelines performed similarly. I’ve seen some people cite Ryan G and Ajeya’s #2 and #3 performance as evidence that we should listen to them, which is maybe good but I think people might be over-updating on the results with so few questions (I certainly pay attention to Ryan and Ajeya’s forecasts, but almost entirely for other reasons).
Overall, it’s unclear to me how this impactful this has been. I decided to run the 2026 survey because it seems at least a bit impactful and it doesn’t take that much time (I logged 18 hours on setting up the 2026 version, I’d guess that some others who helped spent a total of 20-60 hours). But the decision was borderline.
The most recent Tetlockian forecasting style thing I’ve spent substantial time on is the 2025 and 2026 AI forecasting surveys, in which hundreds of people each year have made predictions a year out on benchmarks, and other indicators such as revenue.
The theory of change is to (a) establish common knowledge about how fast things are going relative to people’s expectations (and we collect data on people’s overall views on when AGI will be reached so we can sort of see if we’re “on track” for that), and (b) identify which people seem to be making the most accurate predictions. Importantly, it is not to elicit predictions that are directly useful for important decisions.
I’ve observed some evidence of this working, e.g. re: (a) establishing common knowedge, Anson of Epoch wrote an analysis that I’ve seen referenced a few times. I’m glad to have a data point against the common refrain of “people underpredict benchmark scores and overpredict real-world impact” from revenue outpacing people’s predictions (though it is a narrow and single data point).
Re: (b) identifying who is making the most accurate predictions, I found it informative that in Anson’s analysis (footnote 1), forecasters with pre and post-2030 timelines performed similarly. I’ve seen some people cite Ryan G and Ajeya’s #2 and #3 performance as evidence that we should listen to them, which is maybe good but I think people might be over-updating on the results with so few questions (I certainly pay attention to Ryan and Ajeya’s forecasts, but almost entirely for other reasons).
Overall, it’s unclear to me how this impactful this has been. I decided to run the 2026 survey because it seems at least a bit impactful and it doesn’t take that much time (I logged 18 hours on setting up the 2026 version, I’d guess that some others who helped spent a total of 20-60 hours). But the decision was borderline.