While Q1 2025 results are available for bots, human baseline comparisons haven’t been released yet.
It’s out now : https://www.metaculus.com/notebooks/38673/q1-ai-benchmarking-results/, title is ”Pros Crush Bots”.
Key lesson:
One conclusion we have drawn from this is that the most important factor for good forecasting is the base model, and additional prompting and infrastructure on top of this provide marginal gains.
Scaling remains undefeated.
It’s out now : https://www.metaculus.com/notebooks/38673/q1-ai-benchmarking-results/, title is ”Pros Crush Bots”.
Key lesson:
Scaling remains undefeated.