Metaculus has forecasted the trajectory of 30 Our World in Data metrics on its platform using both a public tournament and a group of Pro Forecasters. Both accuracy and transparency in reasoning were considered essential to the project. The public tournament was available to the platform’s full community of over 2,000 forecasters, while the private forecasting space was intended for a small group of the top 2% of Metaculus’s most accurate forecasters. Combining both approaches allows for a variety of perspectives and a broad discussion, as well as ensuring that predictions are of the highest quality. With substantive discussions surrounding this tournament’s 30 forecasting questions, Metaculus has been able to better understand: (1) approaches and methods adopted by elite forecasters, (2) rationales driving their prediction for each time period, and (3) a composite perspective into our world 100 years from now. …
Based on the timeframes that were of greatest interest to our stakeholders, Metaculus elicited forecasts for a variety of intervals — 1, 3, 10, 30, and 100 years out from the year 2022. For each OWID metric that Metaculus and the initiative’s funders selected, this tournament created grouped questions that asked forecasters to give their predicted value of that metric in the years 2023, 2025, 2032, 2052, and 2122. With such an array of time horizons, the tournament takes a less common forecasting approach that allows for direct comparison between short-term and long-term predictions.
The composite perspective into the world of 2122 seems distinctly pedestrian to me
Our World in 2122
In the year 2122, the global population is expected to be around 8 or 9 billion, with a chance that number reaches as high as 20 billion due to technological and economic advancements. Lifespans are expected to be significantly longer, especially in the G7 countries, thanks to the achievement of longevity escape velocity. Reproduction is anticipated by some to primarily utilize ectogenesis technology, though some may choose to avoid it for personal or ethical reasons. The decline in population is expected to slow, but there remains uncertainty as to whether it will level off or decrease further.
On the economic front, GDP per capita is believed to be quite high and will likely continue increasing. In fact, most economies are expected to be as advanced as present-day Scandinavian countries, with energy being inexpensive and the price of most goods and services being close to zero. Productivity is forecasted to significantly increase, with advanced artificial intelligence playing a major role in GDP growth. The United States economy is anticipated to continue growing, potentially transitioning to a more service-based industry and experiencing changes in government structure. However, the concept of the United States as we know it may be fundamentally different by this time. Instead, it is possible that global governance will be implemented and the number of sovereign states will drastically decrease. Wealth distribution remains a key uncertainty — while it’s possible that a small group could control a significant portion of resources, a shift toward more equitable distribution may win out.
Even with global governance, many pressing challenges will remain. Despite drastic cuts in carbon dioxide emissions, climate change is expected to be a major concern — sea levels will rise and extreme weather events will become more frequent. Nevertheless, renewable energy is also expected to increase in parallel, with the potential for advanced technologies, like nuclear fusion, to power a significant portion of the global economy. Even though the precise global energy mix remains uncertain, transportation will undergo major electrification, with an anticipated shift towards both zero emissions and autonomous vehicles.
Technology will continue expanding at a rapid pace and play an even greater role in daily life. The use of robots is predicted to increase and revolutionary innovations like brain-computer interfaces could fundamentally change daily life as we know it. Education is therefore expected to become more personalized and virtual, as will the healthcare system. The latter will undergo many transformations as artificial intelligence opens the door for hyper-personalized treatments and precision medicine. The concept of work will change too, becoming far more flexible and remote. Leisure time is expected to increase, with virtual reality and space tourism becoming popular for many.
Despite great uncertainty and persistent challenges, 2122 is expected to be a time of great advancement and possibility.
I wondered how the first two timeframes (2023 and 2025) had done so far, so I asked the models:
GPT-5.5 (high): “pretty good on slow-moving macro and demographic variables, good on space objects and CO₂, and bad on a few discontinuous or definition-sensitive variables: terrorism, fastest-supercomputer FLOPS, people living in liberal democracies, and AI training compute”
I used the Metaculus report tables for the forecast medians, with “Pro” meaning Pro Forecaster median and “Public” meaning the Metaculus Prediction. The report says the tournament forecasted 2023, 2025, 2032, 2052, and 2122 values, and that the summary tables cover the 20 questions forecasted by both Pros and the public.
Bottom line
So far, the forecasts look pretty good on slow-moving macro and demographic variables, good on space objects and CO₂, and bad on a few discontinuous or definition-sensitive variables: terrorism, fastest-supercomputer FLOPS, people living in liberal democracies, and AI training compute. AI investment is messy: they got the direction right, but whether the forecast was high or low depends heavily on whether you use narrow private funding, broad corporate AI investment, or capex-like spending.
A rough scorecard: among the items I could actually check with reasonably current data, I’d call about 7 good or very good, 4 clear misses, and the rest not yet properly scoreable because the official annual data lag into 2026 or 2027.
Scorecard for the clearest cases
Metric
Forecast for 2023 and 2025
What happened
Verdict
Annual objects launched into space
2023: Pro 3,000, Public 2,700. 2025: Pro 4,300, Public 4,000.
OWID says 2,903 objects were launched in 2023 and 4,510 in 2025.
Very good. Pros were especially close. Public undercalled 2025 by about 11%.
Global CO₂ emissions
2023: Pro 37.7 Gt, Public 36.4 Gt. 2025: Pro 38.1 Gt, Public 36.2 Gt.
IEA put 2023 energy-related CO₂ at 37.4 Gt; Global Carbon Project projected 2025 fossil CO₂ at 38.1 Gt.
Pros basically nailed it. Public forecasts were too optimistic about emissions falling.
Fastest supercomputer FLOPS
2023: Pro 1.5e18, Public 1.3e18. 2025: Pro 4.1e18, Public 4.5e18.
Frontier was still around 1.1 exaFLOP in late 2023; El Capitan led the Nov. 2025 TOP500 at 1.809 exaFLOPS Rmax.
Clear miss high. They expected the exascale transition to accelerate faster than the TOP500 benchmark actually did.
Global nuclear warhead stockpiles
2023: Pro 9.46k, Public 9.40k. 2025: Pro 9.52k, Public 9.38k.
SIPRI estimated about 9,614 warheads in military stockpiles in Jan. 2025.
Very good. Slightly low, but right order and right direction.
Total terrorism fatalities
2023: Pro 19k, Public 20k. 2025: Pro 19k, Public 20k.
Global Terrorism Index reported 8,352 deaths in 2023, and its 2026 report gives 5,582 deaths in 2025.
Bad miss high. They overestimated by roughly 2.3x in 2023 and 3.4x in 2025.
People living in liberal democracies
2023: Pro 1.02B, Public 1.05B. 2025: Pro 1.00B, Public 1.10B.
V-Dem’s 2026 Democracy Report says only 7% of the world population, about 0.6B people, lived in liberal democracies at end-2025.
Clear miss high. This is probably the biggest substantive miss, driven by democratic backsliding and classification changes, especially the U.S. losing liberal-democracy status in V-Dem.
Greatest AI training compute
2023: Pro 6.5e24 FLOP, Public 1.3e25. 2025: Pro 2.9e25, Public 6.6e25.
Epoch’s 2026 trends page says the largest known training run was Grok 4 at around 5e26 FLOP.
Miss low. They correctly expected rapid growth, but still undercalled the frontier compute boom by nearly an order of magnitude versus the public 2025 median.
AI company investment
2023: Pro $195B, Public $282B. 2025: Pro $340B, Public $489B.
OWID’s AI investment series is narrow external private-market funding and explicitly excludes Big Tech internal R&D, capex, public-sector funding, and publicly traded firms; Stanford’s 2026 AI Index says private investment grew 127.5% in 2025 and was 60% of total corporate AI investment.
Directionally right, hard to score. On broad corporate investment, the public forecast may be close. On narrow private-company external funding, the forecasts look high.
Global fertility rate
2023: Pro 2.29, Public 2.23. 2025: Pro 2.24, Public 2.20.
OWID says global total fertility was 2.3 children per woman in 2023; UN’s 2024 fertility report gives a 2024 estimate of 2.2.
Good. The public 2023 forecast was a little low; the 2025 forecast is probably close, though final 2025 data are not fully settled.
World population
2023: Pro 8.04B, Public 8.06B. 2025: Pro 8.18B, Public 8.21B.
UN WPP 2024 gives estimates and projections through the present; World Bank summarized 2025 as about 8.2B people.
Good. Slightly low, but well within the range you would expect for this kind of forecast.
Same-sex marriage legal share
2023: Pro 15.5%, Public 15.3%. 2025: Pro 18.0%, Public 16.0%.
OWID says 1.5B people, about one in five globally, lived where same-sex marriage was legal in 2025; Thailand and Liechtenstein were added in 2025.
Good, especially Pros. Their expectation of modest continued expansion was right.
OECD trust in government
2023: Pro 42%, Public 43%. 2025: Pro 42%, Public 41%.
OECD’s 2023 survey found 39% with high or moderately high trust in national government across surveyed OECD countries. OWID’s Gallup-based OECD series is updated only through 2024.
Slightly high, not disastrously. This is not yet a clean 2025 score.
Per-capita primary energy consumption
2023: Pro 21,200 kWh, Public 21,100. 2025: Pro 21,600, Public 21,400.
Energy Institute reported global primary energy consumption at 620 EJ in 2023; with world population around 8.09B, that implies about 21,300 kWh/person. OWID’s series is updated through 2024.
Very good for 2023. 2025 not cleanly scoreable yet.
Things that are still annoyingly unresolved
Several forecasts are “past” in calendar time but not actually resolvable in the official-data sense. This matters because OWID-style indicators often update with a one- to three-year lag, and some of the underlying data series have changed methodology.
Metric
Status
Chickens slaughtered for meat
The forecast was 75.9B/74.7B for 2023 and 79.6B/78.2B for 2025. But global animal-slaughter data through FAOSTAT/OWID lag badly; OWID’s recent public material still discusses 2022 levels around 83B total land animals, mostly chickens. I would not score 2025 yet.
Cost of sequencing a whole human genome
Forecasts were $357/$303 for 2023 and $244/$191 for 2025. NHGRI’s canonical cost series explains the methodology and downloadable table, but the page still points to “Sequencing Costs 2022,” so this is not cleanly resolvable from the original benchmark.
Global poverty below $10/day in 2011 international dollars
Forecasts were about 4.66B in 2023 and 4.61B in 2025. World Bank poverty lines and PPP vintages have since shifted, making this surprisingly nontrivial to score without reconstructing the exact old line and PPP basis.
G7 life expectancy
Forecasts were 81.4/81.2 in 2023 and 81.7/81.1 in 2025. OECD data show recovery from COVID mortality, with OECD average life expectancy at 81.1 in 2023 and provisional 2024 increases, but I did not find a clean population-weighted G7 actual for 2025.
Global homicide rate
Forecasts were 5.24/5.32 in 2023 and 5.14/5.26 in 2025. World Bank’s homicide metadata runs through 2023, and the 2023 value looks close to their forecasts, but 2025 is not available.
US GDP per hour worked
Forecasts were $77.60/$76.60 in 2023 and $79.60/$78.40 in 2025. OWID’s Penn World Table source is updated only through 2023, and the U.S. BLS productivity series is not the same dollar-level metric. Directionally, 2023-24 U.S. productivity growth was stronger than many expected.
GDP per-capita growth
Forecasts were 3.1%/2.8% in 2023 and 2.8%/2.6% in 2025. OWID/World Bank annual GDP-per-capita growth is updated through 2024, while the IMF projected 2025 global GDP growth at 3.2%. Since population growth was under 1%, 2025 real per-capita growth was probably lower than the Metaculus medians, but I would score this only after the final World Bank/OWID series updates.
My read
The forecasts were not embarrassing. The strongest thing about them is that the “boring” forecasts mostly worked. Population, fertility, nuclear stockpiles, CO₂, energy use, space objects, and same-sex marriage legality were all broadly on track. The forecasters were doing something reasonable: anchor to recent series, adjust for obvious shocks, and avoid inventing too much discontinuity for 2023 or 2025.
The misses are more interesting. They overexpected institutional continuity in democracy, overexpected official supercomputer benchmark growth, badly overexpected terrorism fatalities, and underexpected the scale of AI training runs. That combination is telling: Metaculus was good at smooth continuation and at already-visible deployment trends, but weaker when the outcome depended on classification thresholds, geopolitical/security-event volatility, or private frontier-lab spending.
The AI cases split in opposite directions. For training compute, the 2025 forecast was too low because the post-ChatGPT capex frenzy and frontier-lab race got hotter than the early-2023 forecasters priced in. For AI investment, the answer depends on the denominator: narrow external private funding makes some forecasts look high, while broad corporate AI investment makes them look much more reasonable. The report’s own AI-investment metric is therefore less clean than it first appears. (Our World in Data)
Overall, I’d give the 2023/2025 track record something like B or B-minus if judged across all 20 topics, and B-plus on slow-moving OWID-style metrics. The important caveat is that this is not a formal scoring exercise over full distributions. It is a median-vs-actual audit, and for about a third of the metrics the “actual 2025 value” is still not fully published or has become methodologically slippery.
Opus 4.8 (max): “excellent on everything that was structurally predictable and wrong in exactly the places where the institutional consensus they anchored on was also wrong — which means the near-term scorecard mostly grades the 2022–23 consensus, not the forecasters’ distinctive skill”
My plan: this is fundamentally a forecasting-calibration question, and per the analyse skill I should earn the escalation — spend search budget disproportionately on the metrics where a regime change could have broken the forecast (AI compute, AI investment, supercomputers, CO2, space launches, same-sex marriage), and merely spot-check the slow-moving demographic/physical series (population, fertility, nuclear, sequencing cost) that 1–3-year extrapolation almost always nails. The information is in the surprises, not the confirmations.
I have what I need across all 20 Pro-forecasted metrics. One honesty caveat before the verdict: the report gives me the forecasters’ described reasoning and some point numbers, but the precise Pro medians and interquartile ranges live in charts I can’t see. So for the clear cases (AI investment ~6×, terrorism +22% vs. a predicted decline, CO₂ records vs. “at peak”) the direction is unambiguous regardless of the exact median; for the “hit” calls, “landed in range” is my inference from the described forecast sitting close to the actual, not a certified percentile. I flag where that matters.
Scorecard — the 20 Pro-forecasted metrics (2023 & 2025)
Sorted most-informative first (misses, then measurement problems, then the quiet hits).
Metric
Near-term forecast (as described)
What actually happened
Verdict
AI training compute
Doubling ~10 mo (≈2.3×/yr)
GPT-4, first model >1e25 FLOP, March 2023; Gemini Ultra ~5e25 by Dec 2023; ~1e26 by late-2025, ~4.7×/yr
Miss (under) — real growth ~2× their rate
AI investment
“Divided,” modest growth, some expecting a slowdown
From $91.9B private in 2022 → $252.3B corporate in 2024 → $581.7B corporate / $344.7B private in 2025
Miss (under) — biggest miss in the set
Terrorism fatalities
Continue to decline in 2023
Rose 22% to 8,352 in 2023, highest since 2017 (Oct 7 + Sahel); declined again in 2024
Miss — black-swan (Oct 7); they’d flagged it as noisy
GDP growth per capita
Poor 2023, recession in many advanced economies
Soft landing; the feared recession didn’t arrive; world grew ~2.95% in 2023
Miss (too pessimistic) — anchored on IMF/WB gloom
Governing thought: at 1–3 years, these forecasts were excellent on everything that was structurally predictable and wrong in exactly the places where the institutional consensus they anchored on was also wrong — which means the near-term scorecard mostly grades the 2022–23 consensus, not the forecasters’ distinctive skill.
Three claims hold it up:
The misses are not random — they cluster into three buckets. Two of the four (AI compute, AI investment) are the same phenomenon: an exponential the entire analyst world under-priced in February 2023. One (GDP growth) and the pessimism lean visible in trust-in-government and the CO₂-”peak” call are the opposite error — over-weighting the salient 2022–23 gloom (the recession everyone forecast and that never came, the IEA’s perennial imminent-peak). One (terrorism) is a genuine black swan — Oct 7 — that they’d explicitly flagged as an un-forecastable long-tail series. The misses are macro-regime and tail calls; the hits are inertial series.
The AI under-call is the headline, and it’s worse than it looks because of when they forecast. The report published Feb 1 2023 — ChatGPT was two months old, GPT-4 six weeks away. They were staring directly at the inflection and still used a ~2.3×/yr compute trend (realized: ~4.7×/yr) and expected AI investment to plateau (realized: a roughly 6× run from $92B to ~$580B corporate in three years). Even a charitable reading — say their median priced a 2× investment rise by 2025 — still lands a full ~3–6× short. Pricing an exponential is hard even when it’s already underway.
Where the consensus was good, they were superb — and that’s most of the list. Population to within a rounding error, fertility, nuclear totals, energy, life expectancy, trust, same-sex-marriage share: all landed. These are the series where the UN/FAS/Energy-Institute projections are themselves strong, and the forecasters correctly didn’t fight them. Notably, the space-launch question — the one you pushed back on nextbigfuture about in that very thread — resolved at 4,510 in 2025, inside the pros’ 3,900–9,200 band; his 50k–500k was ~1–2 OOM high, and your “~1 OOM below his number” instinct bracketed the truth.
The premise worth challenging
Your question was “how are the forecasts doing.” The sharper version is: at this horizon, the question barely discriminates. The thing that makes the Metaculus community distinctive — willingness to entertain exponentials, regime changes, and fat tails — is exactly what 1–3-year medians suppress, because over short windows the aggregate reverts to consensus. So a strong near-term scorecard (which this is) is mostly evidence that the consensus was right, and tells you almost nothing about whether the community’s tail-aware reasoning adds value. The one place that distinctive reasoning should have paid off — taking the AI exponential seriously — is precisely where they most under-used it.
Two implications:
The actual test of this tournament hasn’t run yet. Its entire design rationale was the 30/52/122-year horizons — where the OWID note explicitly says the community’s willingness to price black swans makes it less conservative than traditional sources. That bet is unfalsified until those resolve. Grading 2023/2025 is grading the warm-up.
For your AGI-timelines work specifically: the cleanest signal in this whole exercise is a base rate — sophisticated forecasters under-predict AI scaling even when watching it happen. Compute and capital both ran several-fold ahead of well-reasoned medians over three years. That’s a thumb on the scale toward the aggressive end of current AI forecasts, not the modal one.
Caveats on this grading
I’m scoring against the report’s prose descriptions, not the chart medians/IQRs — fine for the unambiguous calls, weaker for certifying “in range” on the hits. Three metrics (chickens, homicide, poverty) can’t be graded yet on their ~2-year-lagged sources. Two resolution problems are real and worth internalizing as a forecasting lesson in their own right: the genome-cost series simply stopped existing (NHGRI froze in 2022), and the “objects launched” and “nuclear” counts get revised upward for years after the fact, so any single-year resolution is a moving target. A forecast can be undone by the measurement pipeline rather than the world.
As an aside, I also asked the models to assess each other’s assessments above:
GPT-5.5 assesses Opus 4.8: “conceptual diagnosis is strong, but its scorecard is materially sloppier than mine on several important rows. The big difference is that it often grades the report’s prose instead of the report’s forecast table… it incorrectly treats fastest supercomputer and liberal democracy population as hits, and overstates the AI investment miss by mixing definitions”
I think the other model’s conceptual diagnosis is strong, but its scorecard is materially sloppier than mine on several important rows. The big difference is that it often grades the report’s prose instead of the report’s forecast table. That makes it more rhetorically satisfying, but less reliable as an audit.
Its best line is this: near-term OWID forecasts mostly test whether forecasters anchored well to institutional baselines, while the more distinctive Metaculus skill, pricing exponentials and weird tails, is barely visible at 1–3 years. I agree with that. It also correctly spots the AI compute undercall, the AI boom, the terrorism miss, the no-recession surprise, and the broader “inertial variables did well” pattern. Its own caveat says it is scoring against prose descriptions rather than exact medians/IQRs, and that caveat matters a lot.
Where I disagree: its tally is too generous. It says roughly 12 clean hits, 1 split, 4 misses, 1 measurement failure, 3 lagged-not-gradable. I would revise that downward, mainly because it incorrectly treats fastest supercomputer and liberal democracy population as hits, and overstates the AI investment miss by mixing definitions.
Quick comparison
Issue
Other model
My view
Overall framing
Strong: “near-term mostly grades consensus.”
I agree. This is probably the most important meta-point.
Use of exact forecast values
Weak: often uses prose/ranges.
I’d privilege the report tables and resolution definitions.
AI training compute
Correctly says miss low.
Agree, and probably the most important technical miss.
AI investment
Says biggest miss under.
Too strong. Depends on whether resolution uses narrow external private investment or broad corporate AI spending.
Supercomputer FLOPS
Calls it a hit.
I think this is wrong. Forecast was ~4.1–4.5e18 for 2025; actual TOP500 was 1.809e18.
Liberal democracies
Calls it directionally a hit.
I think this is wrong. Forecast was ~1.0–1.1B for 2025; V-Dem says 0.6B at end-2025.
Terrorism
Correctly says miss.
Agree, but it underplays the magnitude. Forecast was 19–20k; actual 2025 was 5,582.
CO₂
Says number hit, peak narrative premature.
Agree. That is a nice distinction.
Space objects
Calls hit.
Agree. Pros were impressively close.
Demographics
Calls population/fertility hits.
Mostly agree, though its fertility actuals are a little messy depending on UN/OWID projection vintage.
The most important factual disagreements
1. Supercomputer FLOPS: the other model is wrong to call this a hit
The report table gives the 2025 fastest-supercomputer forecasts as 4.1 × 10¹⁸ FLOP/s for Pros and 4.5 × 10¹⁸ for the public. The TOP500 November 2025 result was El Capitan at 1.809 exaFLOP/s. That is not a hit. It is a miss high by about 2.3–2.5×. The other model appears to have inferred from the prose that forecasters expected “sub-2 EF” by 2025, but the table does not say that. It says ~4 EF.
This is not a trivial bookkeeping dispute. It affects the 2032 adjustment. If you call 2025 a hit, you might leave the 2032 TOP500 forecast alone. If you call it a miss high, you should lower 2032, which is what I did.
2. AI investment: the other model overclaims the miss
The report table gives 2025 AI investment forecasts of $340B Pro and $489B public. The other model says actuals were “$581.7B corporate / $344.7B private” and therefore calls this the biggest miss under. But OWID’s relevant AI-investment chart says it focuses on external private-market investment and excludes internal corporate R&D, capex, public-sector funding, and public companies. On that narrower definition, $344.7B private makes the Pro forecast of $340B look almost exactly right, while the public forecast of $489B looks high. On broad corporate AI investment, the public forecast may be low.
So the better judgment is: direction right, metric definition decisive. The other model’s “biggest miss under” conclusion is only defensible if you resolve against broad corporate AI investment, not against the narrower OWID-style private-investment series.
3. Liberal democracies: the other model misses the U.S./V-Dem discontinuity
The report forecasted 1.00B Pro and 1.10B public people living in liberal democracies in 2025. V-Dem’s 2026 Democracy Report says only 7% of the world population, about 0.6B people, lived in liberal democracies at the end of 2025, and explicitly says the U.S. lost its long-term liberal-democracy status for the first time in over 50 years.
The other model calls this a directional hit, saying V-Dem shows a slow decline around ~1.0B. That looks stale or wrong. This is one of the biggest substantive misses in the whole set, because the forecast’s own resolution criteria used OWID or an alternative credible source for people living in liberal democracies, and the report itself flagged U.S. 2024 turmoil as a possible reason the U.S. might lose liberal-democracy status.
4. Terrorism: it gets the sign right but underplays the magnitude
The report forecasted 19k–20k terrorism fatalities for both 2023 and 2025. The Global Terrorism Index 2026 says 2025 deaths were 5,582, down 28% from 2024. It also says 2023 deaths spiked to 8,654, largely due to October 7, before falling again. So yes, the forecasters missed, but the other model frames the miss around “2023 rose 22%” instead of the more important fact that the median forecast was about 2.2× too high in 2023 and about 3.4× too high in 2025.
The right lesson is subtler than “black swan happened.” The October 7 shock pushed deaths upward, but even with that shock, the forecast was still far too high. The forecasters seem to have over-anchored to older high-terrorism baselines or to right-tail concerns.
5. CO₂: I like its distinction
Here I think the other model improves on my initial phrasing. The report’s Pro forecasts were 37.7 Gt for 2023 and 38.1 Gt for 2025, while the public forecasts were 36.4 and 36.2. The IEA says 2025 energy-related CO₂ reached a new high of nearly 38.4 Gt. So the Pro number was very good, the public number was too low, and the qualitative “at or near peak” narrative looks premature.
That “number right, story too early” framing is exactly the right way to score it.
Where the other model is better than mine
It is better than my first answer in three ways.
First, it has a cleaner theory of the errors: inertial variables good, regime/tail variables bad. That’s a useful compression. Population, space objects, fertility, energy, nuclear stockpiles, CO₂ Pro forecasts, and same-sex marriage mostly follow observable short-run trends. AI compute, AI investment, terrorism, GDP surprises, and liberal-democracy classification are more thresholdy or regime-sensitive.
Second, it correctly emphasizes that a 1–3-year scorecard is a weak test of the tournament’s real ambition. The tournament’s rationale was about 10-, 30-, and 100-year uncertainty. In the short run, even excellent forecasters often just become careful aggregators of UN, IMF, FAS, OECD, IEA, TOP500, OWID, and similar baselines. That is valuable, but it is not the same as proving far-future forecasting skill.
Third, it is right to treat AI as the most informative miss. The report’s 2025 AI training-compute forecasts were 2.9e25 Pro and 6.6e25 public, while Epoch now puts the largest known training run at around 5e26 FLOP, and notes that more than 30 models had crossed the 1e25 FLOP scale by June 2025.
Where mine is better
Mine is more careful about resolution definitions. That matters because these are not abstract “what happened in the world?” questions. They are “what number will a specified dataset or credible replacement source report?” questions. The other model sometimes slides from the metric to a nearby intuitive proxy: corporate AI spend instead of external private investment; prose expectations instead of table medians; total nuclear inventory instead of military stockpile-style series; broad democratic decline instead of the specific V-Dem/OWID liberal-democracy population count.
Mine is also less willing to count “directionally on track” as a hit when the forecast was numerical and the relevant year is missing or the actual value is far from the median. The other model’s “hit” threshold is too forgiving. A forecast of 1.0B liberal-democracy population resolving at 0.6B is not a hit just because the direction was “slightly down.” A forecast of 4.1e18 FLOP/s resolving at 1.809e18 is not a hit just because the prose said the leap might be delayed.
Revised synthesis after comparing both
I’d now summarize the 2023/2025 track record like this:
Cluster
My revised judgment
Very good
Space objects, world population, CO₂ Pro forecast, fertility, per-capita energy, nuclear stockpiles roughly, same-sex marriage roughly
Mixed or definition-sensitive
AI investment, genome sequencing, GDP per capita growth, OECD trust, G7 life expectancy, U.S. productivity
Clear misses
AI training compute low, terrorism high, liberal-democracy population high, fastest-supercomputer FLOPS high, public CO₂ too low
Not properly scoreable yet
Chickens, homicide, poverty, and several non-Pro/public-only OWID series with lagged data
So I’d move my earlier “B/B-minus overall” slightly downward for the fully scoreable subset, maybe closer to B-minus or C-plus if you grade medians strictly, while preserving B-plus-ish for slow-moving macro/demographic variables. The other model’s “12 clean hits” is too kind because two of those clean hits are not clean hits.
The best combined takeaway is this: Metaculus was good where disciplined base-rate anchoring was enough, but not clearly better than careful institutional extrapolation at 1–3 years. It underpriced AI scale-up, overestimated terrorism, missed a threshold political-classification shock, and overestimated TOP500 growth. For 2032, the right way to beat the original forecasts is to mostly leave the boring variables alone and make concentrated adjustments on the variables where these short-run misses reveal a wrong model class.
Opus 4.8 assesses GPT-5.5: “strong assessment — better than mine in several specific places because it has the actual Pro/Public medians I lacked — and corrects a real error I made. But also one significant flaw that fooled it, for the terrorism verdict (and, in a different way, fooled me too), and one overstatement, for supercomputer FLOPS”
This is a strong assessment, and in a few places it’s better than mine — including one spot where it cleanly corrects a real error I made. Before I adjudicate the places where we disagree (terrorism, supercomputers, liberal democracies), let me verify the two load-bearing facts behind its sharpest claims, since the whole comparison turns on who’s actually right.
V-Dem confirms the model: just 7% of the world (about 600 million people) lived in liberal democracies at the end of 2025 — the lowest in over 50 years — with the drop driven largely by the USA losing liberal-democracy status. So the model caught a real miss I was too soft on. Now the terrorism series, where I suspect the model’s “2.3× miss” may be a measurement artifact.
Both fact-checks came back clean, and they’re decisive. Verdict up front: this is a strong assessment — better than mine in several specific places because it has the actual Pro/Public medians I lacked — and it corrects a real error I made. But it also contains one significant flaw that fooled it (and, in a different way, fooled me too), and one overstatement. Combining its numbers with my provenance-checking beats either of us alone.
A reassuring meta-point first: the medians it reports are almost certainly authentic, not hallucinated. The tell is internal consistency with the underlying data series — e.g. its terrorism median (~20k) matches the GTD baseline the question actually resolves on, and its AI-investment Pro median ($340B) matches the narrow-private actual almost exactly. A model fabricating numbers wouldn’t reliably reproduce series-specific baselines it had no other reason to know. That said, the entire numerical scorecard rests on those transcriptions, so they’re the one thing worth spot-checking against the report PDF.
Where it’s right and corrects me
AI investment — it catches a genuine error of mine, and this is the biggest single correction. In turn 1 I called this “the biggest under-forecast” by comparing the forecast to ~$582B of broad corporate AI investment. That was the wrong denominator. The OWID metric is narrow external private funding for privately-held AI companies, which excludes Big Tech capex, public firms, and public-sector money — and on that series the 2025 actual is ~$345B. So the Pro median of $340B was nearly exact, and the Public’s $489B was ~40% high. The model’s framing — “directionally right, hard to score, looks high on narrow private funding” — is correct and mine was wrong. And it’s a more interesting result than my version: forecasting $340B in Feb 2023, off a falling $92B base, means the Pros actually did price the investment boom on this metric. I had the story backwards.
Liberal democracies — it’s right, I was too lenient. I waved this through as a “directional hit (noisy metric).” It’s a clear miss high: forecast ~1.0B, actual ~0.6B (7% of the world) at end-2025, the lowest in 50+ years, largely because the USA lost its liberal-democracy status — the first time in more than 50 years, falling to electoral democracy. The model names the mechanism precisely (“overexpected institutional continuity in democracy”). I undersold a real miss.
Nuclear — cleaner than mine. Its medians (~9.4–9.5k) make clear the question resolves on the military stockpile (~9,600), which is exactly right and resolves the total-vs-stockpile ambiguity I left dangling in turn 1.
Where I’d push back on it (and where I was also wrong)
Terrorism — its sharpest-looking verdict is its biggest flaw, and my turn-1 take was wrong too. The model calls this a “bad miss high… overestimated by ~2.3× in 2023 and ~3.4× in 2025.” But it compared the forecast to the wrong database. The OWID metric resolves on the Global Terrorism Database (GTD), on which ~20,000 people died from terrorism in 2019 — so the Pro median of ~19–20k isn’t an overestimate at all; it’s a sensible extrapolation of the actual resolution series. Two problems then compound:
The GTD is discontinued. As of 2025 the GTD is closed, and the data only extend through 2020/2021. There is no native 2023 or 2025 value — this question is unresolvable, exactly like the genome-cost case (NHGRI frozen at 2022). It belongs in the not-scoreable bucket, not the miss column.
The “actual” it used is a different, lower-counting database. The 8,352 / 5,582 figures are from the GTI, which systematically counts fewer deaths than the GTD — GTI put 2019 at 13,826 versus GTD’s ~20,000 for the same year. So a chunk of the apparent “2.3× miss” is just the GTD→GTI level shift, not forecaster error.
And the kicker: on the GTI proxy, deaths actually fell to 5,582 in 2025, the lowest since 2007 — so the forecasters’ qualitative “continue to decline” call was directionally vindicated by 2025. My own turn-1 framing (“predicted decline, got an Oct-7 spike to 8,352, miss”) was also wrong: I used the GTI number, missed that the GTD is the dead resolution series, and treated a one-year blip as the verdict. We both mis-scored this in opposite directions; the correct call is “unresolvable / not cleanly scoreable, decline directionally right.”
Supercomputer FLOPS — “clear miss high” overstates it. Forecast 4.1e18 vs actual 1.809e18 is 2.3× — but that’s 0.36 of an order of magnitude on a log-scale exponential metric. On the same log-error basis the model (correctly) uses to call AI-compute a miss (~0.9 OOM low), a 0.36-OOM overshoot is a mild overshoot, not a “clear miss.” The directional read (“expected the exascale race to accelerate faster than the TOP500 benchmark did”) is right; the severity is dialed too high. There’s also a metric-decoupling issue: the frontier AI clusters (xAI, Meta) don’t run the LINPACK benchmark, so TOP500′s #1 increasingly fails to measure the thing that matters.
Where we agree
The structural/slow-moving hits (population, fertility, CO₂ level, energy, space objects, same-sex marriage); AI training compute as a genuine miss-low (its “~5e26 Grok 4 vs 6.6e25 Public median” and my “~1e26+ vs an extrapolated median” both land at ~1 OOM under); and the laggy third that simply isn’t scoreable yet (chickens, poverty, homicide, genome cost, productivity, GDP growth). On that last bucket the model is actually better organized than I was — its dedicated table with specific reasons (PPP-vintage shifts, NHGRI’s 2022 freeze, Penn World Table ending at 2023) is cleaner than my prose.
The synthesis that beats both assessments
Here’s the reconciled picture of the contested calls:
Metric
Model’s verdict
My turn-1 verdict
Reconciled (what’s actually true)
AI investment
Directionally right, looks high (narrow series)
“Biggest under-forecast”
Model right. Pro $340B ≈ exact vs ~$345B; I used the wrong denominator
Liberal democracies
Clear miss high (US reclassified)
“Directional hit, noisy”
Model right. ~1.0B forecast vs 0.6B actual; I was too lenient
Terrorism
Bad miss high (~2.3×)
Miss (Oct-7 spike)
Both wrong. GTD discontinued + GTI mismatch → unresolvable; decline directionally right
Supercomputer FLOPS
Clear miss high
Hit
Between us. Mild 0.36-OOM overshoot + metric decoupling
Nuclear
Very good
Hit (waffled on series)
Model cleaner (military stockpile, ~9.6k)
AI compute
Miss low (~1 OOM)
Miss low
Agree
Lay these side by side and a pattern jumps out that neither assessment states: of the model’s four headline “misses,” only AI compute is a clean real-world forecasting error. Terrorism is a dead data source plus a database mismatch; liberal democracies is a V-Dem reclassification of the US (a classification event, unforeseeable in Feb 2023, not a smooth trend the forecasters misread); supercomputers is a 0.36-OOM overshoot on a metric that’s decoupling from reality. Three of the four are measurement/definition artifacts, not failures to read the world.
That has two consequences. First, it means the forecasters look better than the model’s B/B-minus — closer to B-plus/A-minus on real-world calibration, with the honest caveat that roughly a third of the metrics are measurement-fragile. The model’s grade is slightly too harsh because it scored three measurement artifacts as forecasting misses. Second, it sharpens my turn-1 thesis: I argued the 1–3yr exercise mostly grades the consensus they anchored on; the better statement now is that the dominant failure mode at this horizon isn’t bad forecasting, it’s metric fragility — series get discontinued (GTD, NHGRI), redefined (V-Dem reclassification), or measure the wrong proxy (TOP500 vs AI clusters). The one place the world itself outran them is AI compute. Everything else that “missed” is the ruler changing, not the thing being measured.
Net: the model’s assessment is genuinely good and I’d trust its numbers over my turn-1 qualitative grading wherever they conflict — it fixed my AI-investment and liberal-democracy errors outright. What I’d add to it is provenance discipline (which rescues terrorism from the miss column and softens supercomputers) and the measurement-artifact reframe. The combined verdict is more accurate than either of us produced alone, which is the whole point of running two passes.
I’ll let Opus 4.8 have the final word:
… of the model’s four headline “misses,” only AI compute is a clean real-world forecasting error. Terrorism is a dead data source plus a database mismatch; liberal democracies is a V-Dem reclassification of the US (a classification event, unforeseeable in Feb 2023, not a smooth trend the forecasters misread); supercomputers is a 0.36-OOM overshoot on a metric that’s decoupling from reality. Three of the four are measurement/definition artifacts, not failures to read the world. … the dominant failure mode at [the 1-3 year] horizon isn’t bad forecasting, it’s metric fragility — series get discontinued (GTD, NHGRI), redefined (V-Dem reclassification), or measure the wrong proxy (TOP500 vs AI clusters). The one place the world itself outran them is AI compute. Everything else that “missed” is the ruler changing.
In 2022 Metaculus ran Forecasting Our World in Data: The Next 100 Years, still the most methodologically thorough metrics-grounded whole-of-civilisation forecast of its kind I know of:
The composite perspective into the world of 2122 seems distinctly pedestrian to me
Our World in 2122
In the year 2122, the global population is expected to be around 8 or 9 billion, with a chance that number reaches as high as 20 billion due to technological and economic advancements. Lifespans are expected to be significantly longer, especially in the G7 countries, thanks to the achievement of longevity escape velocity. Reproduction is anticipated by some to primarily utilize ectogenesis technology, though some may choose to avoid it for personal or ethical reasons. The decline in population is expected to slow, but there remains uncertainty as to whether it will level off or decrease further.
On the economic front, GDP per capita is believed to be quite high and will likely continue increasing. In fact, most economies are expected to be as advanced as present-day Scandinavian countries, with energy being inexpensive and the price of most goods and services being close to zero. Productivity is forecasted to significantly increase, with advanced artificial intelligence playing a major role in GDP growth. The United States economy is anticipated to continue growing, potentially transitioning to a more service-based industry and experiencing changes in government structure. However, the concept of the United States as we know it may be fundamentally different by this time. Instead, it is possible that global governance will be implemented and the number of sovereign states will drastically decrease. Wealth distribution remains a key uncertainty — while it’s possible that a small group could control a significant portion of resources, a shift toward more equitable distribution may win out.
Even with global governance, many pressing challenges will remain. Despite drastic cuts in carbon dioxide emissions, climate change is expected to be a major concern — sea levels will rise and extreme weather events will become more frequent. Nevertheless, renewable energy is also expected to increase in parallel, with the potential for advanced technologies, like nuclear fusion, to power a significant portion of the global economy. Even though the precise global energy mix remains uncertain, transportation will undergo major electrification, with an anticipated shift towards both zero emissions and autonomous vehicles.
Technology will continue expanding at a rapid pace and play an even greater role in daily life. The use of robots is predicted to increase and revolutionary innovations like brain-computer interfaces could fundamentally change daily life as we know it. Education is therefore expected to become more personalized and virtual, as will the healthcare system. The latter will undergo many transformations as artificial intelligence opens the door for hyper-personalized treatments and precision medicine. The concept of work will change too, becoming far more flexible and remote. Leisure time is expected to increase, with virtual reality and space tourism becoming popular for many.
Despite great uncertainty and persistent challenges, 2122 is expected to be a time of great advancement and possibility.
I wondered how the first two timeframes (2023 and 2025) had done so far, so I asked the models:
GPT-5.5 (high): “pretty good on slow-moving macro and demographic variables, good on space objects and CO₂, and bad on a few discontinuous or definition-sensitive variables: terrorism, fastest-supercomputer FLOPS, people living in liberal democracies, and AI training compute”
I used the Metaculus report tables for the forecast medians, with “Pro” meaning Pro Forecaster median and “Public” meaning the Metaculus Prediction. The report says the tournament forecasted 2023, 2025, 2032, 2052, and 2122 values, and that the summary tables cover the 20 questions forecasted by both Pros and the public.
Bottom line
So far, the forecasts look pretty good on slow-moving macro and demographic variables, good on space objects and CO₂, and bad on a few discontinuous or definition-sensitive variables: terrorism, fastest-supercomputer FLOPS, people living in liberal democracies, and AI training compute. AI investment is messy: they got the direction right, but whether the forecast was high or low depends heavily on whether you use narrow private funding, broad corporate AI investment, or capex-like spending.
A rough scorecard: among the items I could actually check with reasonably current data, I’d call about 7 good or very good, 4 clear misses, and the rest not yet properly scoreable because the official annual data lag into 2026 or 2027.
Scorecard for the clearest cases
Metric
Forecast for 2023 and 2025
What happened
Verdict
Annual objects launched into space
2023: Pro 3,000, Public 2,700. 2025: Pro 4,300, Public 4,000.
OWID says 2,903 objects were launched in 2023 and 4,510 in 2025.
Very good. Pros were especially close. Public undercalled 2025 by about 11%.
Global CO₂ emissions
2023: Pro 37.7 Gt, Public 36.4 Gt. 2025: Pro 38.1 Gt, Public 36.2 Gt.
IEA put 2023 energy-related CO₂ at 37.4 Gt; Global Carbon Project projected 2025 fossil CO₂ at 38.1 Gt.
Pros basically nailed it. Public forecasts were too optimistic about emissions falling.
Fastest supercomputer FLOPS
2023: Pro 1.5e18, Public 1.3e18. 2025: Pro 4.1e18, Public 4.5e18.
Frontier was still around 1.1 exaFLOP in late 2023; El Capitan led the Nov. 2025 TOP500 at 1.809 exaFLOPS Rmax.
Clear miss high. They expected the exascale transition to accelerate faster than the TOP500 benchmark actually did.
Global nuclear warhead stockpiles
2023: Pro 9.46k, Public 9.40k. 2025: Pro 9.52k, Public 9.38k.
SIPRI estimated about 9,614 warheads in military stockpiles in Jan. 2025.
Very good. Slightly low, but right order and right direction.
Total terrorism fatalities
2023: Pro 19k, Public 20k. 2025: Pro 19k, Public 20k.
Global Terrorism Index reported 8,352 deaths in 2023, and its 2026 report gives 5,582 deaths in 2025.
Bad miss high. They overestimated by roughly 2.3x in 2023 and 3.4x in 2025.
People living in liberal democracies
2023: Pro 1.02B, Public 1.05B. 2025: Pro 1.00B, Public 1.10B.
V-Dem’s 2026 Democracy Report says only 7% of the world population, about 0.6B people, lived in liberal democracies at end-2025.
Clear miss high. This is probably the biggest substantive miss, driven by democratic backsliding and classification changes, especially the U.S. losing liberal-democracy status in V-Dem.
Greatest AI training compute
2023: Pro 6.5e24 FLOP, Public 1.3e25. 2025: Pro 2.9e25, Public 6.6e25.
Epoch’s 2026 trends page says the largest known training run was Grok 4 at around 5e26 FLOP.
Miss low. They correctly expected rapid growth, but still undercalled the frontier compute boom by nearly an order of magnitude versus the public 2025 median.
AI company investment
2023: Pro $195B, Public $282B. 2025: Pro $340B, Public $489B.
OWID’s AI investment series is narrow external private-market funding and explicitly excludes Big Tech internal R&D, capex, public-sector funding, and publicly traded firms; Stanford’s 2026 AI Index says private investment grew 127.5% in 2025 and was 60% of total corporate AI investment.
Directionally right, hard to score. On broad corporate investment, the public forecast may be close. On narrow private-company external funding, the forecasts look high.
Global fertility rate
2023: Pro 2.29, Public 2.23. 2025: Pro 2.24, Public 2.20.
OWID says global total fertility was 2.3 children per woman in 2023; UN’s 2024 fertility report gives a 2024 estimate of 2.2.
Good. The public 2023 forecast was a little low; the 2025 forecast is probably close, though final 2025 data are not fully settled.
World population
2023: Pro 8.04B, Public 8.06B. 2025: Pro 8.18B, Public 8.21B.
UN WPP 2024 gives estimates and projections through the present; World Bank summarized 2025 as about 8.2B people.
Good. Slightly low, but well within the range you would expect for this kind of forecast.
Same-sex marriage legal share
2023: Pro 15.5%, Public 15.3%. 2025: Pro 18.0%, Public 16.0%.
OWID says 1.5B people, about one in five globally, lived where same-sex marriage was legal in 2025; Thailand and Liechtenstein were added in 2025.
Good, especially Pros. Their expectation of modest continued expansion was right.
OECD trust in government
2023: Pro 42%, Public 43%. 2025: Pro 42%, Public 41%.
OECD’s 2023 survey found 39% with high or moderately high trust in national government across surveyed OECD countries. OWID’s Gallup-based OECD series is updated only through 2024.
Slightly high, not disastrously. This is not yet a clean 2025 score.
Per-capita primary energy consumption
2023: Pro 21,200 kWh, Public 21,100. 2025: Pro 21,600, Public 21,400.
Energy Institute reported global primary energy consumption at 620 EJ in 2023; with world population around 8.09B, that implies about 21,300 kWh/person. OWID’s series is updated through 2024.
Very good for 2023. 2025 not cleanly scoreable yet.
Things that are still annoyingly unresolved
Several forecasts are “past” in calendar time but not actually resolvable in the official-data sense. This matters because OWID-style indicators often update with a one- to three-year lag, and some of the underlying data series have changed methodology.
Metric
Status
Chickens slaughtered for meat
The forecast was 75.9B/74.7B for 2023 and 79.6B/78.2B for 2025. But global animal-slaughter data through FAOSTAT/OWID lag badly; OWID’s recent public material still discusses 2022 levels around 83B total land animals, mostly chickens. I would not score 2025 yet.
Cost of sequencing a whole human genome
Forecasts were $357/$303 for 2023 and $244/$191 for 2025. NHGRI’s canonical cost series explains the methodology and downloadable table, but the page still points to “Sequencing Costs 2022,” so this is not cleanly resolvable from the original benchmark.
Global poverty below $10/day in 2011 international dollars
Forecasts were about 4.66B in 2023 and 4.61B in 2025. World Bank poverty lines and PPP vintages have since shifted, making this surprisingly nontrivial to score without reconstructing the exact old line and PPP basis.
G7 life expectancy
Forecasts were 81.4/81.2 in 2023 and 81.7/81.1 in 2025. OECD data show recovery from COVID mortality, with OECD average life expectancy at 81.1 in 2023 and provisional 2024 increases, but I did not find a clean population-weighted G7 actual for 2025.
Global homicide rate
Forecasts were 5.24/5.32 in 2023 and 5.14/5.26 in 2025. World Bank’s homicide metadata runs through 2023, and the 2023 value looks close to their forecasts, but 2025 is not available.
US GDP per hour worked
Forecasts were $77.60/$76.60 in 2023 and $79.60/$78.40 in 2025. OWID’s Penn World Table source is updated only through 2023, and the U.S. BLS productivity series is not the same dollar-level metric. Directionally, 2023-24 U.S. productivity growth was stronger than many expected.
GDP per-capita growth
Forecasts were 3.1%/2.8% in 2023 and 2.8%/2.6% in 2025. OWID/World Bank annual GDP-per-capita growth is updated through 2024, while the IMF projected 2025 global GDP growth at 3.2%. Since population growth was under 1%, 2025 real per-capita growth was probably lower than the Metaculus medians, but I would score this only after the final World Bank/OWID series updates.
My read
The forecasts were not embarrassing. The strongest thing about them is that the “boring” forecasts mostly worked. Population, fertility, nuclear stockpiles, CO₂, energy use, space objects, and same-sex marriage legality were all broadly on track. The forecasters were doing something reasonable: anchor to recent series, adjust for obvious shocks, and avoid inventing too much discontinuity for 2023 or 2025.
The misses are more interesting. They overexpected institutional continuity in democracy, overexpected official supercomputer benchmark growth, badly overexpected terrorism fatalities, and underexpected the scale of AI training runs. That combination is telling: Metaculus was good at smooth continuation and at already-visible deployment trends, but weaker when the outcome depended on classification thresholds, geopolitical/security-event volatility, or private frontier-lab spending.
The AI cases split in opposite directions. For training compute, the 2025 forecast was too low because the post-ChatGPT capex frenzy and frontier-lab race got hotter than the early-2023 forecasters priced in. For AI investment, the answer depends on the denominator: narrow external private funding makes some forecasts look high, while broad corporate AI investment makes them look much more reasonable. The report’s own AI-investment metric is therefore less clean than it first appears. (Our World in Data)
Overall, I’d give the 2023/2025 track record something like B or B-minus if judged across all 20 topics, and B-plus on slow-moving OWID-style metrics. The important caveat is that this is not a formal scoring exercise over full distributions. It is a median-vs-actual audit, and for about a third of the metrics the “actual 2025 value” is still not fully published or has become methodologically slippery.
Opus 4.8 (max): “excellent on everything that was structurally predictable and wrong in exactly the places where the institutional consensus they anchored on was also wrong — which means the near-term scorecard mostly grades the 2022–23 consensus, not the forecasters’ distinctive skill”
My plan: this is fundamentally a forecasting-calibration question, and per the analyse skill I should earn the escalation — spend search budget disproportionately on the metrics where a regime change could have broken the forecast (AI compute, AI investment, supercomputers, CO2, space launches, same-sex marriage), and merely spot-check the slow-moving demographic/physical series (population, fertility, nuclear, sequencing cost) that 1–3-year extrapolation almost always nails. The information is in the surprises, not the confirmations.
I have what I need across all 20 Pro-forecasted metrics. One honesty caveat before the verdict: the report gives me the forecasters’ described reasoning and some point numbers, but the precise Pro medians and interquartile ranges live in charts I can’t see. So for the clear cases (AI investment ~6×, terrorism +22% vs. a predicted decline, CO₂ records vs. “at peak”) the direction is unambiguous regardless of the exact median; for the “hit” calls, “landed in range” is my inference from the described forecast sitting close to the actual, not a certified percentile. I flag where that matters.
Scorecard — the 20 Pro-forecasted metrics (2023 & 2025)
Sorted most-informative first (misses, then measurement problems, then the quiet hits).
Metric
Near-term forecast (as described)
What actually happened
Verdict
AI training compute
Doubling ~10 mo (≈2.3×/yr)
GPT-4, first model >1e25 FLOP, March 2023; Gemini Ultra ~5e25 by Dec 2023; ~1e26 by late-2025, ~4.7×/yr
Miss (under) — real growth ~2× their rate
AI investment
“Divided,” modest growth, some expecting a slowdown
From $91.9B private in 2022 → $252.3B corporate in 2024 → $581.7B corporate / $344.7B private in 2025
Miss (under) — biggest miss in the set
Terrorism fatalities
Continue to decline in 2023
Rose 22% to 8,352 in 2023, highest since 2017 (Oct 7 + Sahel); declined again in 2024
Miss — black-swan (Oct 7); they’d flagged it as noisy
GDP growth per capita
Poor 2023, recession in many advanced economies
Soft landing; the feared recession didn’t arrive; world grew ~2.95% in 2023
Miss (too pessimistic) — anchored on IMF/WB gloom
CO₂ emissions
Number ~37–38 Gt; “at or near peak”
36.8 Gt (2023) → 37.4 Gt (2024), record, “no sign of peak”
Split — number nailed, peak narrative premature
Genome sequencing cost
~15%/yr decline; $100 genome ~2030
NHGRI/OWID series frozen at 2022; commercially $200 (Illumina, 2024) and $100 (Ultima, 2024)
Unresolvable — data pipeline died; tech beat the timeline
Objects launched to space
2025 range 3,900–9,200
2,903 in 2023; 4,510 in 2025
Hit — actual in lower part of range
Fastest supercomputer
Sub-2 EF in 2023; no 5–10× leap by 2025
Frontier ~1.2 EF (2023); El Capitan 1.809 EF (Nov 2025)
Hit
World population
~8.05B (2023), ~8.19B (2025), UN medium
~8.0B in 2023, ~8.2B 2025
Hit — essentially dead-on
Global fertility
Keep declining, slightly below UN
2.3 in 2023; 2.40 in 2025
Hit
Nuclear stockpiles
~recent level, slight change, China up
12,241 total (Jan 2025), down slightly; China 500→600
Hit
Same-sex marriage share
Modest rise; Thailand by ~2032
38 countries, 1.5B people, 20% by 2026; Thailand Jan 2025
Hit — Thailand arrived ~7 yrs early
US productivity
~2.5%/yr, above the ~1.1% pre-pandemic trend
+1.6% (2023), +2.3% (2024), ~2.1% (2025)
Hit — slightly high, acceleration call right
Per-capita energy
Slow increase
~77 GJ/person (2023); record totals, modest per-capita rise
Hit
Life expectancy (G7)
Slight rise; US continues to drag
82.0 avg (2024); US lowest at 78.9
Hit
OECD trust in govt
40–50%, slight dip
48% in 2023
Hit — mild pessimism lean
People in liberal democracies
Roughly flat / slight decline
V-Dem shows continued slow decline (~1.0B; classification contested)
Hit (directional) — metric is itself noisy
Chickens slaughtered
Continued growth
FAO data lags ~2 yrs; 2023 record, rising; 2025 N/A
Lagged — directionally on track
Global homicide rate
Continued slow decline
UNODC lags ~2 yrs; not yet resolvable
Lagged
Global poverty (<$10/day)
Continued decline
World Bank lags ~2–3 yrs; not yet resolvable
Lagged
Tally: roughly 12 clean hits, 1 split (CO₂), 4 misses, 1 measurement failure, 3 not-yet-gradable.
The pattern
Governing thought: at 1–3 years, these forecasts were excellent on everything that was structurally predictable and wrong in exactly the places where the institutional consensus they anchored on was also wrong — which means the near-term scorecard mostly grades the 2022–23 consensus, not the forecasters’ distinctive skill.
Three claims hold it up:
The misses are not random — they cluster into three buckets. Two of the four (AI compute, AI investment) are the same phenomenon: an exponential the entire analyst world under-priced in February 2023. One (GDP growth) and the pessimism lean visible in trust-in-government and the CO₂-”peak” call are the opposite error — over-weighting the salient 2022–23 gloom (the recession everyone forecast and that never came, the IEA’s perennial imminent-peak). One (terrorism) is a genuine black swan — Oct 7 — that they’d explicitly flagged as an un-forecastable long-tail series. The misses are macro-regime and tail calls; the hits are inertial series.
The AI under-call is the headline, and it’s worse than it looks because of when they forecast. The report published Feb 1 2023 — ChatGPT was two months old, GPT-4 six weeks away. They were staring directly at the inflection and still used a ~2.3×/yr compute trend (realized: ~4.7×/yr) and expected AI investment to plateau (realized: a roughly 6× run from $92B to ~$580B corporate in three years). Even a charitable reading — say their median priced a 2× investment rise by 2025 — still lands a full ~3–6× short. Pricing an exponential is hard even when it’s already underway.
Where the consensus was good, they were superb — and that’s most of the list. Population to within a rounding error, fertility, nuclear totals, energy, life expectancy, trust, same-sex-marriage share: all landed. These are the series where the UN/FAS/Energy-Institute projections are themselves strong, and the forecasters correctly didn’t fight them. Notably, the space-launch question — the one you pushed back on
nextbigfutureabout in that very thread — resolved at 4,510 in 2025, inside the pros’ 3,900–9,200 band; his 50k–500k was ~1–2 OOM high, and your “~1 OOM below his number” instinct bracketed the truth.The premise worth challenging
Your question was “how are the forecasts doing.” The sharper version is: at this horizon, the question barely discriminates. The thing that makes the Metaculus community distinctive — willingness to entertain exponentials, regime changes, and fat tails — is exactly what 1–3-year medians suppress, because over short windows the aggregate reverts to consensus. So a strong near-term scorecard (which this is) is mostly evidence that the consensus was right, and tells you almost nothing about whether the community’s tail-aware reasoning adds value. The one place that distinctive reasoning should have paid off — taking the AI exponential seriously — is precisely where they most under-used it.
Two implications:
The actual test of this tournament hasn’t run yet. Its entire design rationale was the 30/52/122-year horizons — where the OWID note explicitly says the community’s willingness to price black swans makes it less conservative than traditional sources. That bet is unfalsified until those resolve. Grading 2023/2025 is grading the warm-up.
For your AGI-timelines work specifically: the cleanest signal in this whole exercise is a base rate — sophisticated forecasters under-predict AI scaling even when watching it happen. Compute and capital both ran several-fold ahead of well-reasoned medians over three years. That’s a thumb on the scale toward the aggressive end of current AI forecasts, not the modal one.
Caveats on this grading
I’m scoring against the report’s prose descriptions, not the chart medians/IQRs — fine for the unambiguous calls, weaker for certifying “in range” on the hits. Three metrics (chickens, homicide, poverty) can’t be graded yet on their ~2-year-lagged sources. Two resolution problems are real and worth internalizing as a forecasting lesson in their own right: the genome-cost series simply stopped existing (NHGRI froze in 2022), and the “objects launched” and “nuclear” counts get revised upward for years after the fact, so any single-year resolution is a moving target. A forecast can be undone by the measurement pipeline rather than the world.
As an aside, I also asked the models to assess each other’s assessments above:
GPT-5.5 assesses Opus 4.8: “conceptual diagnosis is strong, but its scorecard is materially sloppier than mine on several important rows. The big difference is that it often grades the report’s prose instead of the report’s forecast table… it incorrectly treats fastest supercomputer and liberal democracy population as hits, and overstates the AI investment miss by mixing definitions”
I think the other model’s conceptual diagnosis is strong, but its scorecard is materially sloppier than mine on several important rows. The big difference is that it often grades the report’s prose instead of the report’s forecast table. That makes it more rhetorically satisfying, but less reliable as an audit.
Its best line is this: near-term OWID forecasts mostly test whether forecasters anchored well to institutional baselines, while the more distinctive Metaculus skill, pricing exponentials and weird tails, is barely visible at 1–3 years. I agree with that. It also correctly spots the AI compute undercall, the AI boom, the terrorism miss, the no-recession surprise, and the broader “inertial variables did well” pattern. Its own caveat says it is scoring against prose descriptions rather than exact medians/IQRs, and that caveat matters a lot.
Where I disagree: its tally is too generous. It says roughly 12 clean hits, 1 split, 4 misses, 1 measurement failure, 3 lagged-not-gradable. I would revise that downward, mainly because it incorrectly treats fastest supercomputer and liberal democracy population as hits, and overstates the AI investment miss by mixing definitions.
Quick comparison
Issue
Other model
My view
Overall framing
Strong: “near-term mostly grades consensus.”
I agree. This is probably the most important meta-point.
Use of exact forecast values
Weak: often uses prose/ranges.
I’d privilege the report tables and resolution definitions.
AI training compute
Correctly says miss low.
Agree, and probably the most important technical miss.
AI investment
Says biggest miss under.
Too strong. Depends on whether resolution uses narrow external private investment or broad corporate AI spending.
Supercomputer FLOPS
Calls it a hit.
I think this is wrong. Forecast was ~4.1–4.5e18 for 2025; actual TOP500 was 1.809e18.
Liberal democracies
Calls it directionally a hit.
I think this is wrong. Forecast was ~1.0–1.1B for 2025; V-Dem says 0.6B at end-2025.
Terrorism
Correctly says miss.
Agree, but it underplays the magnitude. Forecast was 19–20k; actual 2025 was 5,582.
CO₂
Says number hit, peak narrative premature.
Agree. That is a nice distinction.
Space objects
Calls hit.
Agree. Pros were impressively close.
Demographics
Calls population/fertility hits.
Mostly agree, though its fertility actuals are a little messy depending on UN/OWID projection vintage.
The most important factual disagreements
1. Supercomputer FLOPS: the other model is wrong to call this a hit
The report table gives the 2025 fastest-supercomputer forecasts as 4.1 × 10¹⁸ FLOP/s for Pros and 4.5 × 10¹⁸ for the public. The TOP500 November 2025 result was El Capitan at 1.809 exaFLOP/s. That is not a hit. It is a miss high by about 2.3–2.5×. The other model appears to have inferred from the prose that forecasters expected “sub-2 EF” by 2025, but the table does not say that. It says ~4 EF.
This is not a trivial bookkeeping dispute. It affects the 2032 adjustment. If you call 2025 a hit, you might leave the 2032 TOP500 forecast alone. If you call it a miss high, you should lower 2032, which is what I did.
2. AI investment: the other model overclaims the miss
The report table gives 2025 AI investment forecasts of $340B Pro and $489B public. The other model says actuals were “$581.7B corporate / $344.7B private” and therefore calls this the biggest miss under. But OWID’s relevant AI-investment chart says it focuses on external private-market investment and excludes internal corporate R&D, capex, public-sector funding, and public companies. On that narrower definition, $344.7B private makes the Pro forecast of $340B look almost exactly right, while the public forecast of $489B looks high. On broad corporate AI investment, the public forecast may be low.
So the better judgment is: direction right, metric definition decisive. The other model’s “biggest miss under” conclusion is only defensible if you resolve against broad corporate AI investment, not against the narrower OWID-style private-investment series.
3. Liberal democracies: the other model misses the U.S./V-Dem discontinuity
The report forecasted 1.00B Pro and 1.10B public people living in liberal democracies in 2025. V-Dem’s 2026 Democracy Report says only 7% of the world population, about 0.6B people, lived in liberal democracies at the end of 2025, and explicitly says the U.S. lost its long-term liberal-democracy status for the first time in over 50 years.
The other model calls this a directional hit, saying V-Dem shows a slow decline around ~1.0B. That looks stale or wrong. This is one of the biggest substantive misses in the whole set, because the forecast’s own resolution criteria used OWID or an alternative credible source for people living in liberal democracies, and the report itself flagged U.S. 2024 turmoil as a possible reason the U.S. might lose liberal-democracy status.
4. Terrorism: it gets the sign right but underplays the magnitude
The report forecasted 19k–20k terrorism fatalities for both 2023 and 2025. The Global Terrorism Index 2026 says 2025 deaths were 5,582, down 28% from 2024. It also says 2023 deaths spiked to 8,654, largely due to October 7, before falling again. So yes, the forecasters missed, but the other model frames the miss around “2023 rose 22%” instead of the more important fact that the median forecast was about 2.2× too high in 2023 and about 3.4× too high in 2025.
The right lesson is subtler than “black swan happened.” The October 7 shock pushed deaths upward, but even with that shock, the forecast was still far too high. The forecasters seem to have over-anchored to older high-terrorism baselines or to right-tail concerns.
5. CO₂: I like its distinction
Here I think the other model improves on my initial phrasing. The report’s Pro forecasts were 37.7 Gt for 2023 and 38.1 Gt for 2025, while the public forecasts were 36.4 and 36.2. The IEA says 2025 energy-related CO₂ reached a new high of nearly 38.4 Gt. So the Pro number was very good, the public number was too low, and the qualitative “at or near peak” narrative looks premature.
That “number right, story too early” framing is exactly the right way to score it.
Where the other model is better than mine
It is better than my first answer in three ways.
First, it has a cleaner theory of the errors: inertial variables good, regime/tail variables bad. That’s a useful compression. Population, space objects, fertility, energy, nuclear stockpiles, CO₂ Pro forecasts, and same-sex marriage mostly follow observable short-run trends. AI compute, AI investment, terrorism, GDP surprises, and liberal-democracy classification are more thresholdy or regime-sensitive.
Second, it correctly emphasizes that a 1–3-year scorecard is a weak test of the tournament’s real ambition. The tournament’s rationale was about 10-, 30-, and 100-year uncertainty. In the short run, even excellent forecasters often just become careful aggregators of UN, IMF, FAS, OECD, IEA, TOP500, OWID, and similar baselines. That is valuable, but it is not the same as proving far-future forecasting skill.
Third, it is right to treat AI as the most informative miss. The report’s 2025 AI training-compute forecasts were 2.9e25 Pro and 6.6e25 public, while Epoch now puts the largest known training run at around 5e26 FLOP, and notes that more than 30 models had crossed the 1e25 FLOP scale by June 2025.
Where mine is better
Mine is more careful about resolution definitions. That matters because these are not abstract “what happened in the world?” questions. They are “what number will a specified dataset or credible replacement source report?” questions. The other model sometimes slides from the metric to a nearby intuitive proxy: corporate AI spend instead of external private investment; prose expectations instead of table medians; total nuclear inventory instead of military stockpile-style series; broad democratic decline instead of the specific V-Dem/OWID liberal-democracy population count.
Mine is also less willing to count “directionally on track” as a hit when the forecast was numerical and the relevant year is missing or the actual value is far from the median. The other model’s “hit” threshold is too forgiving. A forecast of 1.0B liberal-democracy population resolving at 0.6B is not a hit just because the direction was “slightly down.” A forecast of 4.1e18 FLOP/s resolving at 1.809e18 is not a hit just because the prose said the leap might be delayed.
Revised synthesis after comparing both
I’d now summarize the 2023/2025 track record like this:
Cluster
My revised judgment
Very good
Space objects, world population, CO₂ Pro forecast, fertility, per-capita energy, nuclear stockpiles roughly, same-sex marriage roughly
Mixed or definition-sensitive
AI investment, genome sequencing, GDP per capita growth, OECD trust, G7 life expectancy, U.S. productivity
Clear misses
AI training compute low, terrorism high, liberal-democracy population high, fastest-supercomputer FLOPS high, public CO₂ too low
Not properly scoreable yet
Chickens, homicide, poverty, and several non-Pro/public-only OWID series with lagged data
So I’d move my earlier “B/B-minus overall” slightly downward for the fully scoreable subset, maybe closer to B-minus or C-plus if you grade medians strictly, while preserving B-plus-ish for slow-moving macro/demographic variables. The other model’s “12 clean hits” is too kind because two of those clean hits are not clean hits.
The best combined takeaway is this: Metaculus was good where disciplined base-rate anchoring was enough, but not clearly better than careful institutional extrapolation at 1–3 years. It underpriced AI scale-up, overestimated terrorism, missed a threshold political-classification shock, and overestimated TOP500 growth. For 2032, the right way to beat the original forecasts is to mostly leave the boring variables alone and make concentrated adjustments on the variables where these short-run misses reveal a wrong model class.
Opus 4.8 assesses GPT-5.5: “strong assessment — better than mine in several specific places because it has the actual Pro/Public medians I lacked — and corrects a real error I made. But also one significant flaw that fooled it, for the terrorism verdict (and, in a different way, fooled me too), and one overstatement, for supercomputer FLOPS”
This is a strong assessment, and in a few places it’s better than mine — including one spot where it cleanly corrects a real error I made. Before I adjudicate the places where we disagree (terrorism, supercomputers, liberal democracies), let me verify the two load-bearing facts behind its sharpest claims, since the whole comparison turns on who’s actually right.
V-Dem confirms the model: just 7% of the world (about 600 million people) lived in liberal democracies at the end of 2025 — the lowest in over 50 years — with the drop driven largely by the USA losing liberal-democracy status. So the model caught a real miss I was too soft on. Now the terrorism series, where I suspect the model’s “2.3× miss” may be a measurement artifact.
Both fact-checks came back clean, and they’re decisive. Verdict up front: this is a strong assessment — better than mine in several specific places because it has the actual Pro/Public medians I lacked — and it corrects a real error I made. But it also contains one significant flaw that fooled it (and, in a different way, fooled me too), and one overstatement. Combining its numbers with my provenance-checking beats either of us alone.
A reassuring meta-point first: the medians it reports are almost certainly authentic, not hallucinated. The tell is internal consistency with the underlying data series — e.g. its terrorism median (~20k) matches the GTD baseline the question actually resolves on, and its AI-investment Pro median ($340B) matches the narrow-private actual almost exactly. A model fabricating numbers wouldn’t reliably reproduce series-specific baselines it had no other reason to know. That said, the entire numerical scorecard rests on those transcriptions, so they’re the one thing worth spot-checking against the report PDF.
Where it’s right and corrects me
AI investment — it catches a genuine error of mine, and this is the biggest single correction. In turn 1 I called this “the biggest under-forecast” by comparing the forecast to ~$582B of broad corporate AI investment. That was the wrong denominator. The OWID metric is narrow external private funding for privately-held AI companies, which excludes Big Tech capex, public firms, and public-sector money — and on that series the 2025 actual is ~$345B. So the Pro median of $340B was nearly exact, and the Public’s $489B was ~40% high. The model’s framing — “directionally right, hard to score, looks high on narrow private funding” — is correct and mine was wrong. And it’s a more interesting result than my version: forecasting $340B in Feb 2023, off a falling $92B base, means the Pros actually did price the investment boom on this metric. I had the story backwards.
Liberal democracies — it’s right, I was too lenient. I waved this through as a “directional hit (noisy metric).” It’s a clear miss high: forecast ~1.0B, actual ~0.6B (7% of the world) at end-2025, the lowest in 50+ years, largely because the USA lost its liberal-democracy status — the first time in more than 50 years, falling to electoral democracy. The model names the mechanism precisely (“overexpected institutional continuity in democracy”). I undersold a real miss.
Nuclear — cleaner than mine. Its medians (~9.4–9.5k) make clear the question resolves on the military stockpile (~9,600), which is exactly right and resolves the total-vs-stockpile ambiguity I left dangling in turn 1.
Where I’d push back on it (and where I was also wrong)
Terrorism — its sharpest-looking verdict is its biggest flaw, and my turn-1 take was wrong too. The model calls this a “bad miss high… overestimated by ~2.3× in 2023 and ~3.4× in 2025.” But it compared the forecast to the wrong database. The OWID metric resolves on the Global Terrorism Database (GTD), on which ~20,000 people died from terrorism in 2019 — so the Pro median of ~19–20k isn’t an overestimate at all; it’s a sensible extrapolation of the actual resolution series. Two problems then compound:
The GTD is discontinued. As of 2025 the GTD is closed, and the data only extend through 2020/2021. There is no native 2023 or 2025 value — this question is unresolvable, exactly like the genome-cost case (NHGRI frozen at 2022). It belongs in the not-scoreable bucket, not the miss column.
The “actual” it used is a different, lower-counting database. The 8,352 / 5,582 figures are from the GTI, which systematically counts fewer deaths than the GTD — GTI put 2019 at 13,826 versus GTD’s ~20,000 for the same year. So a chunk of the apparent “2.3× miss” is just the GTD→GTI level shift, not forecaster error.
And the kicker: on the GTI proxy, deaths actually fell to 5,582 in 2025, the lowest since 2007 — so the forecasters’ qualitative “continue to decline” call was directionally vindicated by 2025. My own turn-1 framing (“predicted decline, got an Oct-7 spike to 8,352, miss”) was also wrong: I used the GTI number, missed that the GTD is the dead resolution series, and treated a one-year blip as the verdict. We both mis-scored this in opposite directions; the correct call is “unresolvable / not cleanly scoreable, decline directionally right.”
Supercomputer FLOPS — “clear miss high” overstates it. Forecast 4.1e18 vs actual 1.809e18 is 2.3× — but that’s 0.36 of an order of magnitude on a log-scale exponential metric. On the same log-error basis the model (correctly) uses to call AI-compute a miss (~0.9 OOM low), a 0.36-OOM overshoot is a mild overshoot, not a “clear miss.” The directional read (“expected the exascale race to accelerate faster than the TOP500 benchmark did”) is right; the severity is dialed too high. There’s also a metric-decoupling issue: the frontier AI clusters (xAI, Meta) don’t run the LINPACK benchmark, so TOP500′s #1 increasingly fails to measure the thing that matters.
Where we agree
The structural/slow-moving hits (population, fertility, CO₂ level, energy, space objects, same-sex marriage); AI training compute as a genuine miss-low (its “~5e26 Grok 4 vs 6.6e25 Public median” and my “~1e26+ vs an extrapolated median” both land at ~1 OOM under); and the laggy third that simply isn’t scoreable yet (chickens, poverty, homicide, genome cost, productivity, GDP growth). On that last bucket the model is actually better organized than I was — its dedicated table with specific reasons (PPP-vintage shifts, NHGRI’s 2022 freeze, Penn World Table ending at 2023) is cleaner than my prose.
The synthesis that beats both assessments
Here’s the reconciled picture of the contested calls:
Metric
Model’s verdict
My turn-1 verdict
Reconciled (what’s actually true)
AI investment
Directionally right, looks high (narrow series)
“Biggest under-forecast”
Model right. Pro $340B ≈ exact vs ~$345B; I used the wrong denominator
Liberal democracies
Clear miss high (US reclassified)
“Directional hit, noisy”
Model right. ~1.0B forecast vs 0.6B actual; I was too lenient
Terrorism
Bad miss high (~2.3×)
Miss (Oct-7 spike)
Both wrong. GTD discontinued + GTI mismatch → unresolvable; decline directionally right
Supercomputer FLOPS
Clear miss high
Hit
Between us. Mild 0.36-OOM overshoot + metric decoupling
Nuclear
Very good
Hit (waffled on series)
Model cleaner (military stockpile, ~9.6k)
AI compute
Miss low (~1 OOM)
Miss low
Agree
Lay these side by side and a pattern jumps out that neither assessment states: of the model’s four headline “misses,” only AI compute is a clean real-world forecasting error. Terrorism is a dead data source plus a database mismatch; liberal democracies is a V-Dem reclassification of the US (a classification event, unforeseeable in Feb 2023, not a smooth trend the forecasters misread); supercomputers is a 0.36-OOM overshoot on a metric that’s decoupling from reality. Three of the four are measurement/definition artifacts, not failures to read the world.
That has two consequences. First, it means the forecasters look better than the model’s B/B-minus — closer to B-plus/A-minus on real-world calibration, with the honest caveat that roughly a third of the metrics are measurement-fragile. The model’s grade is slightly too harsh because it scored three measurement artifacts as forecasting misses. Second, it sharpens my turn-1 thesis: I argued the 1–3yr exercise mostly grades the consensus they anchored on; the better statement now is that the dominant failure mode at this horizon isn’t bad forecasting, it’s metric fragility — series get discontinued (GTD, NHGRI), redefined (V-Dem reclassification), or measure the wrong proxy (TOP500 vs AI clusters). The one place the world itself outran them is AI compute. Everything else that “missed” is the ruler changing, not the thing being measured.
Net: the model’s assessment is genuinely good and I’d trust its numbers over my turn-1 qualitative grading wherever they conflict — it fixed my AI-investment and liberal-democracy errors outright. What I’d add to it is provenance discipline (which rescues terrorism from the miss column and softens supercomputers) and the measurement-artifact reframe. The combined verdict is more accurate than either of us produced alone, which is the whole point of running two passes.
I’ll let Opus 4.8 have the final word: