AI ClaimsRapid review

Philip E. Tetlock and Dan Gardner · 2015

Superforecasting

The Art and Science of Prediction

Cover via Open Library

Rough AI truth score

75/100

The Good Judgment Project establishes real, repeated individual and team differences in short-horizon geopolitical forecasting. Training, decomposition, base rates, updating, aggregation, and feedback contribute to performance, and an independent study reproduces the core result. External validity is bounded: tournament questions exclude many slow, reflexive, structural, and fat-tailed uncertainties that matter most.

Based on three central claims · high confidence

The three claims

01Supported

A small subset of forecasters repeatedly outperforms average participants on resolvable geopolitical questions.

Multi-year tournaments scored thousands of probabilistic forecasts and identified participants with persistently lower Brier scores and useful year-to-year performance stability. Selection and regression to the mean matter, but an independent expedited-identification study found results consistent with the original program.

02Mostly supported

Decomposition, base rates, frequent updating, calibration feedback, and team discussion improve probabilistic forecasts.

Tournament interventions, aggregation studies, and structured-expert work support training, feedback, decomposition, performance weighting, and disciplined updating. Components are bundled, participant selection matters, and teams can share correlated information, so exact causal contributions are not uniform.

03Mixed

These methods generalize from short-horizon tournament questions to most consequential long-range policy and tail-risk problems.

Clear questions, reference classes, decomposition, probability discipline, and updating are broadly useful habits. Long-horizon policy outcomes are often endogenous, reflexive, nonstationary, ambiguously resolved, data-poor, and fat-tailed; tournament evidence does not establish equal gains there.

Other claims worth checking
  • Forecast tournaments can make disagreement explicit and scoreable.
  • Intellectual humility and numeracy predict part of forecasting performance.

What this number means. It is an AI-generated first-pass judgment of three central factual or causal claims—not a rating, exhaustive fact-check, or human peer review. Claim credits are 100% for supported, 75% for mostly supported, 50% for mixed, and 25% for weak, then averaged and rounded. Lower confidence means the score should move more as better evidence arrives.

Method three-central-claims/0.1.0 · checked 2026-09-01 · 3/3 selected claims assessed · method and source audit