Pith. sign in

REVIEW 15 cited by

ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.19839 v5 pith:LB2FRRFU submitted 2024-09-30 cs.LG cs.AIcs.CL

ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities

classification cs.LG cs.AIcs.CL
keywords forecastbenchquestionssystemsbenchmarkforecastingforecastsaccuracycapabilities
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
abstract

Forecasts of future events are essential inputs into informed decision-making. Machine learning (ML) systems have the potential to deliver forecasts at scale, but there is no framework for evaluating the accuracy of ML systems on a standardized set of forecasting questions. To address this gap, we introduce ForecastBench: a dynamic benchmark that evaluates the accuracy of ML systems on an automatically generated and regularly updated set of 1,000 forecasting questions. To avoid any possibility of data leakage, ForecastBench is comprised solely of questions about future events that have no known answer at the time of submission. We quantify the capabilities of current ML systems by collecting forecasts from expert (human) forecasters, the general public, and LLMs on a random subset of questions from the benchmark ($N=200$). While LLMs have achieved super-human performance on many benchmarks, they perform less well here: expert forecasters outperform the top-performing LLM ($p$-value $<0.001$). We display system and human scores in a public leaderboard at www.forecastbench.org.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PolyBench: Benchmarking LLM Forecasting and Trading Capabilities on Live Prediction Market Data

    q-fin.CP 2026-04 conditional novelty 8.0

    Only two of seven LLMs produce positive returns on live Polymarket data, with MiMo-V2-Flash at 17.6% CWR and Gemini-3-Flash at 6.2% CWR while the other five lose money.

  2. Verifiable Rewards for Calibrated Probabilistic Forecasting

    cs.LG 2026-06 unverdicted novelty 7.0

    A verifiable empirical win rate reward combined with gradient masking enables RL training of a 7B model to reach betting-market calibration on NFL win probabilities using only outcome data.

  3. ForecastBench-Sim: A Simulated-World Forecasting Benchmark

    cs.AI 2026-06 unverdicted novelty 7.0

    ForecastBench-Sim is a simulated-world benchmark using Freeciv game rollouts to generate resolvable forecasting questions at arbitrary horizons with paired intervention worlds.

  4. StakeBench: Evaluating Language Understanding Grounded in Market Commitment

    cs.CL 2026-05 unverdicted novelty 7.0

    StakeBench is a new benchmark using market-derived supervision from resolved prediction markets to test LLMs on commitment detection, side identification, action anticipation, and odds projection, revealing partial su...

  5. OracleProto: A Reproducible Framework for Benchmarking LLM Native Forecasting via Knowledge Cutoff and Temporal Masking

    cs.AI 2026-05 conditional novelty 7.0

    OracleProto is a reproducible framework that uses model-cutoff alignment, temporal masking, and leakage detection to create low-leakage benchmarks for LLM native forecasting from past events.

  6. Foresight Arena: An On-Chain Benchmark for Evaluating AI Forecasting Agents

    cs.MA 2026-05 conditional novelty 7.0

    Foresight Arena is an on-chain benchmark using Brier and novel Alpha scores to evaluate AI forecasting agents on live prediction markets via Polygon smart contracts.

  7. Energy-Arena: A Dynamic Benchmark for Operational Energy Forecasting

    econ.EM 2026-04 unverdicted novelty 7.0

    Energy-Arena is a dynamic, forward-looking benchmarking platform that standardizes ex-ante submissions and rolling ex-post evaluations for operational energy forecasting to improve transparency and comparability.

  8. CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction

    cs.AI 2026-04 accept novelty 7.0

    CT Open is a new live platform with an automated LLM-powered decontamination pipeline that supplies uncontaminated benchmarks for predicting clinical trial outcomes.

  9. Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

    cs.AI 2026-08 conditional novelty 6.0

    In two adversarial real-world domains with time-delayed expert ground truth, frontier LLMs over-generate plausible ideas but under-select the specific solutions experts adopted, indicating a filtering and prioritization gap.

  10. WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

    cs.AI 2026-07 conditional novelty 6.0

    A pre-registered benchmark of 13 LLM systems on all 104 World Cup 2026 matches shows fine-grained predictions expose differences that result accuracy hides.

  11. FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches

    cs.LG 2026-07 conditional novelty 6.0

    On 104 World Cup matches, four LLM forecasting agents make identical top picks in 92% of matches, none beats the betting market's Brier score, but their betting ROI spans -18% to +10%.

  12. Beyond Forecasting: The Belief-to-Trade Layer in Prediction-Market Agents

    cs.AI 2026-07 conditional novelty 6.0

    An explicit selection–sizing–risk trading layer on fixed forecasts yields the only positive ROI and stake-weighted Sharpe among five policies on a Polymarket decision archive.

  13. Diverse Evidence, Better Forecasts: Multi-Agent Deliberation Under Information Asymmetry

    cs.AI 2026-07 unverdicted novelty 6.0

    InfoDelphi partitions evidence to induce information asymmetry in multi-agent LLM deliberation, yielding 12-18% Brier score gains and 4-8 pp accuracy gains on a 375-question benchmark.

  14. From Forecasting Leaderboards to Deployment Decisions: A Fail-Closed Certification Protocol

    cs.LG 2026-06 unverdicted novelty 6.0

    Presents a fail-closed certification protocol for determining when forecasting leaderboard winners are deployment-actionable, using a traffic dataset to show friction-induced reversals and an audit to prevent overclaiming.

  15. Coordination as an Architectural Layer for LLM-Based Multi-Agent Systems

    cs.MA 2026-05 unverdicted novelty 6.0

    Coordination treated as a separable architectural layer in LLM multi-agent systems yields distinguishable Murphy-decomposed performance signatures on prediction-market tasks, with some configurations dominating a cost...