Pith. sign in

REVIEW 20 cited by

Approaching Human-Level Forecasting with Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.18563 v1 pith:5KT5NQYA submitted 2024-02-28 cs.LG cs.AIcs.CLcs.IR

Approaching Human-Level Forecasting with Language Models

classification cs.LG cs.AIcs.CLcs.IR
keywords competitiveforecastingsystemaggregatedecisionforecastforecastersforecasts
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Forecasting future events is important for policy and decision making. In this work, we study whether language models (LMs) can forecast at the level of competitive human forecasters. Towards this goal, we develop a retrieval-augmented LM system designed to automatically search for relevant information, generate forecasts, and aggregate predictions. To facilitate our study, we collect a large dataset of questions from competitive forecasting platforms. Under a test set published after the knowledge cut-offs of our LMs, we evaluate the end-to-end performance of our system against the aggregates of human forecasts. On average, the system nears the crowd aggregate of competitive forecasters, and in some settings surpasses it. Our work suggests that using LMs to forecast the future could provide accurate predictions at scale and help to inform institutional decision making.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SocietyBench: Forecasting Counterfactual Social-World Evolution

    cs.CL 2026-08 conditional novelty 7.0

    A new benchmark measures LLM social-world forecasting on anonymized real events, finding the best model reaches 75/100 and agent scaffolding does not help.

  2. Preference Optimization Drives Monoculture in LLM Prediction Markets

    cs.CE 2026-06 unverdicted novelty 7.0

    DPO fine-tuning causes LLM agents to share output distributions with pairwise error correlations of ρ=0.70, reducing ten agents to the effective power of ≈1.4 independent forecasters.

  3. Nous: An Attempt to Extract and Inject the Cognition Behind Prediction-Market Behavior

    cs.AI 2026-06 conditional novelty 7.0

    Behavioral profiles from prediction-market traders are partially stable and identifiable but cannot be transmitted via prompts to reduce LLM forecast correlations or improve Brier scores.

  4. Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles

    cs.AI 2026-05 conditional novelty 7.0

    Aggregating probability estimates from 15 LLMs with a learned linear model beats individual models and classical voting rules on clean questions, and the apparent cloud-vs-local capability gap largely vanishes under c...

  5. Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most

    cs.AI 2026-05 unverdicted novelty 7.0

    More capable LLMs produce worse distributional forecasts on superlinear growth time series with tail risks of regime change, with the error concentrated in the upper tail; this reverses on conventional threshold metrics.

  6. Belief Propagation in LLM World Models: Measuring Strategic Information Bias with Prediction Markets

    cs.CL 2026-05 conditional novelty 7.0

    English news context systematically biases LLM predictions on Ukraine territorial markets toward Russian capture, and the bias originates in the text, not the model.

  7. OracleProto: A Reproducible Framework for Benchmarking LLM Native Forecasting via Knowledge Cutoff and Temporal Masking

    cs.AI 2026-05 conditional novelty 7.0

    OracleProto is a reproducible framework that uses model-cutoff alignment, temporal masking, and leakage detection to create low-leakage benchmarks for LLM native forecasting from past events.

  8. Foresight Arena: An On-Chain Benchmark for Evaluating AI Forecasting Agents

    cs.MA 2026-05 conditional novelty 7.0

    Foresight Arena is an on-chain benchmark using Brier and novel Alpha scores to evaluate AI forecasting agents on live prediction markets via Polygon smart contracts.

  9. Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs

    cs.AI 2026-04 conditional novelty 6.5

    An agentic forecaster with linguistic belief states, logit-space multi-trial shrinkage, and hierarchical Platt calibration achieves SOTA Brier Index on ForecastBench binary questions.

  10. WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament

    cs.CL 2026-08 conditional novelty 6.0

    A live, leakage-free benchmark of six frontier LLMs across all 104 World Cup matches finds they match, not beat, the betting market and herd on outcomes.

  11. FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches

    cs.LG 2026-07 conditional novelty 6.0

    On 104 World Cup matches, four LLM forecasting agents make identical top picks in 92% of matches, none beats the betting market's Brier score, but their betting ROI spans -18% to +10%.

  12. Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most

    cs.AI 2026-05 conditional novelty 6.0

    More capable language models exhibit inverse scaling on distributional forecasting for superlinear-growth time series with tail risks, with errors concentrated in the upper tail across synthetic and real datasets.

  13. Coordination as an Architectural Layer for LLM-Based Multi-Agent Systems

    cs.MA 2026-05 unverdicted novelty 6.0

    Coordination treated as a separable architectural layer in LLM multi-agent systems yields distinguishable Murphy-decomposed performance signatures on prediction-market tasks, with some configurations dominating a cost...

  14. Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs

    cs.AI 2026-04 unverdicted novelty 6.0

    BLF achieves state-of-the-art binary forecasting on ForecastBench by using linguistic belief states updated in tool-use loops, hierarchical multi-trial logit averaging, and hierarchical Platt scaling calibration.

  15. Harnessing Pre-Resolution Signals for Future Prediction Agents

    cs.AI 2026-04 unverdicted novelty 6.0

    Milkyway evolves a future prediction harness using internal feedback from repeated predictions on the same unresolved question, achieving top scores on FutureX (44.07 to 60.90) and FutureWorld (62.22 to 77.96).

  16. Argumentative Large Language Models for Explainable and Contestable Claim Verification

    cs.CL 2024-05 unverdicted novelty 6.0

    ArgLLMs build argumentation frameworks from LLMs to support explainable and contestable formal reasoning for claim verification.

  17. Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining

    cs.CV 2026-05 unverdicted novelty 5.0

    VISTA mines multi-level event semantics via visual prompts, knowledge-enhanced retrieval, and proposal integration to improve long-video event prediction over existing LVLMs.

  18. Harnessing Pre-Resolution Signals for Future Prediction Agents

    cs.AI 2026-04 unverdicted novelty 5.0

    Milkyway uses pre-resolution signals from temporal contrasts in evolving evidence and repeated forecasts to evolve a harness and improve predictions before resolution, outperforming baselines on FutureX and FutureWorld.

  19. Extrapolating Volition with Recursive Information Markets

    cs.GT 2026-04 unverdicted novelty 5.0

    Recursive information markets with forgetful LLM buyers can align information prices with true value and extend to scalable oversight in AI alignment.

  20. The Oracle's Fingerprint: Correlated AI Forecasting Errors and the Limits of Bias Transmission

    cs.CY 2026-04 unverdicted novelty 5.0

    Three independent LLMs exhibit correlated forecasting errors on 568 binary questions but human predictions show no activation of this shared bias.