Pith. sign in

REVIEW 4 cited by

Large Language Model Prediction Capabilities: Evidence from a Real-World Forecasting Tournament

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.13014 v1 pith:X5QGPPJJ submitted 2023-10-17 cs.CY cs.AIcs.CLcs.LG

classification cs.CYcs.AIcs.CLcs.LG
keywords forecastingforecastsgpt-4real-worldcapabilitieslanguagelargeprediction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Accurately predicting the future would be an important milestone in the capabilities of artificial intelligence. However, research on the ability of large language models to provide probabilistic predictions about future events remains nascent. To empirically test this ability, we enrolled OpenAI's state-of-the-art large language model, GPT-4, in a three-month forecasting tournament hosted on the Metaculus platform. The tournament, running from July to October 2023, attracted 843 participants and covered diverse topics including Big Tech, U.S. politics, viral outbreaks, and the Ukraine conflict. Focusing on binary forecasts, we show that GPT-4's probabilistic forecasts are significantly less accurate than the median human-crowd forecasts. We find that GPT-4's forecasts did not significantly differ from the no-information forecasting strategy of assigning a 50% probability to every question. We explore a potential explanation, that GPT-4 might be predisposed to predict probabilities close to the midpoint of the scale, but our data do not support this hypothesis. Overall, we find that GPT-4 significantly underperforms in real-world predictive tasks compared to median human-crowd forecasts. A potential explanation for this underperformance is that in real-world forecasting tournaments, the true answers are genuinely unknown at the time of prediction; unlike in other benchmark tasks like professional exams or time series forecasting, where strong performance may at least partly be due to the answers being memorized from the training data. This makes real-world forecasting tournaments an ideal environment for testing the generalized reasoning and prediction capabilities of artificial intelligence going forward.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Predicting Empirical AI Research Outcomes with Language Models

    cs.AI 2025-06 conditional novelty 7.0 of 10

    A fine-tuned, retrieval-augmented language model predicts which of two AI research ideas will perform better empirically, beating human experts and frontier models on a new contamination-controlled benchmark.

  2. WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A live, leakage-free benchmark of six frontier LLMs across all 104 World Cup matches finds they match, not beat, the betting market and herd on outcomes.

  3. Argumentatively Coherent Judgmental Forecasting

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Filtering forecasts that are incoherent with their argumentative reasoning improved accuracy in human and LLM experiments, though users do not naturally follow the proposed coherence notion.

  4. Large Language Models for Predictive Analysis: How Far Are They?

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Existing LLMs perform poorly on predictive analysis, with the best model scoring 24.11/28 and most models failing to generate executable code.

Pith tools