Pith. sign in

REVIEW 10 cited by

Bench to the Future: A Pastcasting Benchmark for Forecasting Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.21558 v1 pith:ABW2GBME submitted 2025-06-11 cs.CL cs.AIcs.LG

Bench to the Future: A Pastcasting Benchmark for Forecasting Agents

classification cs.CL cs.AIcs.LG
keywords forecastingbenchmarkpastcastingquestionsresultsbenchchallengingenvironment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Forecasting is a challenging task that offers a clearly measurable way to study AI systems. Forecasting requires a large amount of research on the internet, and evaluations require time for events to happen, making the development of forecasting benchmarks challenging. To date, no forecasting benchmark provides a realistic, hermetic, and repeatable environment for LLM forecasters. We introduce Bench To the Future (BTF), a "pastcasting" benchmark with hundreds of high-quality questions for which the resolution is already known. Each question is accompanied by a large offline corpus of tens of thousands of relevant web pages, enabling a way to elicit realistic "forecasts" on past events from LLMs. Results suggest that our pastcasting environment can produce results comparable to those based on forecasts using the internet on at-the-time unresolved questions. We show results benchmarking agent and chain-of-thought forecasting approaches using several LLMs, including the recently-released Claude 4 models, and demonstrate BTF's ability to track steady forecasting capability progress over time. We intend this to be a living benchmark, with new questions added continually to account for increasing training data cutoff dates. We invite researchers to contact us at hello@futuresearch.ai to utilize our benchmark or tooling for their own research.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Agentic Time Machine as an Infrastructure for Future-Event Forecasting

    cs.AI 2026-06 unverdicted novelty 7.0

    Agentic Time Machine reconstructs historical web states for offline evaluation of forecasting agents, with a multi-agent framework achieving top ranks on FutureX live and past benchmarks.

  2. ForecastBench-Sim: A Simulated-World Forecasting Benchmark

    cs.AI 2026-06 unverdicted novelty 7.0

    ForecastBench-Sim is a simulated-world benchmark using Freeciv game rollouts to generate resolvable forecasting questions at arbitrary horizons with paired intervention worlds.

  3. WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning

    cs.CL 2026-06 unverdicted novelty 7.0

    WorldReasoner supplies 345 resolved forecasting tasks built from 14,141 articles to score LM agents on outcome quality, evidence quality, and reasoning quality against time-bounded evidence and hindsight graphs.

  4. Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most

    cs.AI 2026-05 unverdicted novelty 7.0

    More capable LLMs produce worse distributional forecasts on superlinear growth time series with tail risks of regime change, with the error concentrated in the upper tail; this reverses on conventional threshold metrics.

  5. Evaluating Strategic Reasoning in Forecasting Agents

    cs.AI 2026-04 unverdicted novelty 7.0

    BTF-2 benchmark shows frontier AI forecasters lag humans mainly in assessing political and business leaders' incentives, plan follow-through likelihood, and institutional processes, with a composite agent gaining 0.01...

  6. Global Merger-Arbitrage Forecasting with Language Models

    cs.CL 2026-07 conditional novelty 6.5

    Expert-context research agents plus hindsight-guided finetuning cut class-balanced Brier score on merger outcomes to 0.151, beating calibrated market prices, XGBoost, and frontier LLMs.

  7. Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most

    cs.AI 2026-05 conditional novelty 6.0

    More capable language models exhibit inverse scaling on distributional forecasting for superlinear-growth time series with tail risks, with errors concentrated in the upper tail across synthetic and real datasets.

  8. Harnessing Pre-Resolution Signals for Future Prediction Agents

    cs.AI 2026-04 unverdicted novelty 6.0

    Milkyway evolves a future prediction harness using internal feedback from repeated predictions on the same unresolved question, achieving top scores on FutureX (44.07 to 60.90) and FutureWorld (62.22 to 77.96).

  9. Temporal Leakage in Search-Engine Date-Filtered Web Retrieval: A Retrospective Forecasting Case Study

    cs.CL 2026-01 conditional novelty 6.0

    Date filters on major search engines frequently leak post-cutoff information, inflating Brier scores in retrospective forecasting from 0.24 to 0.10.

  10. Harnessing Pre-Resolution Signals for Future Prediction Agents

    cs.AI 2026-04 unverdicted novelty 5.0

    Milkyway uses pre-resolution signals from temporal contrasts in evolving evidence and repeated forecasts to evolve a harness and improve predictions before resolution, outperforming baselines on FutureX and FutureWorld.