Agentic Time Machine reconstructs historical web states for offline evaluation of forecasting agents, with a multi-agent framework achieving top ranks on FutureX live and past benchmarks.
arXiv preprint arXiv:2506.21558 , year=
7 Pith papers cite this work. Polarity classification is still indexing.
years
2026 7representative citing papers
ForecastBench-Sim is a simulated-world benchmark using Freeciv game rollouts to generate resolvable forecasting questions at arbitrary horizons with paired intervention worlds.
WorldReasoner supplies 345 resolved forecasting tasks built from 14,141 articles to score LM agents on outcome quality, evidence quality, and reasoning quality against time-bounded evidence and hindsight graphs.
More capable LLMs produce worse distributional forecasts on superlinear growth time series with tail risks of regime change, with the error concentrated in the upper tail; this reverses on conventional threshold metrics.
BTF-2 benchmark shows frontier AI forecasters lag humans mainly in assessing political and business leaders' incentives, plan follow-through likelihood, and institutional processes, with a composite agent gaining 0.011 Brier by focusing on blind spots.
Date filters on major search engines frequently leak post-cutoff information, inflating Brier scores in retrospective forecasting from 0.24 to 0.10.
Milkyway uses pre-resolution signals from temporal contrasts in evolving evidence and repeated forecasts to evolve a harness and improve predictions before resolution, outperforming baselines on FutureX and FutureWorld.
citing papers explorer
-
Agentic Time Machine as an Infrastructure for Future-Event Forecasting
Agentic Time Machine reconstructs historical web states for offline evaluation of forecasting agents, with a multi-agent framework achieving top ranks on FutureX live and past benchmarks.
-
ForecastBench-Sim: A Simulated-World Forecasting Benchmark
ForecastBench-Sim is a simulated-world benchmark using Freeciv game rollouts to generate resolvable forecasting questions at arbitrary horizons with paired intervention worlds.
-
WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning
WorldReasoner supplies 345 resolved forecasting tasks built from 14,141 articles to score LM agents on outcome quality, evidence quality, and reasoning quality against time-bounded evidence and hindsight graphs.
-
Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most
More capable LLMs produce worse distributional forecasts on superlinear growth time series with tail risks of regime change, with the error concentrated in the upper tail; this reverses on conventional threshold metrics.
-
Evaluating Strategic Reasoning in Forecasting Agents
BTF-2 benchmark shows frontier AI forecasters lag humans mainly in assessing political and business leaders' incentives, plan follow-through likelihood, and institutional processes, with a composite agent gaining 0.011 Brier by focusing on blind spots.
-
Temporal Leakage in Search-Engine Date-Filtered Web Retrieval: A Retrospective Forecasting Case Study
Date filters on major search engines frequently leak post-cutoff information, inflating Brier scores in retrospective forecasting from 0.24 to 0.10.
-
Harnessing Pre-Resolution Signals for Future Prediction Agents
Milkyway uses pre-resolution signals from temporal contrasts in evolving evidence and repeated forecasts to evolve a harness and improve predictions before resolution, outperforming baselines on FutureX and FutureWorld.