Pith. sign in

Futurex: An advanced live benchmark for llm agents in future prediction

13 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.

13 Pith papers citing it
1 external citations · external index

citation-role summary

background 1 baseline 1

citation-polarity summary

years

2026 12 2025 1

representative citing papers

KellyBench: A Benchmark for Long-Horizon Sequential Decision Making

cs.AI · 2026-04-30 · unverdicted · novelty 7.0

KellyBench reveals that frontier language models lose money on average when making sequential betting decisions over a full soccer season, with the best model returning -8% and scoring only 26.5% on a human expert rubric for strategy sophistication.

Harnessing Pre-Resolution Signals for Future Prediction Agents

cs.AI · 2026-04-17 · unverdicted · novelty 5.0 · 2 refs

Milkyway uses pre-resolution signals from temporal contrasts in evolving evidence and repeated forecasts to evolve a harness and improve predictions before resolution, outperforming baselines on FutureX and FutureWorld.

citing papers explorer

Showing 13 of 13 citing papers.