REVIEW 3 major objections 6 minor 38 references
AI forecasters match the market, then bet to -18% to +10% ROI
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
On 104 World Cup matches, four LLM forecasting agents make identical top picks in 92% of matches, none beats the betting market's Brier score, but their betting ROI spans -18% to +10%.
T0 review reviewed 2026-08-01 challenge →
load-bearing objection A genuinely useful contamination-free LLM forecasting benchmark with a real market baseline, but the economic scoring rests on opening lines from one book, so treat the ROI and 'beats the market' numbers as provisional. the 3 major comments →
FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
On the paper's own terms, the discovery is that four independently built frontier LLM agents, each running the same retrieval-and-reasoning loop on the same public information, behave as near-clones when asked what will happen but as distinct individuals when asked what to do. Their top picks coincide in 92% of matches, their accuracy spans only three points, and every agent's probability estimates track the market-implied probabilities with correlations of 0.97–0.99; as forecasters, none beats the market's Brier score (best agent 0.471 vs market 0.469). Yet scored as bettors at real 1X2 odds, their return on investment spans -18.1% to +10.3%, a spread driven not by which side they pick but
What carries the argument
The load-bearing device is a fixed search–act–reflect loop: identical prompts instruct each agent to gather web evidence, output a 1X2 probability distribution and a staking decision, and later reflect given only the final score. The pre-match betting market, with the vigorish removed from published 1X2 odds, is the machinery that makes the separation visible: it acts as an independent fifth forecaster, as the settlement price for every virtual bet, and as a price anchor that the agents either follow (market-conforming bets) or override (contrarian bets). Together they allow the same 104 events to be scored on three axes—calibration (Brier, log-loss, ECE), decision quality (ROI and hit-rate
Load-bearing premise
The economic scoring and the market baseline both depend on treating pre-match bookmaker odds—collected mostly as opening lines from a single public sportsbook—as the price at which every virtual bet settles and as the independent 'market' competitor; if those lines were not actually available to the agents at the time of betting, drifted before kickoff, or were distorted by the single manually corrected line, the ROI spread and the market-beats-agents finding would be artifa
What would settle it
Run the identical search–act–reflect protocol on a future tournament (or re-settle the same 104 matches) using closing consensus odds instead of opening lines from one book; if any agent then beats the market's Brier score, or if the ROI spread collapses to noise, the claim that prediction quality and decision quality are separable axes would be undercut.
If this is right
- Selecting an LLM for a forecasting-and-acting pipeline should weigh calibration, staking discipline, and self-assessment, not headline accuracy.
- Any benchmark claiming super-human forecasting must include an economically grounded market baseline, since web-enabled agents here match but do not beat a vig-bearing bookmaker.
- The same protocol transfers to any scheduled, market-priced event stream (elections, product launches), enabling live, rolling, contamination-free evaluation of successive model releases.
- Because agents sharing a retrieval surface share errors, ensembling them inherits rather than cancels the common bias, limiting 'wisdom of the silicon crowd' when the crowd reads the same web.
- The released stakes allow other researchers to re-score the same forecasts under arbitrary staking policies, testing whether decision quality is a fixed trait or a policy artifact.
Where Pith is reading between the lines
- If the 92% pick consensus is driven by all agents retrieving the same market odds, then a search-free variant of the benchmark would likely reveal wider disagreement and might expose private reasoning; this is testable with the released dataset by re-running agents without web access.
- If the ROI spread replicates on other tournaments, staking discipline would look like a trainable skill: an agent could be optimized to size bets separately from predicting outcomes, with potentially large economic value.
- Because the market baseline uses mostly opening lines from one book, the 'no agent beats the market' result is probably conservative; closing consensus odds are typically more efficient and would be even harder to beat.
- The self-knowledge axis rests on a single lexical coding of free-text reflections; prompting models for numeric confidence before revealing the outcome, or coding reflections with a second scheme, would test whether the 36–86% spread is a stable trait or an artifact of the prompt.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces WC2026-Agents, a benchmark and dataset for evaluating LLM forecasting agents on all 104 matches of the 2026 FIFA World Cup. Four frontier models (Claude Opus 4.8, ChatGPT GPT-5.5, Gemini 3.1 Pro, Grok Expert Mode) run the same search-act-reflect loop: at T-24h they produce a 1X2 probability distribution and a virtual $100 bet, and after the match they reflect on their own forecast. Because every match kicked off after the models' training cutoffs, the benchmark is claimed to be contamination-free by construction. The release contains 416 forecasts, 414 reflections, per-match pre-match 1X2 odds, ground truth, and a one-command reproducible evaluation pipeline. The reference evaluation reports that the agents choose the same top outcome in 92% of matches, that none beats the de-vigged market's Brier score (0.4688 vs best agent 0.4705), and that betting ROI ranges from -18.1% (Claude) to +10.3% (Grok), while a flat stake on the market favorite would have returned +$1,041. The paper interprets these results as separating predictive calibration, decision quality, and self-knowledge.
Significance. If the results hold, this is a valuable benchmark contribution. The contamination-free design is credible, the release is unusually complete (verbatim reasoning, cited URLs, parse-issue flags, one-command pipeline), and the pairing with a market baseline is a genuinely useful template for evaluating agentic forecasting. The finding that models converge on picks yet differ in staking and self-assessment is interesting and likely robust. However, the central economic comparisons rest on opening lines from mostly one book, and the headline Brier/ROI differences are accompanied by no uncertainty quantification. Both are addressable, but until they are, the strong claims about 'no agent beats the market' and the ROI spread should be treated as provisional.
major comments (3)
- [§4.5, §10; Tables 2–3; §6.2, §6.4–6.5] The market baseline is load-bearing, and its provenance is too thin. §4.5 says the odds are opening lines from mostly one public sportsbook, with one line manually corrected, and the only external validation is a single match (Brazil–Japan, matching a cited '56/25/19'). The Limitation section explicitly concedes 'opening lines from mostly one book, not closing consensus.' If these lines drifted before kickoff or were not tradable at T-24h, then (i) the 'market' Brier in Table 2 is not necessarily the efficient price the paper claims, and (ii) every ROI number in Table 3 and the flat-favorite baseline in §6.4 are settled at prices the agents could not actually trade. Please add systematic evidence (e.g., comparison to closing consensus from multiple books, drift distributions, tradability checks), or re-scope all claims to 'against the recorded opening lines.'
- [Table 2 / §6.2] The claim that no agent beats the market is based on a Brier gap of 0.4688 vs 0.4705/0.4706 over 104 matches. No confidence intervals or significance tests are reported; a difference of ~0.0017 is within plausible sampling noise. The same holds for the ECE ordering (§6.3) and the ROI spread (Table 3), where the paper reports point estimates only. Please add paired comparisons (e.g., sign test or bootstrap on per-match Brier differences) and uncertainty intervals. Where intervals overlap, soften statements such as 'none beats the market' and the efficient-market interpretation in §6.2.
- [§4.1–4.2, §6.1] The 'identical search–act–reflect loop' is byte-identical in prompt text, but the agents run through different consumer interfaces and web tools, and §10 acknowledges that providers may update models mid-window. This is acceptable for a deployed-agent study, but the word 'identical' overstates control. More substantively, because the loop includes web search, agents can read the same opening odds that later define the market baseline; the near-perfect correlation in Fig. 2 (r = 0.97–0.99) is therefore partly a property of the benchmark design. The paper mostly acknowledges this, but §6.1's conclusion should be phrased as 'agents recover the information in the market baseline,' not as independent evidence for market efficiency.
minor comments (6)
- [§6.1] Typo: 'fouragents' should be 'four agents.'
- [Fig. 1] The bar chart counts are hard to read; please add a small table with unanimous correct / unanimous wrong / split counts by stage.
- [§4.5] Specify whether the reported mean overround of 1.05 is arithmetic or weighted, and report the range across matches.
- [Abstract/intro] The abstract's 'none beats the market's Brier score' should carry the qualifier 'on this tournament's recorded opening-line baseline,' consistent with the limitation.
- [§6.4] The sentence 'net results of -$275 to +$650' is clear, but the comparison to the flat-favorite baseline should state explicitly that this baseline is not a single agent and is computed after the fact.
- [§6.6] The Pearson correlations between stake and confidence are reported without CIs; given n=104 and the multiple comparisons across four agents, consider reporting bootstrap intervals.
Circularity Check
Mild circular validation of market odds via agents' own citations; central benchmark comparisons are otherwise independent empirical evaluations.
specific steps
-
other
[Section 4.5, 'Market odds' paragraph]
"we validated them against the market probabilities the agents themselves cited in their reasoning (e.g. our de-vigged Brazil–Japan line, 55/25/20, matches a cited “56/25/19”)"
The market odds are validated using the agents' cited probabilities, but these agents are the very subjects being scored against the market. Section 6.1 shows agent probabilities track the market almost perfectly (r=0.97–0.99), so the agents are 'to first order, reading and lightly re-expressing the same price.' An agent's citation is therefore not an independent confirmation of the odds; it is the same market information looped through the model. This validation cannot establish that the opening-line odds were current or tradable, and it is a self-referential data-quality check. However, the central findings do not derive from this validation—the odds come from public sportsbook previews with source URLs—so this is a minor non-load-bearing circularity rather than a definitional equivalenc
full rationale
The paper's central claims are empirical comparisons against an external market baseline (bookmaker odds), not derivations from assumptions. The four agents' forecasts and the market's implied probabilities are independently collected; Brier score, ROI, and reflection metrics are straightforward scoring rules applied after the fact. No parameter is fitted to the data and then renamed as a prediction. The 'contamination-free' claim is a temporal property (matches occurred after training cutoffs), not a circular derivation. The main circularity-adjacent issue is in §4.5, where the market odds are validated using the agents' own cited probabilities; since §6.1 shows the agents' probabilities track the market almost perfectly, this validation is self-referential and cannot independently confirm the odds. It is a quality check, not a load-bearing derivation: the odds themselves come from public previews, and the core findings stand or fall on the odds' accuracy, which the paper itself flags as a limitation ('Odds are opening lines from mostly one book, not closing consensus'). Additionally, the protocol lets agents read the market during search, which makes 'no agent beats the market' partly a benchmark-design property; however, it is not a definitional equivalence because agents can and do diverge (Gemini cites the market only 12% of the time). Overall, the derivation chain is self-contained, with only minor, non-load-bearing circular elements.
Axiom & Free-Parameter Ledger
axioms (6)
- domain assumption De-vigged published odds approximate true match probabilities.
- domain assumption Opening lines from mostly one public sportsbook are a valid settlement price for bets placed ~24h before kickoff.
- domain assumption The 104 fixtures, scores, and advancement outcomes are real and correctly hand-resolved.
- domain assumption The four proprietary assistants' training cutoffs predate the tournament and no mid-window backend updates occurred.
- domain assumption Virtual betting profit at the released odds measures decision quality.
- domain assumption Reflection labels (outcome_vs_prediction) are a meaningful proxy for self-knowledge.
Cite this review
Pith. "Pith review of FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches." pith.science (2026). https://pith.science/paper/DNKGNBEN
@misc{pith2026260717765,
author = {Pith},
title = {Pith review of: FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches},
year = {2026},
howpublished = {\url{https://pith.science/paper/DNKGNBEN}},
note = {Machine review of arXiv:2607.17765}
}
read the original abstract
We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events. For every one of the 104 matches of the 2026 FIFA World Cup, four frontier models -- Claude Opus 4.8, ChatGPT (GPT-5.5, high reasoning), Gemini 3.1 Pro, and Grok (Expert Mode) -- ran an identical search-act-reflect loop: gather evidence with a web tool, commit to a 1X2 (team-A win / draw / team-B win) distribution and a virtual 100-USD bet, and, after the match, reflect given only the final score. Because every match kicked off after the models' training cutoffs, the benchmark is contamination-free by construction. Crucially, we pair the four agents with a fifth competitor drawn from the same information environment -- the pre-match betting market -- collected as per-match 1X2 odds, giving an economically grounded baseline and letting us score not just what an agent predicts but what it does with money. The release contains 416 forecasts and 414 reflections with verbatim reasoning, ground truth (including penalty shootouts), odds, and a reproducible evaluation suite. A reference evaluation surfaces findings that raw accuracy hides: the four agents issue an identical top pick in 92% of matches and none beats the market's Brier score; indeed, a naive flat stake on the market favorite out-earns all four agents. Yet the agents diverge sharply as decision-makers: betting return-on-investment ranges from -18% to +10%, fading the market is unprofitable for all four, the share of forecasts that cite the market ranges from 12% to 100%, and self-reported error rates on wrong picks range from 36% to 86%. The benchmark thus measures calibration, decision quality, and self-knowledge -- axes on which frontier models differ even when their predictions do not. Data and code: https://github.com/graphuofm/FIFA2026LLM
Figures
Reference graph
Works this paper leans on
-
[1]
Giovanni Angelini and Luca De Angelis. 2019. Efficiency of online football betting markets.International Journal of Forecasting35, 2 (2019), 712–721. Ding, Guo, and Xu
2019
-
[2]
Anthropic. 2026. Claude Opus 4.8. https://www.anthropic.com/claude
2026
-
[3]
Rahul Baboota and Harleen Kaur. 2019. Predictive analysis and modelling football results using machine learning approach for English Premier League.Interna- tional Journal of Forecasting35, 2 (2019), 741–755
2019
-
[4]
Glenn W. Brier. 1950. Verification of forecasts expressed in terms of probability. Monthly Weather Review78, 1 (1950), 1–3
1950
-
[5]
Constantinou, Norman E
Anthony C. Constantinou, Norman E. Fenton, and Martin Neil. 2012. pi-football: A Bayesian network model for forecasting Association Football match outcomes. Knowledge-Based Systems36 (2012), 322–339
2012
-
[6]
Dietvorst, Joseph P
Berkeley J. Dietvorst, Joseph P. Simmons, and Cade Massey. 2015. Algorithm aversion: People erroneously avoid algorithms after seeing them err.Journal of Experimental Psychology: General144, 1 (2015), 114–126
2015
-
[7]
Eugene F. Fama. 1970. Efficient capital markets: A review of theory and empirical work.The Journal of Finance25, 2 (1970), 383–417
1970
-
[8]
Shahriar Golchin and Mihai Surdeanu. 2024. Time travel in LLMs: Tracing data contamination in large language models. InInternational Conference on Learning Representations (ICLR)
2024
-
[9]
Google DeepMind. 2026. Gemini 3.1 Pro model card. https://deepmind.google/ models/model-cards/gemini-3-1-pro/
2026
-
[10]
Weinberger
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks. InInternational Conference on Machine Learning (ICML). 1321–1330
2017
-
[11]
Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt. 2024. Ap- proaching human-level forecasting with language models.arXiv preprint arXiv:2402.18563(2024)
Pith/arXiv arXiv 2024
-
[12]
Dan Hendrycks, Collin Burns, Steven Basart, et al . 2021. Measuring massive multitask language understanding. InInternational Conference on Learning Rep- resentations (ICLR)
2021
-
[13]
Jie Huang, Xinyun Chen, Swaroop Mishra, et al . 2024. Large language mod- els cannot self-correct reasoning yet. InInternational Conference on Learning Representations (ICLR)
2024
-
[14]
Saurav Kadavath, Tom Conerly, Amanda Askell, et al. 2022. Language models (mostly) know what they know.arXiv preprint arXiv:2207.05221(2022)
Pith/arXiv arXiv 2022
-
[15]
Ezra Karger, Houtan Bastani, Chen Yueh-Han, et al . 2024. ForecastBench: A dynamic benchmark of AI forecasting capabilities.arXiv preprint arXiv:2409.19839 (2024)
Pith/arXiv arXiv 2024
-
[16]
Steven D. Levitt. 2004. Why are gambling markets organised so differently from financial markets?The Economic Journal114, 495 (2004), 223–246
2004
-
[17]
Percy Liang, Rishi Bommasani, Tony Lee, et al . 2023. Holistic evaluation of language models.Transactions on Machine Learning Research (TMLR)(2023)
2023
-
[18]
Aman Madaan, Niket Tandon, Prakhar Gupta, et al. 2023. Self-Refine: Iterative refinement with self-feedback. InAdvances in Neural Information Processing Systems (NeurIPS)
2023
-
[19]
Moore and Paul J
Don A. Moore and Paul J. Healy. 2008. The trouble with overconfidence.Psycho- logical Review115, 2 (2008), 502–517
2008
-
[20]
Cooper, and Milos Hauskrecht
Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht. 2015. Ob- taining well calibrated probabilities using Bayesian binning. InAAAI Conference on Artificial Intelligence
2015
-
[21]
OpenAI. 2026. GPT-5.5. https://openai.com/index/gpt-5-5-instant/
2026
-
[22]
O’Brien, Carrie J
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, et al. 2023. Generative agents: Interactive simulacra of human behavior. InACM Symposium on User Interface Software and Technology (UIST)
2023
-
[23]
Park, and Philip E
Philipp Schoenegger, Indre Tuminauskaite, Peter S. Park, and Philip E. Tetlock
-
[24]
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learn- ing. InAdvances in Neural Information Processing Systems (NeurIPS)
2023
-
[25]
Erik Snowberg and Justin Wolfers. 2010. Explaining the favorite–long shot bias: Is it risk-love or misperceptions?Journal of Political Economy118, 4 (2010), 723–746
2010
-
[26]
Tetlock and Dan Gardner
Philip E. Tetlock and Dan Gardner. 2015.Superforecasting: The Art and Science of Prediction. Crown
2015
-
[27]
Katherine Tian, Eric Mitchell, Allan Zhou, et al. 2023. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models. In Empirical Methods in Natural Language Processing (EMNLP)
2023
-
[28]
Erik Štrumbelj. 2014. On determining probability forecasts from betting odds. International Journal of Forecasting30, 4 (2014), 934–943
2014
-
[29]
Lei Wang, Chen Ma, Xueyang Feng, et al . 2024. A survey on large language model based autonomous agents.Frontiers of Computer Science18, 6 (2024)
2024
-
[30]
Jason Wei, Xuezhi Wang, Dale Schuurmans, et al. 2022. Chain-of-thought prompt- ing elicits reasoning in large language models. InAdvances in Neural Information Processing Systems (NeurIPS)
2022
-
[31]
Justin Wolfers and Eric Zitzewitz. 2004. Prediction markets.Journal of Economic Perspectives18, 2 (2004), 107–126
2004
-
[32]
Fabian Wunderlich and Daniel Memmert. 2018. The betting odds rating system: Using soccer forecasts to forecast soccer.PLOS ONE13, 6 (2018), e0198668
2018
-
[33]
xAI. 2026. Grok Expert Mode. https://x.ai
2026
-
[34]
Zhiheng Xi, Wenxiang Chen, Xin Guo, et al. 2023. The rise and potential of large language model based agents: A survey.arXiv preprint arXiv:2309.07864(2023)
Pith/arXiv arXiv 2023
-
[35]
Miao Xiong, Zhiyuan Hu, Xinyang Lu, et al . 2024. Can LLMs express their uncertainty? An empirical evaluation of confidence elicitation. InInternational Conference on Learning Representations (ICLR)
2024
-
[36]
Shunyu Yao, Jeffrey Zhao, Dian Yu, et al. 2023. ReAct: Synergizing reasoning and acting in language models. InInternational Conference on Learning Representations (ICLR)
2023
-
[37]
Andy Zou, Tristan Xiao, Ryan Jia, et al. 2022. Forecasting future world events with neural networks. InAdvances in Neural Information Processing Systems (NeurIPS). A Reproducibility One command (python src/run_all.py ) rebuilds every processed table, analysis CSV, and figure from the raw transcripts and odds; the paper’s numbers are read directly from tho...
2022
-
[2024]
Wisdom of the silicon crowd: LLM ensemble prediction capabilities rival human crowd accuracy.Science Advances10, 45 (2024)
2024
This paper was first reviewed by deepseek-v4-flash on August 1, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.