Pith. sign in

REVIEW 3 major objections 6 minor 38 references

AI forecasters match the market, then bet to -18% to +10% ROI

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 17:03 UTC pith:DNKGNBEN

load-bearing objection A genuinely useful contamination-free LLM forecasting benchmark with a real market baseline, but the economic scoring rests on opening lines from one book, so treat the ROI and 'beats the market' numbers as provisional. the 3 major comments →

arxiv 2607.17765 v1 pith:DNKGNBEN submitted 2026-07-20 cs.LG cs.AIcs.DB

FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches

classification cs.LG cs.AIcs.DB
keywords LLM forecasting agentscontamination-free benchmarkcalibrationdecision qualityself-knowledgebetting marketBrier scoreWorld Cup 2026
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper introduces a contamination-free benchmark built from all 104 matches of the 2026 FIFA World Cup, in which four frontier LLMs ran an identical search–act–reflect loop: search the web, commit to a 1X2 probability distribution and a virtual $100 bet, then reflect on the outcome. The pre-match betting market is treated as a fifth competitor, serving both as an independent baseline and as the settlement price for every virtual bet. The central claim is that predictive accuracy, decision quality, and self-knowledge are separable axes: the agents agree on the most likely winner in 92% of matches and none beats the market's Brier score, yet their betting returns range from -18% to +10%, their rate of citing the market ranges from 12% to 100%, and their self-reported error rates on wrong picks range from 36% to 86%. A sympathetic reader would care because headline accuracy alone would have reported a null result; here it hides large, stable behavioral differences that matter when a model is actually deployed to act on a prediction.

Core claim

On the paper's own terms, the discovery is that four independently built frontier LLM agents, each running the same retrieval-and-reasoning loop on the same public information, behave as near-clones when asked what will happen but as distinct individuals when asked what to do. Their top picks coincide in 92% of matches, their accuracy spans only three points, and every agent's probability estimates track the market-implied probabilities with correlations of 0.97–0.99; as forecasters, none beats the market's Brier score (best agent 0.471 vs market 0.469). Yet scored as bettors at real 1X2 odds, their return on investment spans -18.1% to +10.3%, a spread driven not by which side they pick but

What carries the argument

The load-bearing device is a fixed search–act–reflect loop: identical prompts instruct each agent to gather web evidence, output a 1X2 probability distribution and a staking decision, and later reflect given only the final score. The pre-match betting market, with the vigorish removed from published 1X2 odds, is the machinery that makes the separation visible: it acts as an independent fifth forecaster, as the settlement price for every virtual bet, and as a price anchor that the agents either follow (market-conforming bets) or override (contrarian bets). Together they allow the same 104 events to be scored on three axes—calibration (Brier, log-loss, ECE), decision quality (ROI and hit-rate

Load-bearing premise

The economic scoring and the market baseline both depend on treating pre-match bookmaker odds—collected mostly as opening lines from a single public sportsbook—as the price at which every virtual bet settles and as the independent 'market' competitor; if those lines were not actually available to the agents at the time of betting, drifted before kickoff, or were distorted by the single manually corrected line, the ROI spread and the market-beats-agents finding would be artifa

What would settle it

Run the identical search–act–reflect protocol on a future tournament (or re-settle the same 104 matches) using closing consensus odds instead of opening lines from one book; if any agent then beats the market's Brier score, or if the ROI spread collapses to noise, the claim that prediction quality and decision quality are separable axes would be undercut.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Selecting an LLM for a forecasting-and-acting pipeline should weigh calibration, staking discipline, and self-assessment, not headline accuracy.
  • Any benchmark claiming super-human forecasting must include an economically grounded market baseline, since web-enabled agents here match but do not beat a vig-bearing bookmaker.
  • The same protocol transfers to any scheduled, market-priced event stream (elections, product launches), enabling live, rolling, contamination-free evaluation of successive model releases.
  • Because agents sharing a retrieval surface share errors, ensembling them inherits rather than cancels the common bias, limiting 'wisdom of the silicon crowd' when the crowd reads the same web.
  • The released stakes allow other researchers to re-score the same forecasts under arbitrary staking policies, testing whether decision quality is a fixed trait or a policy artifact.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the 92% pick consensus is driven by all agents retrieving the same market odds, then a search-free variant of the benchmark would likely reveal wider disagreement and might expose private reasoning; this is testable with the released dataset by re-running agents without web access.
  • If the ROI spread replicates on other tournaments, staking discipline would look like a trainable skill: an agent could be optimized to size bets separately from predicting outcomes, with potentially large economic value.
  • Because the market baseline uses mostly opening lines from one book, the 'no agent beats the market' result is probably conservative; closing consensus odds are typically more efficient and would be even harder to beat.
  • The self-knowledge axis rests on a single lexical coding of free-text reflections; prompting models for numeric confidence before revealing the outcome, or coding reflections with a second scheme, would test whether the 36–86% spread is a stable trait or an artifact of the prompt.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces WC2026-Agents, a benchmark and dataset for evaluating LLM forecasting agents on all 104 matches of the 2026 FIFA World Cup. Four frontier models (Claude Opus 4.8, ChatGPT GPT-5.5, Gemini 3.1 Pro, Grok Expert Mode) run the same search-act-reflect loop: at T-24h they produce a 1X2 probability distribution and a virtual $100 bet, and after the match they reflect on their own forecast. Because every match kicked off after the models' training cutoffs, the benchmark is claimed to be contamination-free by construction. The release contains 416 forecasts, 414 reflections, per-match pre-match 1X2 odds, ground truth, and a one-command reproducible evaluation pipeline. The reference evaluation reports that the agents choose the same top outcome in 92% of matches, that none beats the de-vigged market's Brier score (0.4688 vs best agent 0.4705), and that betting ROI ranges from -18.1% (Claude) to +10.3% (Grok), while a flat stake on the market favorite would have returned +$1,041. The paper interprets these results as separating predictive calibration, decision quality, and self-knowledge.

Significance. If the results hold, this is a valuable benchmark contribution. The contamination-free design is credible, the release is unusually complete (verbatim reasoning, cited URLs, parse-issue flags, one-command pipeline), and the pairing with a market baseline is a genuinely useful template for evaluating agentic forecasting. The finding that models converge on picks yet differ in staking and self-assessment is interesting and likely robust. However, the central economic comparisons rest on opening lines from mostly one book, and the headline Brier/ROI differences are accompanied by no uncertainty quantification. Both are addressable, but until they are, the strong claims about 'no agent beats the market' and the ROI spread should be treated as provisional.

major comments (3)
  1. [§4.5, §10; Tables 2–3; §6.2, §6.4–6.5] The market baseline is load-bearing, and its provenance is too thin. §4.5 says the odds are opening lines from mostly one public sportsbook, with one line manually corrected, and the only external validation is a single match (Brazil–Japan, matching a cited '56/25/19'). The Limitation section explicitly concedes 'opening lines from mostly one book, not closing consensus.' If these lines drifted before kickoff or were not tradable at T-24h, then (i) the 'market' Brier in Table 2 is not necessarily the efficient price the paper claims, and (ii) every ROI number in Table 3 and the flat-favorite baseline in §6.4 are settled at prices the agents could not actually trade. Please add systematic evidence (e.g., comparison to closing consensus from multiple books, drift distributions, tradability checks), or re-scope all claims to 'against the recorded opening lines.'
  2. [Table 2 / §6.2] The claim that no agent beats the market is based on a Brier gap of 0.4688 vs 0.4705/0.4706 over 104 matches. No confidence intervals or significance tests are reported; a difference of ~0.0017 is within plausible sampling noise. The same holds for the ECE ordering (§6.3) and the ROI spread (Table 3), where the paper reports point estimates only. Please add paired comparisons (e.g., sign test or bootstrap on per-match Brier differences) and uncertainty intervals. Where intervals overlap, soften statements such as 'none beats the market' and the efficient-market interpretation in §6.2.
  3. [§4.1–4.2, §6.1] The 'identical search–act–reflect loop' is byte-identical in prompt text, but the agents run through different consumer interfaces and web tools, and §10 acknowledges that providers may update models mid-window. This is acceptable for a deployed-agent study, but the word 'identical' overstates control. More substantively, because the loop includes web search, agents can read the same opening odds that later define the market baseline; the near-perfect correlation in Fig. 2 (r = 0.97–0.99) is therefore partly a property of the benchmark design. The paper mostly acknowledges this, but §6.1's conclusion should be phrased as 'agents recover the information in the market baseline,' not as independent evidence for market efficiency.
minor comments (6)
  1. [§6.1] Typo: 'fouragents' should be 'four agents.'
  2. [Fig. 1] The bar chart counts are hard to read; please add a small table with unanimous correct / unanimous wrong / split counts by stage.
  3. [§4.5] Specify whether the reported mean overround of 1.05 is arithmetic or weighted, and report the range across matches.
  4. [Abstract/intro] The abstract's 'none beats the market's Brier score' should carry the qualifier 'on this tournament's recorded opening-line baseline,' consistent with the limitation.
  5. [§6.4] The sentence 'net results of -$275 to +$650' is clear, but the comparison to the flat-favorite baseline should state explicitly that this baseline is not a single agent and is computed after the fact.
  6. [§6.6] The Pearson correlations between stake and confidence are reported without CIs; given n=104 and the multiple comparisons across four agents, consider reporting bootstrap intervals.

Circularity Check

1 steps flagged

Mild circular validation of market odds via agents' own citations; central benchmark comparisons are otherwise independent empirical evaluations.

specific steps
  1. other [Section 4.5, 'Market odds' paragraph]
    "we validated them against the market probabilities the agents themselves cited in their reasoning (e.g. our de-vigged Brazil–Japan line, 55/25/20, matches a cited “56/25/19”)"

    The market odds are validated using the agents' cited probabilities, but these agents are the very subjects being scored against the market. Section 6.1 shows agent probabilities track the market almost perfectly (r=0.97–0.99), so the agents are 'to first order, reading and lightly re-expressing the same price.' An agent's citation is therefore not an independent confirmation of the odds; it is the same market information looped through the model. This validation cannot establish that the opening-line odds were current or tradable, and it is a self-referential data-quality check. However, the central findings do not derive from this validation—the odds come from public sportsbook previews with source URLs—so this is a minor non-load-bearing circularity rather than a definitional equivalenc

full rationale

The paper's central claims are empirical comparisons against an external market baseline (bookmaker odds), not derivations from assumptions. The four agents' forecasts and the market's implied probabilities are independently collected; Brier score, ROI, and reflection metrics are straightforward scoring rules applied after the fact. No parameter is fitted to the data and then renamed as a prediction. The 'contamination-free' claim is a temporal property (matches occurred after training cutoffs), not a circular derivation. The main circularity-adjacent issue is in §4.5, where the market odds are validated using the agents' own cited probabilities; since §6.1 shows the agents' probabilities track the market almost perfectly, this validation is self-referential and cannot independently confirm the odds. It is a quality check, not a load-bearing derivation: the odds themselves come from public previews, and the core findings stand or fall on the odds' accuracy, which the paper itself flags as a limitation ('Odds are opening lines from mostly one book, not closing consensus'). Additionally, the protocol lets agents read the market during search, which makes 'no agent beats the market' partly a benchmark-design property; however, it is not a definitional equivalence because agents can and do diverge (Gemini cites the market only 12% of the time). Overall, the derivation chain is self-contained, with only minor, non-load-bearing circular elements.

Axiom & Free-Parameter Ledger

0 free parameters · 6 axioms · 0 invented entities

No numerical parameters were fitted to the outcome data; the ledger records the domain assumptions the benchmark's validity rests on.

axioms (6)
  • domain assumption De-vigged published odds approximate true match probabilities.
    Used in §4.5 to construct the market baseline; mean overround 1.05 is reported, but the de-vigging method is not specified beyond removing the overround.
  • domain assumption Opening lines from mostly one public sportsbook are a valid settlement price for bets placed ~24h before kickoff.
    Section 4.5 and Section 10 state knockout odds are opening lines from public previews; P&L in Table 3 is settled against these lines without evidence they were tradable at bet time.
  • domain assumption The 104 fixtures, scores, and advancement outcomes are real and correctly hand-resolved.
    Section 4.4 hand-resolved knockout transcripts and cross-checked four ways; this is a data-integrity claim, not independently verifiable from the paper.
  • domain assumption The four proprietary assistants' training cutoffs predate the tournament and no mid-window backend updates occurred.
    Abstract and §4.3 claim contamination-free 'by construction'; for closed consumer interfaces this is unverifiable and only partially acknowledged in Section 10.
  • domain assumption Virtual betting profit at the released odds measures decision quality.
    Section 5 (T2) assumes odds were available at the time of each agent's bet and that there are no execution constraints, which is not established.
  • domain assumption Reflection labels (outcome_vs_prediction) are a meaningful proxy for self-knowledge.
    Section 4.1 and §6.8 treat the model's free-text self-label as honesty; it could also reflect prompt sycophancy or mislabeling.

pith-pipeline@v1.3.0-alltime-deepseek · 11158 in / 12064 out tokens · 138081 ms · 2026-08-01T17:03:05.161099+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches." pith.science (2026). https://pith.science/paper/DNKGNBEN

@misc{pith2026260717765,
  author       = {Pith},
  title        = {Pith review of: FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DNKGNBEN}},
  note         = {Machine review of arXiv:2607.17765}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events. For every one of the 104 matches of the 2026 FIFA World Cup, four frontier models -- Claude Opus 4.8, ChatGPT (GPT-5.5, high reasoning), Gemini 3.1 Pro, and Grok (Expert Mode) -- ran an identical search-act-reflect loop: gather evidence with a web tool, commit to a 1X2 (team-A win / draw / team-B win) distribution and a virtual 100-USD bet, and, after the match, reflect given only the final score. Because every match kicked off after the models' training cutoffs, the benchmark is contamination-free by construction. Crucially, we pair the four agents with a fifth competitor drawn from the same information environment -- the pre-match betting market -- collected as per-match 1X2 odds, giving an economically grounded baseline and letting us score not just what an agent predicts but what it does with money. The release contains 416 forecasts and 414 reflections with verbatim reasoning, ground truth (including penalty shootouts), odds, and a reproducible evaluation suite. A reference evaluation surfaces findings that raw accuracy hides: the four agents issue an identical top pick in 92% of matches and none beats the market's Brier score; indeed, a naive flat stake on the market favorite out-earns all four agents. Yet the agents diverge sharply as decision-makers: betting return-on-investment ranges from -18% to +10%, fading the market is unprofitable for all four, the share of forecasts that cite the market ranges from 12% to 100%, and self-reported error rates on wrong picks range from 36% to 86%. The benchmark thus measures calibration, decision quality, and self-knowledge -- axes on which frontier models differ even when their predictions do not. Data and code: https://github.com/graphuofm/FIFA2026LLM

Figures

Figures reproduced from arXiv: 2607.17765 by Cong Guo, Jason Xu, Jiacheng Ding.

Figure 1
Figure 1. Figure 1: Convergence. Left: pairwise same-pick rate. Right: match counts by consensus; genuine disagreement is rare and [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Each agent’s probability for a team-A win against the market-implied probability, all 104 matches. Points hug the [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Cumulative virtual profit over the 104-match tournament, each agent settling its own stakes at real 1X2 odds. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 5
Figure 5. Figure 5: Share of pre-match forecasts citing each factor, per [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 4
Figure 4. Figure 4: Against-the-market analysis. Top: share of each [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 6
Figure 6. Figure 6: The eight matches mispredicted by all four agents. Every agent backed the favourite; four ties were lost on penalties [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Self-assessment on each agent’s own wrong picks. [PITH_FULL_IMAGE:figures/full_fig_p007_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Reliability of each agent’s top pick (App.). All lie [PITH_FULL_IMAGE:figures/full_fig_p008_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

38 extracted references · 4 linked inside Pith

  1. [1]

    Giovanni Angelini and Luca De Angelis. 2019. Efficiency of online football betting markets.International Journal of Forecasting35, 2 (2019), 712–721. Ding, Guo, and Xu

  2. [2]

    Anthropic. 2026. Claude Opus 4.8. https://www.anthropic.com/claude

  3. [3]

    Rahul Baboota and Harleen Kaur. 2019. Predictive analysis and modelling football results using machine learning approach for English Premier League.Interna- tional Journal of Forecasting35, 2 (2019), 741–755

  4. [4]

    Glenn W. Brier. 1950. Verification of forecasts expressed in terms of probability. Monthly Weather Review78, 1 (1950), 1–3

  5. [5]

    Constantinou, Norman E

    Anthony C. Constantinou, Norman E. Fenton, and Martin Neil. 2012. pi-football: A Bayesian network model for forecasting Association Football match outcomes. Knowledge-Based Systems36 (2012), 322–339

  6. [6]

    Dietvorst, Joseph P

    Berkeley J. Dietvorst, Joseph P. Simmons, and Cade Massey. 2015. Algorithm aversion: People erroneously avoid algorithms after seeing them err.Journal of Experimental Psychology: General144, 1 (2015), 114–126

  7. [7]

    Eugene F. Fama. 1970. Efficient capital markets: A review of theory and empirical work.The Journal of Finance25, 2 (1970), 383–417

  8. [8]

    Shahriar Golchin and Mihai Surdeanu. 2024. Time travel in LLMs: Tracing data contamination in large language models. InInternational Conference on Learning Representations (ICLR)

  9. [9]

    Google DeepMind. 2026. Gemini 3.1 Pro model card. https://deepmind.google/ models/model-cards/gemini-3-1-pro/

  10. [10]

    Weinberger

    Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks. InInternational Conference on Machine Learning (ICML). 1321–1330

  11. [11]

    Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt. 2024. Ap- proaching human-level forecasting with language models.arXiv preprint arXiv:2402.18563(2024)

  12. [12]

    Dan Hendrycks, Collin Burns, Steven Basart, et al . 2021. Measuring massive multitask language understanding. InInternational Conference on Learning Rep- resentations (ICLR)

  13. [13]

    Jie Huang, Xinyun Chen, Swaroop Mishra, et al . 2024. Large language mod- els cannot self-correct reasoning yet. InInternational Conference on Learning Representations (ICLR)

  14. [14]

    Saurav Kadavath, Tom Conerly, Amanda Askell, et al. 2022. Language models (mostly) know what they know.arXiv preprint arXiv:2207.05221(2022)

  15. [15]

    Ezra Karger, Houtan Bastani, Chen Yueh-Han, et al . 2024. ForecastBench: A dynamic benchmark of AI forecasting capabilities.arXiv preprint arXiv:2409.19839 (2024)

  16. [16]

    Steven D. Levitt. 2004. Why are gambling markets organised so differently from financial markets?The Economic Journal114, 495 (2004), 223–246

  17. [17]

    Percy Liang, Rishi Bommasani, Tony Lee, et al . 2023. Holistic evaluation of language models.Transactions on Machine Learning Research (TMLR)(2023)

  18. [18]

    Aman Madaan, Niket Tandon, Prakhar Gupta, et al. 2023. Self-Refine: Iterative refinement with self-feedback. InAdvances in Neural Information Processing Systems (NeurIPS)

  19. [19]

    Moore and Paul J

    Don A. Moore and Paul J. Healy. 2008. The trouble with overconfidence.Psycho- logical Review115, 2 (2008), 502–517

  20. [20]

    Cooper, and Milos Hauskrecht

    Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht. 2015. Ob- taining well calibrated probabilities using Bayesian binning. InAAAI Conference on Artificial Intelligence

  21. [21]

    OpenAI. 2026. GPT-5.5. https://openai.com/index/gpt-5-5-instant/

  22. [22]

    O’Brien, Carrie J

    Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, et al. 2023. Generative agents: Interactive simulacra of human behavior. InACM Symposium on User Interface Software and Technology (UIST)

  23. [23]

    Park, and Philip E

    Philipp Schoenegger, Indre Tuminauskaite, Peter S. Park, and Philip E. Tetlock

  24. [24]

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learn- ing. InAdvances in Neural Information Processing Systems (NeurIPS)

  25. [25]

    Erik Snowberg and Justin Wolfers. 2010. Explaining the favorite–long shot bias: Is it risk-love or misperceptions?Journal of Political Economy118, 4 (2010), 723–746

  26. [26]

    Tetlock and Dan Gardner

    Philip E. Tetlock and Dan Gardner. 2015.Superforecasting: The Art and Science of Prediction. Crown

  27. [27]

    Katherine Tian, Eric Mitchell, Allan Zhou, et al. 2023. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models. In Empirical Methods in Natural Language Processing (EMNLP)

  28. [28]

    Erik Štrumbelj. 2014. On determining probability forecasts from betting odds. International Journal of Forecasting30, 4 (2014), 934–943

  29. [29]

    Lei Wang, Chen Ma, Xueyang Feng, et al . 2024. A survey on large language model based autonomous agents.Frontiers of Computer Science18, 6 (2024)

  30. [30]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, et al. 2022. Chain-of-thought prompt- ing elicits reasoning in large language models. InAdvances in Neural Information Processing Systems (NeurIPS)

  31. [31]

    Justin Wolfers and Eric Zitzewitz. 2004. Prediction markets.Journal of Economic Perspectives18, 2 (2004), 107–126

  32. [32]

    Fabian Wunderlich and Daniel Memmert. 2018. The betting odds rating system: Using soccer forecasts to forecast soccer.PLOS ONE13, 6 (2018), e0198668

  33. [33]

    xAI. 2026. Grok Expert Mode. https://x.ai

  34. [34]

    Zhiheng Xi, Wenxiang Chen, Xin Guo, et al. 2023. The rise and potential of large language model based agents: A survey.arXiv preprint arXiv:2309.07864(2023)

  35. [35]

    Miao Xiong, Zhiyuan Hu, Xinyang Lu, et al . 2024. Can LLMs express their uncertainty? An empirical evaluation of confidence elicitation. InInternational Conference on Learning Representations (ICLR)

  36. [36]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, et al. 2023. ReAct: Synergizing reasoning and acting in language models. InInternational Conference on Learning Representations (ICLR)

  37. [37]

    Andy Zou, Tristan Xiao, Ryan Jia, et al. 2022. Forecasting future world events with neural networks. InAdvances in Neural Information Processing Systems (NeurIPS). A Reproducibility One command (python src/run_all.py ) rebuilds every processed table, analysis CSV, and figure from the raw transcripts and odds; the paper’s numbers are read directly from tho...

  38. [2024]

    Wisdom of the silicon crowd: LLM ensemble prediction capabilities rival human crowd accuracy.Science Advances10, 45 (2024)