{"id":"387b55e1-d653-4e6f-81f2-12ee9eddf7df","arxiv_id":"2607.18084","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A pre-registered benchmark of 13 LLM systems on all 104 World Cup 2026 matches shows fine-grained predictions expose differences that result accuracy hides.","lead":"WorldCupArena asks language models and deep-research agents to predict football matches before kickoff — winner, exact score, lineups, events, statistics, and the tournament champion. Across 104 matches of the 2026 World Cup, models with similar win-accuracy separate on the finer details, and web search did not reliably improve forecasts.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Temporal leakage is the load-bearing concern: A.5 admits search-source publication dates are imperfect, so the pre-match claim could be an artifact of contaminated evidence.","rationale":"The reader's verdict (CONDITIONAL) already identifies the temporal protocol as the weakest assumption, and I agree. The central claim is that fine-grained evaluation reveals differences that result accuracy hides. For that claim to be evidence of model capability, every prediction must be genuinely pre-outcome. The paper's own A.5 admits that search-source publication dates are imperfect and leakage checks require manual review. Given 104 matches and 13 systems, a single undetected post-deadline source could inflate a model's detailed predictions, producing the exact phenomenon claimed. This is not a minor methodological detail; it is the benchmark's core validity. The proposed audit of a sample of saved URLs with independent timestamp verification is a direct, feasible test. If no post-deadline sources are found, the concern is mitigated; if they are found, the headline claim likely requires reanalysis or rejection. I did not elevate the other statistical concerns (absence of error bars, post-hoc calibration) to the same level because they affect the strength of the evidence, not the validity of the benchmark. Thus the verdict should remain CONDITIONAL, pending this audit.","tokens_in":16100,"tokens_out":4751,"duration_ms":340798,"concrete_test":"Take a random sample of 20 matches from the 104. For each S2 prediction in those matches, retrieve the saved list of URLs and timestamps from the artifact repository. Independently verify the true publication time of each URL via archive.org snapshots, site metadata, or server 'Last-Modified' headers. Flag any source whose verified publication time is after the 24-hour-before-kickoff lock. Recompute the headline metrics (Result accuracy, Scoreline, T2–T4, and Composite) after excluding every prediction that used at least one flagged source. If the pattern of 'similar result accuracy but clearly different detailed predictions' survives this stricter exclusion, the leakage concern is substantially mitigated; if it disappears or changes ranking leaders, the central claim is compromised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that fine-grained predictions separate models that appear similar on result accuracy—requires that every prediction was actually made before kickoff, with no model seeing outcome-revealing information. The paper's only safeguard is the automatic leakage check described in §3.1, which excludes predictions containing information published after the deadline. But Appendix A.5 concedes: 'Search-source publication times are also imperfect, so automatic leakage checks need periodic manual review.' For 13 systems over 104 matches, manual review of all search sources is not documented as systematic; one missed post-deadline source in an S2 system could inflate its Scoreline and player/event scores, creating the exact kind of 'clearer difference' the headline claims. The concern is not that leakage is proven, but that the benchmark's central validity guarantee rests on an unverified assumption. If differential leakage occurred, the comparison between S1 (given evidence) and S2 (self-search) would be biased, and the claim that detailed predictions reveal capability would be undermined.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"WorldCupArena introduces a dynamic, pre-event benchmark for evaluating language models and deep-research agents on football forecasting. For all 104 matches of the 2026 FIFA World Cup, models either receive a common evidence package (S1) or search for evidence themselves (S2); they predict result, score, lineups, events, statistics, and the full competition, with predictions locked 24 hours before kickoff. The paper reports result accuracy, exact-score accuracy, a Scoreline metric that rewards close misses, and layer-level scores, and compares 13 systems against betting-market and human-fan baselines. The central claim is that models with similar result accuracy differ more clearly on fine-grained predictions, and that the benchmark reveals capabilities not captured by coarse result accuracy. The paper also reports an in-play forecasting track and an open-source artifact release.","tokens_in":16383,"tokens_out":4156,"duration_ms":51962,"significance":"If the central claim is sustained, the benchmark is a useful and reusable contribution: it evaluates genuinely pre-event forecasting, distinguishes result accuracy from finer-grained predictive skill, and provides a protocol that can be applied to future leagues. The paper has substantial strengths: saved predictions and evidence, availability-aware aggregation that distinguishes missing data from zero outcomes, explicit leakage checks as part of the pipeline, a published scoring/evaluation codebase, and an honest discussion of implementation boundaries. These are real contributions to evaluation methodology for LLM forecasting. However, the headline quantitative claims are not yet fully supported: the main comparisons lack confidence intervals and raw Scoreline values, the display calibration is applied after observing the score distribution, and the temporal-leakage safeguard is acknowledged to be imperfect without a documented systematic audit. These issues are fixable and do not, in my view, invalidate the benchmark design.","major_comments":[{"comment":"The benchmark's central validity guarantee is that all predictions were made before kickoff and contain no outcome-revealing information. The paper states in §3.1 that a prediction is excluded if it contains information published after the deadline, but §A.5 concedes that 'search-source publication times are also imperfect, so automatic leakage checks need periodic manual review.' The manuscript does not document any systematic manual audit of the 104-match, 13-system corpus. For S2 systems, a single missed post-deadline source could inflate Scoreline, player, or event scores and create the kind of fine-grained separation the headline claims. I am not asserting that leakage occurred; I am saying the central claim rests on an unverified assumption. Please provide a leak-audit trail: number of runs flagged and excluded per system, the manual-review protocol actually followed, and a sensiti","section":"§3.1 and §A.5"},{"comment":"The headline comparisons are reported without confidence intervals, raw Scoreline values, or paired tests. For example, the 15.14 display-point Scoreline gap between Claude Opus 4.7 (Thinking + Search) and BetVictor compares a 58-match system with a 104-match baseline; the result-accuracy gap is only 2.4 points. Similarly, several systems in Table 2 have result accuracies of 65–68% and Scoreline values spanning 36.95–68.49, but no uncertainty is attached to these numbers. The paper could be over-reading noise. The authors state that comparisons are 'descriptive rather than paired,' but the phrase 'models with similar result accuracy differ more clearly on detailed predictions' is a comparative claim that needs statistical support. Please report raw Scoreline means and per-system standard errors or confidence intervals, and perform paired shared-match analyses for the S1/S2 contrasts and","section":"§5.2, Table 2"},{"comment":"The display calibration in Eq (3) is chosen after observing the distribution of raw scores and is applied to the Scoreline values in Table 2. The example in the text shows that raw scores 44 and 47 become displayed scores 23.15 and 35.43: a 3-point raw gap becomes a 12.28-point displayed gap. This nonlinear expansion can make a small raw difference look like a 'clearer gain' in Scoreline. The transformation is fixed, but the center and temperature are free parameters selected on the basis of the observed result cluster. Because the central 'clearer gain' claim is expressed in display-calibrated Scoreline values, please report the raw Scoreline numbers throughout the leaderboard and provide a sensitivity analysis showing that the main ordering and the headline gaps are stable under reasonable alternative calibration choices.","section":"§4.2, Eq (3)"}],"minor_comments":[{"comment":"The variables e_d, e_t, and e_team are used in Eq (1) but not formally defined. Please define them explicitly (e.g., absolute error in goal difference, total goals, and team-wise goals) in one sentence before or after the equation.","section":"§4.1, Eq (1)"},{"comment":"It is easy to misread Table 2's Scoreline column as a raw mean rather than a display-calibrated value. Please add a note under the table or in the column header stating that Scoreline is reported after the display calibration in Eq (3), and that raw values are available in the artifact.","section":"§4.2 and Table 2"},{"comment":"The comparison with the human-fan baseline uses 94 matches for fans and 58 for the best Claude system, and the Polymarket baseline has no Scoreline value in 39 fixtures because of 'Any Other Score.' Please make these coverage differences more prominent, since they affect the interpretation of the baseline gaps.","section":"§5.2, RQ5"},{"comment":"The shared-match analysis is reported only in the text as three numbers. Consider a small table or figure showing the shared-match comparison for the three models, including the number of shared matches, so readers can verify the direction and magnitude of the search effect.","section":"§5.2, RQ3"},{"comment":"The sentence 'Champion, runner-up, third place, and top scorers are derived from the saved match list. rather than accepted as unconstrained independent claims.' contains a misplaced period before 'rather' and reads awkwardly. Please rephrase.","section":"§A.2"},{"comment":"The Polymarket baseline retrieves prices in the window [T−24h, T), while other baselines use the latest stored pre-match odds. Please clarify whether this timing difference matters for the reported result-accuracy comparison, especially for matches with late news.","section":"§5.1"}],"recommendation":"major_revision","confidential_remarks":"The central idea is sound and the artifact seems genuinely open and reproducible, but the manuscript currently overstates the strength of its headline comparisons. The temporal-leakage issue and the lack of statistical quantification are both addressable within the paper's scope: a manual audit summary, raw Scoreline values, confidence intervals, and sensitivity analyses. I would not reject, but I would ask for a revised version before accepting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear X,\n\nWorldCupArena is worth a look if you care about evaluating LLMs and agents on genuinely pre-event forecasting. The core deliverable is solid: a pre-registered, multi-layer benchmark (result/score, lineups, events, stats, competition path) with predictions locked 24 hours before kickoff, saved evidence, and open code, prompts, predictions, and evaluation scripts. That combination is new—existing sports benchmarks are retrospective and forecasting benchmarks mostly ask for coarse single answers. The authors handle missing truth sensibly (availability-aware aggregation), report coverage counts, and include betting-market and human-fan baselines. The observation that different models lead different layers (Claude on T1/T2, its search version on T3, Gemini Deep Research on T4) is a useful counterexample to relying on result accuracy alone.\n\nThe soft spots are real but addressable. First, the headline claim—'models with similar result accuracy differ more clearly on detailed predictions'—is overstated. The spreads on T2/T3/T4 are a few points, and no confidence intervals or significance tests are reported. Second, the Scoreline display values in Table 2 are post-hoc S-curve amplified (Eq. 3), and the raw means are not shown; the 'clearer gain' over baselines is in display points, and the calibration was chosen after seeing the score distribution. The monotonic transform preserves ranking, so this is not fatal, but it makes the central comparison look more dramatic than the raw numbers support.\n\nThe temporal leakage concern deserves referee scrutiny. The protocol saves URLs and timestamps and excludes leaked predictions, but Appendix A.5 concedes search-source publication times are imperfect and manual review is needed 'periodically.' For 104 matches and 13 systems, no systematic manual audit is documented. That weakens the S2 (self-search) results specifically—differential leakage could bias the S1-vs-S2 comparison. I don't think it sinks the benchmark: the finding that S1 systems with similar result accuracy differ on detailed metrics doesn't depend on search, and the paper already frames its search comparison as descriptive. But the reliability of the S2 numbers is an open question.\n\nOverall: a solid, honest benchmark paper with a few post-hoc choices and one unverified assumption. The fixes are straightforward—report raw Scoreline, add CIs or bootstrap, and either document the manual leakage audit or soften the S2 claims. I'd send it to serious review; with revisions it would be a useful reference for anyone building pre-event evaluations.","headline":"A genuinely reusable pre-registered benchmark for fine-grained football forecasting, with a headline claim that is a bit ahead of the evidence.","tokens_in":16828,"tokens_out":4355,"would_cite":true,"duration_ms":41680,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"WorldCupArena sets out to show that evaluating football forecasts needs fine-grained, pre-match predictions: across 104 World Cup matches, models with nearly identical result accuracy separate sharply on scorelines, lineups, events, and sta","keywords":["football forecasting","large language models","deep research agents","dynamic benchmark","pre-registered prediction","scoreline metric","World Cup 2026","fine-grained evaluation"],"falsifier":"Audit the publicly saved source URLs and snapshots for the search-enabled systems: if any page retrieved before the lock already contains the final score or a post-match report, the temporal guarantee is broken; a simple diff against the official match record would reveal it.","tokens_in":1199,"feed_emoji":"⚽","tokens_out":1206,"duration_ms":78930,"temperature":0.7,"pith_summary":"The paper tries to establish that a model's football forecasting skill cannot be read off result accuracy alone. To show this, it builds WorldCupArena, a benchmark that freezes each model's forecast—result, score, lineups, events, statistics, and full competition path—24 hours before kickoff and scores it only after the official record exists. Across all 104 matches of the 2026 World Cup and 13 systems, result accuracy differed little between many systems, while the finer layers separated them clearly; the top system's clearest edge over betting-market and human-fan baselines was a much higher scoreline score, meaning its misses were closer rather than its hits more frequent. The same locked-prediction protocol is reusable for future leagues and cups, so models released later can be tested on genuinely unknown outcomes. A sympathetic reader would care because it supplies a method—and evidence—for measuring whether AI forecasts are useful before an event, not just accurate after the fact.","feed_headline":"Win-rate alone hides which AI forecasters are better","feed_subtitle":"Across 104 World Cup matches, near-miss scoreline quality separates systems that look equal on win rate.","key_machinery":"The load-bearing mechanism is the prediction lock plus the five-layer scoring taxonomy. Every forecast is frozen 24 hours before kickoff together with the evidence snapshot and, for search-enabled agents, the retrieved source URLs; scoring happens only after the official record exists. The five layers—result and score, players and lineups, events, tactics and statistics, and competition outcome—are combined with fixed weights, and missing truth fields are excluded rather than scored as zeros. The scoreline score does the differentiating work: it awards 100 for an exact score and otherwise gives partial credit for a correct result class and small errors in goal difference, total goals, and te","core_discovery":"On the paper's own terms, the discovery is that fine-grained pre-match predictions carry signal that coarse result accuracy hides. Over all 104 matches and 13 systems, result accuracy fell in a narrow band while detailed scores varied widely; the leading system's clearest advantage over betting-market and human-fan baselines was not more strict hits but a much higher scoreline score, meaning its wrong scoreline predictions were consistently closer. At the competition level, four systems predicted champion Spain, and only two also produced the exact Spain–Argentina final pairing.","pith_inferences":["Inference: If locked-forecast scoring generalizes, the same design could expose hidden differences in other domains where a coarse category is the usual metric—election outcomes, economic releases, or injury reports—where a near miss is informative.","Inference: The finding that search does not consistently help is a statement about current commercial systems, not a law; a controlled comparison holding the base model fixed while varying only the search tool would test whether better search evidence ever pays off.","Inference: The consensus failures on the same upsets suggest that model diversity, not just average quality, matters for forecasting; ensembling diverse models might avoid the misses that a single strong system cannot.","Inference: Because the scoreline metric deliberately rewards closeness, it could be used as a training or selection signal—optimizing for scoreline quality may yield forecasts more useful to bettors or planners, who care about the size of the miss, not just whether the favorite won."],"forward_implications":["Result accuracy alone is an incomplete measure of forecasting skill: systems that pick the same number of winners can be far apart on scorelines, lineups, events, and statistics.","Exact-score prediction remains very hard for all systems (roughly 10 to 17 percent), and near-miss credit is what separates the leading systems from betting-market and fan baselines.","Adding web search did not consistently improve match forecasting in the systems compared; observed changes on shared matches ranged from slightly negative to moderately negative.","Competition-level evaluation adds signal: four systems predicted the champion, but only two also predicted the exact final pairing, distinguishing a correct champion reached through the right bracket from one reached via the wrong path.","The four-step protocol—save evidence and prediction before kickoff, score after the official record—is reusable for future leagues and cups and for models released after the 2026 World Cup."],"fun_headline_variants":["Scoreline quality, not win rate, reveals AI's forecasting edge","Detailed predictions tell which football AI is better","WorldCupArena: fine-grained forecasts expose AI skill gaps","Closer wrong scores separate AI forecasters","Beyond win rate: scoreline score differentiates AI systems"],"cache_read_input_tokens":18304,"weakest_assumption_plain":"The entire pre-match validity rests on the assumption that no system saw the final outcome before its prediction was locked, which depends on imperfect publication-date checks and manual review of search sources.","fun_headline_variants_meta":{"raw":{"variants":["Scoreline quality, not win rate, reveals AI's forecasting edge","Detailed predictions tell which football AI is better","WorldCupArena: fine-grained forecasts expose AI skill gaps","Closer wrong scores separate AI forecasters","Beyond win rate: scoreline score differentiates AI systems"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000173,"raw_usage":{"total_tokens":1118,"prompt_tokens":746,"completion_tokens":372,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":490,"completion_tokens_details":{"reasoning_tokens":293}},"tokens_in":490,"tokens_out":372,"duration_ms":4739,"temperature":1.0,"reasoning_tokens":293,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T16:01:46.178692+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Audit the publicly saved source URLs and snapshots for the search-enabled systems: if any page retrieved before the lock already contains the final score or a post-match report, the temporal guarantee is broken; a simple diff against the official match record would reveal it.","supporting_citations":[],"review_version":1}