A pre-registered benchmark of 13 LLM systems on all 104 World Cup 2026 matches shows fine-grained predictions expose differences that result accuracy hides.
Title resolution pending
1 Pith paper cite this work, alongside 162 external citations. Polarity classification is still indexing.
1
Pith paper citing it
162
external citations · OpenAlex
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
A pre-registered benchmark of 13 LLM systems on all 104 World Cup 2026 matches shows fine-grained predictions expose differences that result accuracy hides.