{"id":"ec7ef14d-5656-48bb-aa13-86d13637623a","arxiv_id":"2608.01395","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across 30 languages, commercial LLMs outscore all open-weight models in every EU language, and non-English service costs more and scores lower, suggesting equality requires resources beyond public web crawls.","lead":"This paper tests nine AI language models as agents playing goal-directed dialogue games in 30 languages, including all 24 official EU languages. It finds the two commercial models beat every open-weight model in every EU language, and that non-English languages cost more to run and score lower.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-run, no-seed measurements leave the universal claim 'both commercial models outscore every open-weight model in every EU language' unsupported; the tightest margin (Greek, 5.6 pts) is smaller than the paper itself warns readers to ignore.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: single-run measurements without variance estimates are being used to support the universal claim that both commercial models beat every open-weight model in every EU-24 language. The Greek margin of 5.6 points is the tightest and is explicitly within the range the paper itself says should not be read as meaningful. My independent check of Table 4 confirms this is the smallest commercial-vs-best-open margin across EU-24. The proposed concrete test—repeated seeded runs on the closest contested languages—would distinguish a robust ordering from a stochastic artifact. Since the reader's verdict is already CONDITIONAL and this concern reinforces it without overturning the broader contributions (the benchmark construction, the average performance gaps, the cost and tokenization analyses), I recommend no change to the verdict. The concern is substantive but does not invalidate the paper's central methodology or its main aggregate findings; it does mean the universal 'every language' phrasing should be treated as provisional until variance is measured.","tokens_in":38751,"tokens_out":3625,"duration_ms":33554,"concrete_test":"Rerun the Greek condition (the tightest margin, 5.6 points) and the Spanish condition (next tightest, 8.0 points) for Claude Opus 4.8 and GLM-5.2 at least 10 times each, using identical localized game files and the same game instances but different random seeds at temperature 1, following the paper's default decoding. Compute per-run clemscore and the Opus-minus-GLM margin for each language, then report the 95% confidence interval or bootstrap distribution of the margin. If the lower bound exceeds zero, the universal claim survives this test; if the interval includes zero or negative margins occur, the claim must be downgraded to 'commercial models usually/on average outperform open-weight models,' and the paper should report variance estimates for all EU-24 language-model cells.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim (Abstract; Section 5.1, Table 4) is that in every one of the 24 official EU languages, both commercial models score above every open-weight model. The margin that most strains this claim is Greek: Claude Opus 4.8 scores 75.5 and the best open-weight model (GLM-5.2) scores 69.9, a 5.6-point gap. Under the stated protocol (Section 4.3), each language is run once at temperature 1 with no fixed seed, and the Limitations explicitly say 'we report no variance estimate and small differences between adjacent cells should not be read as meaningful.' 5.6 points is precisely the kind of small difference the authors tell readers not to interpret. Per-game tables (Tables 11-15) show within-model swings of 10-30 points across languages and games, so run-to-run variance at temperature 1 could plausibly exceed 5.6 points. If the true margin is sometimes negative, the universal 'every' statement fails, even though the aggregate pattern (commercial systems leading on average) may remain. Greek is among the six manually verified languages, so localization quality is not the issue here; the issue is purely the lack of variance evidence for a universal quantifier. The cost claim (median +31% cost, -10% score) is also computed from single runs and list prices, but the universality claim is the most load-bearing because it underpins the conclusion that 'no open-weight model covers the EU-24 well.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a multilingual extension of the clembench dialogue-game evaluation framework to 30 languages (the 24 official EU languages plus six others), and evaluates nine LLMs (two commercial, seven open-weight) as self-playing agents in 14 goal-directed games. Scores are computed programmatically from rule compliance and task success, without reference answers. The main empirical claims are that (i) in every EU-24 language both commercial models outscore every open-weight model; (ii) no open-weight model covers the EU-24 well; (iii) performance correlates with web-text availability and with a newly defined Linguistic Economic Footprint for open-weight but not commercial models; and (iv) non-English languages cost more and score lower than English, with a median cost premium of 31% and a score deficit of 10%.","tokens_in":39082,"tokens_out":5269,"duration_ms":53248,"significance":"If the central findings hold, this is a valuable contribution: it provides a systematic, multi-turn, programmatically scored benchmark covering the full EU-24 language set, with cost and tokenizer analysis, and it has direct policy relevance for the EU's language-equality commitments. The paper's strengths include its transparent localisation pipeline, public code and leaderboard, and the absence of fitted parameters or reference-answer dependency. The comparison with external benchmarks is descriptive rather than circular. However, the strongest universal claim—that both commercial models beat every open-weight model in every EU language—rests on single-run measurements with no variance estimates, and the tightest margin is small enough that the authors themselves caution against interpreting such differences. The cost-premium headline is similarly derived from single runs and list prices. These issues do not undermine the overall pattern, but they do require either additional evidence or a more qualified statement.","major_comments":[{"comment":"The abstract and §5.1 state that in every EU-24 language both commercial models outscore every open-weight model. The tightest margin is Greek: Claude Opus 4.8 at 75.5 vs. GLM-5.2 at 69.9, a 5.6-point gap. The Limitations explicitly state that 'each language is run once, at temperature 1 and without a fixed seed, so we report no variance estimate and small differences between adjacent cells should not be read as meaningful.' A 5.6-point difference is exactly the kind of small difference the authors tell readers not to interpret, and the per-game tables (Tables 11–15) show swings of 10–30 points, so run-to-run variance at temperature 1 could plausibly exceed this margin. The universal quantifier is therefore not supported by the reported measurements. I ask for repeated runs or per-episode bootstrap/CI estimates for at least the tightest cells, or for rewriting the claim as the observed r","section":"§5.1, Table 4; Limitations"},{"comment":"The benchmark's portability claim depends on localised game files being equally playable in all 30 languages, but only six languages were manually verified by native speakers. The Limitations acknowledge 'a small number of game–language pairs still show near-zero completion for reasons we attribute to parsing rather than to competence.' Several unverified EU languages are exactly the ones where open-weight models collapse (Irish, Maltese, Latvian), so a localisation or parsing artifact could depress scores and materially affect the coverage conclusion. Please report which pairs are affected, how the parsing attribution was established, and whether the main EU-24 conclusions survive when those cells are removed or corrected.","section":"§3.3, §5.1, Limitations"},{"comment":"The headline cost claim ('the median non-English language costs 31% more to run than English, and scores 10% lower') is computed from single runs per language–model cell and from listed API prices. Token counts are objective, but generation at temperature 1 is stochastic, and the same variance caveat applies. Please provide a measure of run-to-run or episode-level variability for the cost ratios, or present the 31% and 10% figures as rough central tendencies rather than precise estimates.","section":"§5.3, Table 8"}],"minor_comments":[{"comment":"The definition of clemscore as a 'normalised product' of %Played and Quality should be made explicit with a formula, including how the scaling to [0,100] is applied.","section":"§4.4"},{"comment":"FineWeb-2 is referenced only by footnote URL; a full citation should be added to the reference list, as is done for other datasets.","section":"Figure 8 / Appendix H"},{"comment":"The text says tokeniser profiles are 'near-identical across providers' and then immediately reports that Claude Opus 4.8 averages roughly 50% more tokens per word than the median. Please reconcile these statements or clarify that Claude is the exception.","section":"§5.2"},{"comment":"The two-to-six point inflation of Chinese scores due to dropping Wordle is reported in both the main text and Appendix B; consider consolidating to avoid redundancy.","section":"Appendix B"},{"comment":"The caption lists sources for speaker shares and GDP only indirectly via Appendix G; state the specific data versions and access dates in the caption or immediately below the table.","section":"Table 16"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid empirical contribution and I see no circularity: the benchmark is self-authored, but the scoring is programmatic, the code is public, and the external-benchmark comparison is explicitly descriptive. The main risk is over-claiming from single-run measurements; the universal 'every EU language' claim is more precise than the protocol supports. If the authors add variance evidence or soften the claim, the paper would be a strong candidate for acceptance. The self-citation pattern is appropriate given the authors' prior work on clembench."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, short take: read it if you care about multilingual evaluation or EU digital language policy. It extends clembench's dialogue-game paradigm to all 24 official EU languages plus six others, runs nine LLMs (two commercial, seven open-weight) with programmatic scoring, and finds a clean commercial-over-open-weight gap in every EU language, plus a median 31% cost premium and 10% score penalty for non-English. That coverage alone is a real contribution.\n\nWhat it does well: the localization pipeline (translate-then-validate with two different models) makes the extension cheap, the scoring is reference-free and rule-based rather than author-judged, the LEF metric adds a useful economic dimension, and the token-fertility analysis is a nice, concrete mechanism for why non-English is costlier. Code and data appear to be public. No fitted parameters, no hand-scored subjective outcomes. Good.\n\nThe soft spots are real but localized. The headline universal—'in every EU language both commercial models outscore every open-weight model'—rests on a single run per language at temperature 1 with no fixed seed. The tightest margin is Greek: 5.6 points, and the paper's own limitations section says small adjacent differences shouldn't be read as meaningful. At temperature 1, run-to-run variance could plausibly exceed that. So the 'every' claim is not established, even if the aggregate pattern is likely right. The remedy is easy: repeated seeds or variance reporting, or temper the abstract's universal quantifier. The other soft spots are minor: only six of 30 languages were natively verified, cost uses list prices, and Chinese dropped Wordle (which the authors disclose and quantify). None of these break the central finding that commercial models are currently ahead on this interactive benchmark and that smaller languages cost more per unit of service.\n\nWho's it for: people building multilingual evaluation suites, and anyone advising on EU language technology procurement. It deserves a serious referee—send it out, but the authors should be pushed to run repeats or soften the universal claims before publication. If I'm on the committee, I'd accept after minor revision.","headline":"First systematic multi-turn EU-24 LLM benchmark; the coverage and cost analysis are genuinely useful, but the universal 'every language' claim rests on single-run measurements and should be revised or re-run.","tokens_in":39587,"tokens_out":2992,"would_cite":true,"duration_ms":29624,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Commercial LLMs outscore open-weight rivals in all 24 EU languages","keywords":["multilingual evaluation","dialogue games","EU languages","open-weight vs commercial LLMs","tokenizer cost","language equality","clemscore","low-resource languages"],"falsifier":"Rerun the nine models across the EU-24 with fixed seeds and multiple repetitions; if any commercial model ever falls below an open-weight model in any official language, or if the Greek margin (5.6 points) vanishes, the universal ordering claim is falsified. Separately, have native speakers audit all 30 localised game packs; if the near-zero-completion cells turn out to be parsing failures, the benchmark overstates some language gaps.","tokens_in":38659,"feed_emoji":"🌐","tokens_out":9854,"duration_ms":77987,"temperature":0.7,"pith_summary":"The paper tries to establish that, under an interactive multi-turn test, no open-weight large language model currently serves all 24 official EU languages adequately. It reports that two commercial systems outscore every one of the seven open-weight models in all 24 official languages, by margins from 5.6 points in Greek to 35.8 in Irish. It also claims that equal access is not equal service: pooled over models and languages, the median non-English language costs 31% more to run than English and scores 10% lower. The evaluation method makes full institutional coverage cheap, because adding a language means localising a fixed set of game files rather than writing reference answers. The reason to care is that EU law promises citizens service in their own language, and the paper's numbers test whether current AI systems can deliver that.","feed_headline":"Commercial LLMs outscore open-weight rivals in all 24 EU languages","feed_subtitle":"Interactive game tests show non-English EU languages cost 31% more and score 10% lower.","key_machinery":"The load-bearing object is a dialogue-game evaluation harness built on clembench: 14 goal-directed games (Taboo, Codenames, Wordle, Reference Game, Image Game, and others) played in self-play, with each episode scored by $\\%\\mathrm{Played}$ (share of episodes completed without aborting) and $\\mathrm{Quality}$ (task score over completed episodes), combined into $\\mathrm{clemscore} = \\%\\mathrm{Played} \\times \\mathrm{Quality}$ scaled to $[0,100]$. Because the game mechanics are language-agnostic, a new language is added by localising prompt text, response-parsing rules, feedback messages and word lists; the paper does this with a two-stage machine-localisation pipeline in which one non-evaluate","core_discovery":"Across 30 languages and 14 dialogue games, both commercial systems—GPT-5.4 and Claude Opus 4.8—score above every open-weight model in all 24 official EU languages; the margins run from 5.6 points in Greek to 35.8 in Irish. The open-weight deficit is concentrated: for all seven open models, the weakest EU language is Irish or Maltese, and the commercial systems keep 80.7% and 83.0% of their English score in their weakest EU language. Open-weight scores track public web-text volume ($\\rho=0.72$, $p<0.001$) and a language's economic footprint (up to $\\rho=0.78$); commercial models show no such correlation. Pooled, the median non-English language costs 31% more than English and scores 10% lower,","pith_inferences":["A direct stress test would be to rerun all nine models over the EU-24 with fixed seeds; the paper's single-run protocol means the 5.6-point Greek margin is the point where the universal ordering claim is most exposed.","If the open-weight gap is a training-data effect, then publicly funded non-web corpora (broadcast archives, parliamentary records, local-government text) could narrow it; the paper gestures at this policy conclusion but does not test it.","The near-zero completion cells the paper attributes to parsing, rather than competence, imply the benchmark may underestimate ability in some language–game pairs; a native-speaker audit of all 30 localised game packs would settle this.","The strong LEF correlation for open-weight models suggests market incentives alone will not serve the smallest official languages, which is the implicit case for public intervention the paper leaves to the reader."],"forward_implications":["Any EU-wide deployment today cannot get full EU-24 coverage from an evaluated open-weight model; the realistic options are commercial APIs or building language-specific resources.","Linguistic parity is achievable, not a pipe dream: the commercial models score about as well on Maltese, Estonian and Latvian as on English, so the gap reflects under-provisioning rather than intrinsic difficulty.","Public web crawls alone cannot close the gap: open-weight performance tracks crawled text volume, so parity requires resources beyond the open web, such as public broadcast archives.","Tokenisers charge a hidden price before inference: the median non-English language costs 31% more to run than English while scoring 10% lower, and low-resource EU languages use roughly twice the tokens per word.","The two commercial systems deliver markedly less value-per-dollar outside English, so a uniform per-token price does not buy equal service across languages."],"supporting_citations":[{"why":"Provides the clembench dialogue-game framework and the 14 games the benchmark is built on.","marker":"Chalamalasetti et al. (2023)"},{"why":"Argues for dialogue game-based evaluation as a third paradigm and supplies the reference-free, rule-scored design.","marker":"Schlangen et al. (2025)"},{"why":"Supplies ConceptNet 5.7 as the source of Taboo target and taboo word lists.","marker":"Speer et al. (2017)"},{"why":"Supplies Universal Dependencies v2.17 treebanks used to build Wordle target lists.","marker":"Nivre et al. (2020)"},{"why":"Establishes that tokenization makes languages cost different amounts, grounding the cost-premium claim.","marker":"Ahia et al. (2023)"},{"why":"Shows language model tokenizers introduce unfairness between languages, supporting the fertility-as-pricing mechanism.","marker":"Petrov et al. (2023)"},{"why":"Documents systematic inequalities in language technology across languages, motivating the LEF and data-availability analysis.","marker":"Blasi et al. (2022)"},{"why":"Explains tokeniser quality as a consequence of training-text availability, supporting the claim that fertility is a symptom rather than a cause.","marker":"Rust et al. (2021)"}],"fun_headline_variants":["Commercial LLMs outscore open-weight in all 24 EU languages","Commercial beats open-weight in every EU-24 language","EU-24 non-English costs 31% more, scores 10% lower","Language equality has a price: non-English EU costs 31% more","Open-weight LLMs trail commercial in all EU-24 languages"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The claim rests on treating each language's single-run clemscore as a stable measurement: the paper runs every language once, at temperature 1 with no fixed seed, and itself warns that small differences between adjacent cells should not be read as meaningful, so a 5.6-point margin may sit within run-to-run noise.","fun_headline_variants_meta":{"raw":{"variants":["Commercial LLMs outscore open-weight in all 24 EU languages","Commercial beats open-weight in every EU-24 language","EU-24 non-English costs 31% more, scores 10% lower","Language equality has a price: non-English EU costs 31% more","Open-weight LLMs trail commercial in all EU-24 languages"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001205,"raw_usage":{"total_tokens":4823,"prompt_tokens":790,"completion_tokens":4033,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":3942}},"tokens_in":534,"tokens_out":4033,"duration_ms":27146,"temperature":1.0,"reasoning_tokens":3942,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T00:12:57.434652+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the nine models across the EU-24 with fixed seeds and multiple repetitions; if any commercial model ever falls below an open-weight model in any official language, or if the Greek margin (5.6 points) vanishes, the universal ordering claim is falsified. Separately, have native speakers audit all 30 localised game packs; if the near-zero-completion cells turn out to be parsing failures, the benchmark overstates some language gaps.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Argues for dialogue game-based evaluation as a third paradigm and supplies the reference-free, rule-scored design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies ConceptNet 5.7 as the source of Taboo target and taboo word lists."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows language model tokenizers introduce unfairness between languages, supporting the fertility-as-pricing mechanism."}],"review_version":1}