{"id":"36e0bab6-ff6e-4840-a759-670b8eb84eaa","arxiv_id":"2505.12201","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM-as-a-Judge is inconsistent across languages: five models grading the same parallel ground-truth content in 25 languages agreed poorly (Fleiss' Kappa around 0.2 on average), and the proposed majority-vote ensemble only partially improves consistency.","lead":"This paper tests whether LLM judges give the same verdict when the same question and correct answer are presented in 25 different languages; they do not, with low cross-language agreement for all five models tested. The finding warns that multilingual automatic evaluation may be unreliable and suggests ensemble voting as a partial fix.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fleiss' Kappa under the all-gold design is not interpretable as 'inconsistent'; chance agreement is inflated, and raw cross-language agreement is never reported.","rationale":"The paper's central claim rests on the average Fleiss' Kappa near 0.3, interpreted as evidence that LLM judges produce inconsistent results across languages. The reader's conditional verdict focuses on external validity: gold answers may not represent real candidate distributions. However, there is a more fundamental internal problem arising from the same design choice. Because all judged answers are gold, the true rating is identical for all items, so the marginal rate of the positive category is simply the per-language accuracy, which is high. Kappa corrects for chance agreement, and with such high marginals, chance agreement is high; independent errors across languages yield Kappa near zero even when raw agreement exceeds 90%. Thus, values around 0.3 are not interpretable using conventional Kappa benchmarks. The paper never reports raw cross-language agreement, so the reader cannot determine whether the 'unreliable' conclusion follows from the data. This is not a minor concern: it targets the single headline number that supports the abstract and conclusion. The proposed test, computing PABAK or percent agreement on the existing data, is cheap and would settle whether the low Kappa is an artifact. If PABAK is high, the paper's interpretation is invalid; if PABAK is low, the original claim gains support. Because the paper as presented does not address this, the verdict should remain conditional, but the condition should be revised to require prevalence-adjusted agreement metrics or an evaluation on mixed correct/incorrect candidates. This differs from the reader's identified weakest assumption in that it is an internal measurement flaw, not merely a question of external transfer, though both stem from the all-correct design.","tokens_in":14473,"tokens_out":8520,"duration_ms":89467,"concrete_test":"On the existing data, recompute each Table 2 cell using prevalence-adjusted bias-adjusted kappa (PABAK) and the mean pairwise percent agreement across languages. If PABAK is high (e.g., > 0.8) while FK stays near 0.3, the headline value is an artifact of the all-correct design; if PABAK is also low, the reliability concern survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative evidence is the average Fleiss' Kappa of about 0.3 reported in Table 2. But every item in every dataset is a ground-truth answer (Section 3.3), so the true rating is constant ('correct' or grade 5) for all items. With this design, the marginal probability of a positive rating equals the per-language accuracy, typically 0.85-0.97. Fleiss' Kappa corrects for chance agreement; when the positive marginal rates are high and similar, chance agreement P_e is very high, so Kappa is structurally compressed. Concretely, if two 'raters' (languages) each label gold answers correctly with probability p=0.95 and make errors independently, raw pairwise agreement is p^2+(1-p)^2=0.905, but the expected Kappa is 0. Therefore values around 0.3 cannot be read directly as 'inconsistent'; they indicate only that errors across languages are moderately correlated. The paper never reports raw agreement, percent agreement, or prevalence-adjusted coefficients (e.g., PABAK), and it applies conventional Kappa benchmarks (Section 4.1) that assume variable true labels. For example, Aya-Expanse on XQuAD has per-language accuracy 96.86% and Kappa 0.30; without raw agreement, the reader cannot tell whether this reflects 6% or 60% cross-language disagreement. This is a measurement validity problem, not just an external validity limitation. It directly undermines the headline claim that 'LLMs are not yet reliable for evaluating multilingual predictions.'","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether LLM-as-a-Judge is reliable in multilingual settings by measuring cross-language consistency of judgments on parallel multilingual data. For five models, five tasks, and 25 languages, each item's ground-truth answer is judged separately in each language, and Fleiss' Kappa is computed across languages, treating each language as a rater. The authors report an average Fleiss' Kappa of about 0.3, conclude that LLMs are not yet reliable multilingual judges, analyze factors such as low-resource languages, task type, prompt design, and model scale, and propose a majority-vote ensemble of open-source models to improve consistency.","tokens_in":14738,"tokens_out":3922,"duration_ms":43309,"significance":"If the headline result were supported by an appropriate measure, the paper would be a useful, broad empirical caution for multilingual LLM-as-a-Judge and a practical contribution via the ensemble strategy. The study has real strengths: it uses parallel data across many languages and tasks, covers closed and open models, provides full prompt templates in the appendix, and involves no fitted parameters or circular derivations. However, the central evidence is currently compromised by a measurement-validity problem in the choice and interpretation of Fleiss' Kappa, so the significance of the conclusion is conditional on reanalysis.","major_comments":[{"comment":"The central consistency metric is not interpretable for the all-gold design. Because every judged item is a ground-truth answer (Section 3.3), the true rating is constant ('correct' or grade 5), and each language's marginal positive rate is the per-language accuracy. When the marginal rate is high, say p=0.95, two independently labeling languages agree with raw probability p^2+(1-p)^2≈0.905, yet the expected Fleiss' Kappa is 0. Thus values around 0.3 for Aya-Expanse on XQuAD, where accuracy is 96.86%, do not by themselves indicate large cross-language disagreement; they may indicate only that the few errors are not perfectly correlated across languages. The paper never reports raw pairwise agreement, percent agreement, or prevalence-adjusted coefficients such as PABAK, and it applies conventional Kappa benchmarks that assume variable true labels. This affects not only the headline absolute values but also the comparisons in Sections 4.2, 5.1, 5.4, and Table 5, because Kappa is compressed differently whenever per-language accuracy differs. The claim that LLMs are inconsistent across languages needs to be re-established with an appropriate agreement measure or raw agreement rates before it can support the paper's conclusions.","section":"Section 3.4 and Table 2"},{"comment":"The reported average Fleiss' Kappa of about 0.3 is not consistent with Table 2. The arithmetic mean of the 25 Yes/No Kappa values is about 0.24, the mean of the 25 Grade Kappa values is about 0.17, and the overall mean of all 50 cells is about 0.21. No model or task subset in the table produces an average near 0.3 except isolated cells. The abstract and Section 4.1 should either report the actual aggregate value or explain which subset is being averaged.","section":"Abstract and Section 4.1"},{"comment":"The Cohen's Kappa analysis between English and other languages inherits the same prevalence artifact. For low-resource languages, per-language judgment accuracy is typically lower, which changes the expected chance agreement and therefore compresses or inflates Kappa regardless of pairwise agreement. The claim that consistency is particularly poor for low-resource languages (for example, the MGSM Telugu result near 0.002 in Section 5.1) is not supported without reporting raw agreement between the English and non-English judgments or using a prevalence-adjusted coefficient.","section":"Section 5.1 and Figure 4"}],"minor_comments":[{"comment":"The metric name 'BLUE' should be 'BLEU'.","section":"Section 1"},{"comment":"The phrase 'State-ot-the-art' should be 'state-of-the-art'.","section":"Section 3.1"},{"comment":"The column header 'Acc' appears to report average grades, not accuracy; the header should be aligned with the Grade setting or the values should be described consistently in the caption.","section":"Table 4"},{"comment":"The legend contains 'Column2', which appears to be a placeholder, and the axis label is garbled; both should be cleaned up.","section":"Figure 5"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper asks a good question, but the headline number is wrong and the main statistic is the wrong tool for its own design. The average Fleiss' Kappa reported in the abstract (~0.3) is not what Table 2 shows; the actual average across the table is about 0.21. More importantly, every judged answer is a gold answer (Section 3.3), so the true label is always \"correct\" (or grade 5). With per-language accuracy mostly 0.85–0.97, chance agreement is high and Fleiss' Kappa is structurally compressed. A Kappa of 0.3 can sit on top of 90%+ raw agreement. The paper never reports raw agreement or PABAK, so \"inconsistent\" does not follow from the Kappa values as presented.\n\nWhat is genuinely new: measuring the judge's cross-language consistency on parallel gold items, rather than comparing LLM scores to human annotations, is a fresh angle. I don't know earlier work that does exactly this. The per-language Cohen's Kappa vs English (Figure 4) is the most informative part; low-resource language degradation is visible without relying on the problematic aggregate. The prompt experiments are small but reasonable, and the limitations section is honest about pointwise-only judgment and model scale. The citation pattern looks normal; Hada et al. is the relevant comparator and is cited.\n\nSoft spots, in proportion. The Kappa/raw-agreement issue is the main one and it is load-bearing. For example, Aya-Expanse on XQuAD has 96.86% accuracy and Kappa 0.2999; if errors across languages were independent, raw agreement would be around 94% and Kappa near zero. Low Kappa here means errors are somewhat correlated, not that the judge disagrees 70% of the time. Without raw agreement, the reader cannot distinguish 6% disagreement from 60%. The abstract's \"~0.3\" is also simply inconsistent with the table, whose average is closer to 0.21. The ensemble claim is overreached: it is compared only against the minimum of three open models, sometimes beats none of them (WMT23 Yes/No), and is never compared with the best single model or GPT-4o. The model-scale conclusion rests on two families, which the authors acknowledge; that is a minor limitation.\n\nWho this is for: people building multilingual leaderboards or using LLM judges in more than one language. But before trusting the conclusions, the authors need to reanalyze the data with raw agreement and a prevalence-adjusted coefficient, and reframe the claims accordingly. If raw agreement turns out high, the story changes from \"LLMs are unreliable judges\" to \"LLMs are fairly consistent on gold answers but Kappa is low due to prevalence.\" That is still worth knowing, but it is a different paper.\n\nI would send this to a serious referee. The question is important, the parallel data are reusable, and the flaw is correctable. As it stands I would not cite it, but a revised version with raw agreement reported and the ensemble claim toned down could make a solid contribution. Do not desk reject; ask for major revision.","headline":"Useful measurement buried under a statistically compromised headline: the all-gold design makes Fleiss' Kappa uninterpretable, and the abstract average doesn't match the table.","tokens_in":15267,"tokens_out":4058,"would_cite":false,"duration_ms":40168,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM-as-a-Judge is not yet reliable for multilingual predictions: the same parallel answer gets inconsistent verdicts across languages, with average Fleiss' Kappa around 0.3.","keywords":["multilingual LLM-as-a-Judge","judgment consistency","Fleiss' Kappa","cross-lingual evaluation","low-resource languages","pointwise evaluation","ensemble majority vote"],"falsifier":"Build a parallel corpus of model-generated answers of deliberately mixed quality in the same 25 languages, have the same five judges score them, and recompute Fleiss' $\\kappa$ across languages; if agreement rises well above the reported 0.3, the paper's reliability conclusion may not transfer to real judging situations.","tokens_in":14250,"feed_emoji":"⚖️","tokens_out":8324,"duration_ms":73585,"temperature":0.7,"pith_summary":"The paper tests a foundational assumption behind multilingual LLM-as-a-Judge: that a good judge gives the same verdict for the same content regardless of the language it is expressed in. The authors build parallel test sets in 25 languages across five tasks, run five LLM judges, and treat each language as one rater to compute Fleiss' $\\kappa$. They report an average $\\kappa$ near 0.3, with some model-task pairs far lower, meaning the same judge frequently disagrees with itself about semantically identical answers. Consistency is worst for low-resource languages, and neither multilingual pretraining nor larger model size fixes it. If this holds, automatic evaluation results cannot be treated as language-neutral, and multilingual benchmarks need a separate reliability check for their judges.","feed_headline":"Same LLM judge gives different verdicts across 25 languages","feed_subtitle":"Average Fleiss' Kappa near 0.3 shows multilingual LLM-as-a-Judge is not yet trustworthy.","key_machinery":"The load-bearing device is parallel multilingual data combined with Fleiss' Kappa computed across languages: each language's judgment is treated as a rater, so that agreement measures whether the verdict depends on content or on language. Since the same underlying item appears in every language, language is the only variable. The judge prompt is held fixed in English with a placeholder naming the evaluation language, following prior practice; pointwise evaluation (one candidate judged against a rubric) is used rather than pairwise comparison because parallel incorrect candidates are hard to obtain. The result is a measurement scheme that isolates cross-lingual judgment reliability from task accuracy.","core_discovery":"The paper's central discovery is that LLM-as-a-Judge is not language-invariant: when the same question-answer pair is translated into 25 languages and presented in a parallel format, a single judge model frequently disagrees with itself. Treating each language's verdict as one rater, the authors compute Fleiss' $\\kappa$ and find average agreement around 0.3 across five models and five tasks, with some model-task combinations falling below 0.1. Good task accuracy does not rescue consistency: a judge can be mostly correct in every language while still disagreeing across languages, and Spearman correlations between accuracy and Kappa vary in sign across tasks. Low-resource languages show particularly low agreement with English judgments, and neither multilingually trained models such as Aya-Expanse nor larger parameter counts reliably close the gap. The paper also reports that an ensemble majority vote of three open-source judges improves consistency relative to the worst single judge in most settings, and that binary Yes/No judgments are more consistent than 1-5 grade judgments.","pith_inferences":["Pith inference: because the prompt is written in English with a target-language placeholder, part of the measured inconsistency may come from the prompt language rather than the judged content; a direct test would rerun the experiment with fully native prompts per language.","Pith inference: the pointwise-only design leaves pairwise comparative judging untested; constructing parallel incorrect-candidate pairs and measuring Kappa there would reveal whether the instability is specific to absolute grading.","Pith inference: the near-zero or negative correlation between accuracy and Kappa suggests a judge could be uniformly wrong yet 'consistent' in appearance, so practical reliability checks should also calibrate against human labels, not just cross-language agreement.","Pith inference: the task-dependence of consistency suggests a certification scheme in which a judge is validated per task-language cell before its outputs are used, rather than trusted on the strength of its average multilingual competence."],"forward_implications":["Multilingual evaluation results obtained by a single LLM judge in one language cannot be assumed to hold in another; reported numbers should be accompanied by a cross-language consistency check.","Low-resource languages are the weakest link, so evaluations reported for languages like Telugu or Swahili need the most scrutiny.","Since larger models and multilingual pretraining do not systematically improve consistency, simply upgrading the judge is unlikely to make multilingual evaluation trustworthy.","The ensemble majority vote of open-source judges is a practical remedy that usually beats the worst single judge, and binary judgments with explanatory prompts are more consistent than fine-grained grades.","A judge that scores well on accuracy in each language can still be inconsistent across languages, so accuracy and consistency should be reported as separate axes of judge quality."],"supporting_citations":[{"why":"Defines LLM-as-a-Judge, the evaluation paradigm whose cross-lingual reliability this paper tests.","marker":"Zheng et al. (2023)"},{"why":"Provides the XQuAD parallel question-answering dataset used as one of the five evaluation tasks.","marker":"Artetxe et al. (2020)"},{"why":"Provides the MGSM parallel multilingual math dataset, the setting with the lowest judged consistency.","marker":"Shi et al. (2023)"},{"why":"Provides the WMT23 parallel machine-translation test instances scored by the judges.","marker":"Kocmi et al. (2023)"},{"why":"Provides the WikiLingua parallel summarization dataset, the task with the highest observed Kappa.","marker":"Ladhak et al. (2020)"},{"why":"Provides the XDailyDialog parallel dialogue corpus used to test dialogue-generation judging.","marker":"Liu et al. (2023)"},{"why":"Introduces Aya-Expanse, the multilingual-specific model whose failure to improve consistency supports the claim that multilingual training is not sufficient.","marker":"Dang et al. (2024)"},{"why":"Motivates the English prompt with a target-language placeholder, the prompt design used throughout the experiments.","marker":"Ahuja et al. (2023)"}],"fun_headline_variants":["LLM judges disagree with themselves across languages","Multilingual LLM judges: only 0.3 agreement on average","LLM-as-a-Judge fails consistency test in 25 languages","Ensemble voting improves multilingual LLM judge consistency","Multilingual training doesn't fix LLM judge inconsistency"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that judging perfectly correct, parallel reference answers behaves like judging the mixed-quality, sometimes wrong candidate outputs that LLM-as-a-Judge meets in real use, and that judging one candidate at a time captures the reliability of judging pairs of candidates as well.","fun_headline_variants_meta":{"raw":{"variants":["LLM judges disagree with themselves across languages","Multilingual LLM judges: only 0.3 agreement on average","LLM-as-a-Judge fails consistency test in 25 languages","Ensemble voting improves multilingual LLM judge consistency","Multilingual training doesn't fix LLM judge inconsistency"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000722,"raw_usage":{"total_tokens":3237,"prompt_tokens":937,"completion_tokens":2300,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":2220}},"tokens_in":553,"tokens_out":2300,"duration_ms":15600,"temperature":1.0,"reasoning_tokens":2220,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:38:24.474761+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a parallel corpus of model-generated answers of deliberately mixed quality in the same 25 languages, have the same five judges score them, and recompute Fleiss' $\\kappa$ across languages; if agreement rises well above the reported 0.3, the paper's reliability conclusion may not transfer to real judging situations.","supporting_citations":[],"review_version":1}