REVIEW 3 major objections 4 minor 1 cited by
The Roles of English in Evaluating Multilingual Language Models
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Using English as an interface in mixed-language prompts evaluates more than the intended task and is therefore imprecise for measuring multilingual understanding, the paper argues.
desk verdict A useful two-role framing of English in multilingual eval, but the blanket 'inherently imprecise' conclusion overreaches for task-performance goals. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the distinction between two roles of English in multilingual evaluation, and the named "mixed-prompt" as the setup that conflates them. A mixed-prompt is an English instruction or label frame interleaved with target-language content, unlike natural code-switching. The argument works by showing that this setup loads multiple uncontrolled factors onto a single score: label token representations, instruction-following in English, script switching, and unnatural language switching, all in addition to the target-language task. The paper also uses the programming-language analogy to show that English-as-interface would only be clean if prompt meaning were language-agnostic, which it is not.
What would settle it
A controlled benchmark comparing mixed-prompts with fully native target-language prompts, matched for formatting and label representation, that finds no systematic score gap would undercut the claim that mixed-prompts add extraneous factors, because the interface could then be considered as clean as native prompts.
Extended reading notes
Core claim
The paper's central claim is that using English as an interface in mixed-prompts evaluates more than just the task or multilingual understanding, and is therefore imprecise or misleading. The authors distinguish English as interface, whose goal is task performance, from English as natural language, whose goal is language understanding. A mixed-prompt such as MaLa-500's "The topic of the news {sentence} is {topic}" embeds a target-language sentence inside an English frame; with English words as labels, the model is simultaneously evaluated on English instruction following, script switching for non-Latin scripts, the unnaturalness of the switch, and the task itself. Because English is a natural language, not a programming language, these extra factors cannot be separated from task performance. The conclusion follows: "This all results in imprecise or misleading evaluations, even if the ultimate goal was to evaluate and improve task performance."
Load-bearing premise
The argument assumes that language understanding, not task performance, should be the primary goal of multilingual LM evaluation; if task performance alone is a legitimate goal, mixed-prompts remain a defensible method.
Editorial extensions
If this is right
- Model rankings from benchmarks that embed target-language sentences in English frames and use English label words should not be read as rankings of multilingual understanding.
- Translation evaluations that prompt in English when English is neither source nor target measure prompt-following and interface compatibility in addition to translation quality.
- Adopting native-language or natural code-switched prompts will require more instruction-tuning data in target languages, shifting the bottleneck from prompt design to data and modeling.
- Reporting "multilingual understanding" from mixed-prompt results should be replaced with a description that the model understands English instructions interleaved with target-language passages.
Reading between the lines
- A testable extension the paper does not run is to quantify the confound: compare mixed-prompt and fully native-prompt scores on the same tasks and regress the gap on an English instruction-following benchmark; the paper's position predicts a systematic, not incidental, gap.
- The argument generalizes beyond English: any high-resource language used as a bridge for low-resource languages would carry the same interface/natural-language ambiguity, so the critique applies to evaluation schemes that translate into a pivot language.
- If mixed-prompts are imprecise at evaluation time, the same imprecision likely transfers to training regimes that fine-tune on mixed-prompt data, implying the recommendation to move away from English-as-interface has consequences for data creation, not just for evaluation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that English used in multilingual language-model evaluation can play two distinct roles: as an interface (a prompt format chosen to maximize task performance) and as a natural language (a target of linguistic understanding). The authors contend that mixing an English instruction frame with a target-language passage, which they call a mixed-prompt, is unnatural rather than genuine code-switching, and that it evaluates more than the intended task or than multilingual language understanding (MLU). They illustrate this with examples from SIB-200, IrokoBench, and a machine-translation prompt from Hendy et al. They conclude that mixed-prompt evaluations are imprecise or misleading, even when the stated goal is task performance, and recommend moving toward native target-language prompts or natural code-switched prompts.
Significance. If the conceptual distinction were fully sustained, the paper would make a useful contribution by giving researchers a vocabulary for separating interface-level task performance from language understanding in multilingual benchmarks. Its strengths are the clear two-role taxonomy, the concrete worked examples, the explicit acknowledgment of the task-performance perspective, and the transparency about nonstandard terminology. The paper makes no empirical claims and ships no code, which is appropriate for a position paper. The central weakness is that the argument's strongest conclusion—that mixed prompts are imprecise even for task performance—is not supported by the definitions provided; the analysis establishes imprecision only relative to MLU, not relative to a task-performance construct. If that overreach is corrected, the paper can serve as a helpful caution against conflating benchmark scores with language understanding.
major comments (3)
- [§3, first paragraph; Conclusion] The paper's strongest conclusion, that mixed prompts yield 'imprecise or misleading evaluations, even if the ultimate goal was to evaluate and improve task performance,' is not actually established by the preceding argument. In §3, the imprecision is defined relative to multilingual language understanding ('we arguably do not test MLU'), while the task-performance perspective in §2 is acknowledged as practical but never given a target definition. If the evaluation construct is task performance in a system that is deployed with an English instruction template, a mixed-prompt evaluation measures exactly the behavior of interest; the English frame is the operational environment, not an extra variable. The authors need to either reject the task-performance perspective on substantive grounds or qualify the conclusion so that 'imprecise or misleading' applies only to evaluations that claim to measure MLU. As written, the recommendation to abandon mixed prompts overreaches its premises.
- [§3, AfriMMLU example] The bullet list stating that the IrokoBench prompt tests 'code-switching, script-switching, instruction following in English, grammatical error correction in English' in addition to the task rests on an implicit decomposition of the task in which the English instruction frame is outside the task. For a benchmark whose input is exactly the English template plus a target-language question, processing that template is part of the input–output mapping. The paper should define 'the task' at the level of abstraction it assumes; otherwise the phrase 'more than just the task' is a terminological claim about where the task boundary is drawn, not an argument. This matters because the recommendation to move away from mixed prompts depends on that boundary.
- [§3, representation paragraph] The sentence 'Without target language words or (to an extent) language-agnostic labels, the evaluation method and goal will be inherently imprecise' is too strong without qualification. In a cross-lingual evaluation that fixes the same English label set for all languages, the label representation is constant across conditions and is not a source of cross-lingual imprecision; it may be a confound for measuring MLU, but it is a deliberate part of the interface for a task-performance evaluation. The claim should be restricted to the MLU setting or to comparisons where label language is itself a variable.
minor comments (4)
- [Figure 1] Figure 1 is informative but is never referenced in the body text; please add a pointer, for example at the end of §2 where the diagram is first discussed.
- [Footnote 2] The nonstandard definition of MLU is important enough to appear in the main text; as written, a reader who skips footnotes may misunderstand the scope of the argument.
- [Footnote 6] The GitHub URL in footnote 6 points to a specific commit; for archival purposes, consider citing the repository version or a DOI rather than an ephemeral blob URL.
- [Throughout] The terms 'imprecise,' 'task,' and 'MLU' are used as technical terms without explicit definitions in one place; a short 'terminology' paragraph in §2 would improve readability and would also help resolve the major concerns above.
Circularity Check
No significant circularity: the paper is an argumentative position piece with no fitted parameters, predictions, or self-citation chain that reduces its conclusion to its inputs.
full rationale
This is a position paper, not an empirical derivation. It introduces two conceptual roles for English in multilingual LM evaluation (as an interface vs. as a natural language) and argues that mixed prompts evaluate extra factors beyond the intended task or beyond multilingual language understanding. The supporting examples (MaLa-500, AfriMMLU, and the Hendy et al. translation prompts) are external evaluation setups, not quantities derived from the paper's own assumptions, so there is no fitted-parameter or prediction-style circularity. The paper's central conclusion, that mixed prompts are imprecise or misleading, is indeed conditional on treating language understanding as the primary evaluation goal, and the paper grants but does not refute the task-performance perspective; however, that is an argumentative gap or a normative priority issue, not a circular reduction of the conclusion to its premises. The authors' self-citation to Ploeger et al. (2024) is only a passing reference to prior methodological discussion and is not load-bearing. No equation, definition, or fitted input is reused as the claimed result, so no specific circular step can be exhibited. The honest finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Multilingual evaluation should aim at language understanding (MLU), not at task performance as an end in itself.
- ad hoc to paper A mixed-prompt (English instruction frame plus target-language passage) is not natural language use and therefore does not test MLU.
- domain assumption Prompting is so sensitive to wording that English cannot serve as a language-agnostic interface.
- domain assumption The extra factors in mixed-prompts (script-switching, English instruction following, English grammar errors, code-switching) are outside the intended construct of task performance or MLU.
Cite this review
Pith. "Pith review of The Roles of English in Evaluating Multilingual Language Models." pith.science (2026). https://pith.science/paper/KY2BPEEU
@misc{pith2026241208392,
author = {Pith},
title = {Pith review of: The Roles of English in Evaluating Multilingual Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/KY2BPEEU}},
note = {Machine review of arXiv:2412.08392}
}
read the original abstract
Multilingual natural language processing is getting increased attention, with numerous models, benchmarks, and methods being released for many languages. English is often used in multilingual evaluation to prompt language models (LMs), mainly to overcome the lack of instruction tuning data in other languages. In this position paper, we lay out two roles of English in multilingual LM evaluations: as an interface and as a natural language. We argue that these roles have different goals: task performance versus language understanding. This discrepancy is highlighted with examples from datasets and evaluation setups. Numerous works explicitly use English as an interface to boost task performance. We recommend to move away from this imprecise method and instead focus on furthering language understanding.
Forward citations
Cited by 1 Pith paper
-
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
A meta-review of about 110 critical studies finds nine systemic weaknesses in AI benchmarking and concludes that benchmarks are receiving disproportionate trust in AI governance.
Reference graph
Works this paper leans on
-
[1]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
David Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev, Jesujoba Alabi, Yanke Mao, Haonan Gao, and En-Shiun Lee. 2024 a . https://aclanthology.org/2024.eacl-long.14 SIB-200 : A Simple , Inclusive , and Big Evaluation Dataset for Topic Classification in 200+ Languages and Dialects . In Proceedings of the 18th Conference of the European Chapter of the Ass...
work page 2024
-
[4]
David Ifeoluwa Adelani, Jessica Ojo, Israel Abebe Azime, Jian Yun Zhuang, Jesujoba O. Alabi, Xuanli He, Millicent Ochieng, Sara Hooker, Andiswa Bukula, En-Shiun Annie Lee, Chiamaka Chukwuneke, Happy Buzaaba, Blessing Sibanda, Godson Kalipe, Jonathan Mukiibi, Salomon Kabongo, Foutse Yuehgoh, Mmasibidi Setaka, Lolwethu Ndolela, Nkiruka Odu, Rooweither Mabuy...
arXiv 2024
-
[5]
Mikel Artetxe, Sebastian Ruder, Dani Yogatama, Gorka Labaka, and Eneko Agirre. 2020. https://aclanthology.org/2020.acl-main.658 A Call for More Rigor in Unsupervised Cross-lingual Learning . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 7375--7388
work page 2020
-
[6]
Akari Asai, Sneha Kudugunta, Xinyan Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, and Hannaneh Hajishirzi. 2024. https://aclanthology.org/2024.naacl-long.100 BUFFET : Benchmarking Large Language Models for Few-shot Cross-lingual Transfer . In Proceedings of the 2024 Conference of the North American Chapter of the Association...
work page 2024
-
[7]
Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Jeffrey Hsu, Mimansa Jaiswal, Wilson Y. Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick,...
arXiv 2024
-
[8]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss , Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott G...
work page 2020
Show all 29 references
-
[9]
Julen Etxaniz, Gorka Azkune, Aitor Soroa, Oier Lacalle, and Mikel Artetxe. 2024. https://aclanthology.org/2024.naacl-short.46 Do Multilingual Language Models Think Better in English ? In Proceedings of the 2024 Conference of the North American Chapter of the Association for Co...
2024
-
[10]
Jinlan Fu, See-Kiong Ng, and Pengfei Liu. 2022. https://aclanthology.org/2022.emnlp-main.674 Polyglot Prompt : Multilingual Multitask Prompt Training . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages 9919--9935
2022
-
[11]
Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023. http://arxiv.org/abs/2302.09210v1 How Good Are GPT Models at Machine Translation ? A Comprehensive Evaluation . arXiv prepr...
2023 arXiv
-
[12]
Lianzhe Huang, Shuming Ma, Dongdong Zhang, Furu Wei, and Houfeng Wang. 2022. https://aclanthology.org/2022.emnlp-main.790 Zero-shot Cross-lingual Transfer of Prompt-based Tuning with a Unified Multilingual Prompt . In Proceedings of the 2022 Conference on Empirical Methods in ...
2022
-
[13]
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. https://aclanthology.org/2020.acl-main.560 The State and Fate of Linguistic Diversity and Inclusion in the NLP World . In Proceedings of the 58th Annual Meeting of the Association for Compu...
2020
-
[14]
o pf, Yannic Kilcher, Dimitri von R \
Andreas K \"o pf, Yannic Kilcher, Dimitri von R \"u tte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Minh Nguyen, Oliver Stanley, Rich \'a rd Nagyfi, Shahul Es, Sameer Suri, David Alexandrovich Glushkov, Arnav Varma Dantuluri, Andrew Maguire, Christ...
2023
-
[15]
o rg Tiedemann, Andr \'e F. T. Martins, and Hinrich Sch \
Peiqin Lin, Shaoxiong Ji, J \"o rg Tiedemann, Andr \'e F. T. Martins, and Hinrich Sch \"u tze. 2024. http://arxiv.org/abs/2401.13303v2 MaLA-500 : Massive Language Adaptation of Large Language Models . arXiv preprint, arXiv:2401.13303v2
2024 arXiv
-
[16]
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mon...
2022
-
[17]
Lesley Milroy and Pieter Muysken, editors. 1995. https://www.cambridge.org/core/books/one-speaker-two-languages/864B8AA6972F95CB5603629264CF8324 One Speaker , Two Languages : Cross-Disciplinary Perspectives on Code-Switching . Cambridge University Press
1995
-
[18]
Fred Philippy, Siwen Guo, and Shohreh Haddadan. 2023. https://aclanthology.org/2023.acl-long.323 Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models : A Review . In Proceedings of the 61st Annual Meeting of the Asso...
2023
-
[19]
Esther Ploeger, Wessel Poelman, Miryam de Lhoneux , and Johannes Bjerva. 2024. https://aclanthology.org/2024.emnlp-main.326 What is `` Typological Diversity '' in NLP ? In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages 5681--5700
2024
-
[20]
Sebastian Ruder, Ivan Vuli \'c , and Anders S gaard. 2022. https://aclanthology.org/2022.findings-acl.184 Square One Bias in NLP : Towards a Multi-Dimensional Exploration of the Research Manifold . In Findings of the Association for Computational Linguistics : ACL 2022 , pages...
2022
-
[21]
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2023. https://openreview.net/forum?id=RIu5lyNXjT Quantifying Language Models ' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting . In The Twelfth Internationa...
2023
-
[22]
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. 2022. https://openreview.net/forum?id=fR3wGCk-IXp Language models are multilingual chain-of-thought reasone...
2022
-
[23]
Shivalika Singh, Freddie Vargus, Daniel D'souza, B \"o rje Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura O'Mahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Moura, Dominik Krzemi \'n ski, Hakimeh ...
2024
-
[24]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...
2023 arXiv
-
[25]
David L. Waltz. 1978. https://dl.acm.org/doi/10.1145/359545.359550 An English language question answering system for a large relational database . Communications of the ACM, 21(7):526--539
1978
-
[26]
Genta Winata, Alham Fikri Aji, Zheng Xin Yong, and Thamar Solorio. 2023. https://aclanthology.org/2023.findings-acl.185 The Decades Progress on Code-Switching Research in NLP : A Systematic Survey on Trends and Challenges . In Findings of the Association for Computational Ling...
2023
-
[27]
Terry Winograd. 1972. https://www.sciencedirect.com/science/article/pii/0010028572900023 Understanding natural language . Cognitive Psychology, 3(1):1--191
1972
-
[28]
Daniel Zeman and Philip Resnik. 2008. https://aclanthology.org/I08-3008 Cross- Language Parser Adaptation between Related Languages . In Proceedings of the IJCNLP-08 Workshop on NLP for Less Privileged Languages
2008
-
[29]
Xiang Zhang, Senyu Li, Bradley Hauer, Ning Shi, and Grzegorz Kondrak. 2023. https://aclanthology.org/2023.emnlp-main.491 Don't Trust ChatGPT when your Question is not in English : A Study of Multilingual Abilities and Types of LLMs . In Proceedings of the 2023 Conference on Em...
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.