REVIEW 4 major objections 5 minor 7 cited by
Multilingual Large Language Models: A Systematic Survey
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A systematic survey organizes multilingual large language models into six domains and argues that evaluation, interpretability, and cultural awareness are the open frontiers.
desk verdict Useful entry-level map of the MLLM landscape, but the 'comprehensive' claim outruns an undocumented and visibly selective curation process. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central instrument is the paper's six-domain taxonomy and roadmap, which classifies MLLM research into corpora, architectures, pre-training and tuning, evaluation, interpretability, and applications. This taxonomy organizes the literature and exposes cross-cutting mechanisms such as translation-assisted tuning, cross-lingual alignment, tokenizer evaluation, and per-language safety benchmarking. The companion curated paper list is meant to let readers trace each category back to primary sources.
What would settle it
A systematic literature search of the same period that finds a substantial body of MLLM research that cannot be placed in any of the six domains, such as multilingual agent frameworks that span both application and reasoning, would show the taxonomy is incomplete.
Extended reading notes
Core claim
The paper's central claim is that the MLLM research landscape can be captured by a six-domain taxonomy and roadmap, and that the central threads are cross-lingual transfer, language bias, and multilingual evaluation. On its own terms, the survey establishes that multilingual capability comes from large vocabularies and multilingual pre-training, that tuning transfers abilities from English-centric pivot languages to others, that evaluation must be done per language across benchmarks for reasoning, alignment, and safety, and that interpretability work is turning MLLMs from black boxes into white boxes. It also argues that domain-specific applications in medicine, computer science, mathematics, and law demonstrate real value, while low-resource languages and cultural adaptation remain open.
Load-bearing premise
The survey's load-bearing assumption is that its curated selection of papers is representative and its six-domain taxonomy complete, because no systematic search or inclusion criteria are documented.
Editorial extensions
If this is right
- A newcomer can use the six-domain taxonomy as a checklist for situating a new method or dataset.
- Multilingual evaluation needs language-by-language coverage, from high-resource to low-resource, across holistic, task-specific, alignment, and safety benchmarks.
- LLMs can act as multilingual evaluators, but their judgments are biased for low-resource and non-Latin script languages.
- Adapting MLLMs to new languages works through vocabulary expansion, continual pre-training, romanization, or encoder bridges, depending on data availability.
- Domain-specific MLLMs in medicine, law, code, and mathematics are a viable path to practical deployment.
Reading between the lines
- The six domains can be read as a lifecycle from data to deployment, an ordering the authors do not explicitly claim.
- The reported finding that MLLMs produce more unsafe responses in non-English languages suggests a testable follow-up: measure whether multilingual preference tuning closes that safety gap.
- Tokenizer design, treated as an evaluation topic, also acts as a cost and fairness lever; comparing downstream gains against tokenizer fertility and parity would quantify that link.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a survey of multilingual large language models (MLLMs), organized into six domains: architectures, multilingual corpora, pre-training and tuning, evaluation, interpretability, and applications. It provides taxonomies, tables of datasets and models, a GitHub repository of related papers, and a discussion of challenges and future directions. The survey's central claim is to be a comprehensive and systematic overview of the latest research on MLLMs.
Significance. If the comprehensiveness claim holds, the survey would be a valuable reference map of MLLM research, especially its structured taxonomy, dataset tables, and interpretability sections. The manuscript includes a substantial amount of curated information, including a public GitHub list of related papers. However, the absence of a documented selection methodology and the presence of internal inconsistencies mean the significance is currently conditional on these issues being addressed.
major comments (4)
- [Abstract, Section 2] The paper claims to be a 'comprehensive survey' and the title says 'Systematic Survey', but no methodology is provided for how the surveyed papers were selected. There is no description of search databases, date ranges, inclusion/exclusion criteria, or screening processes. Without this, the reader cannot assess whether the coverage is representative, which is load-bearing for the survey's central claim. The authors should add a methodology section and state the limitations of their selection.
- [Section 9] The application section is heavily skewed toward English and Chinese domain-specific models (e.g., DoctorGLM, HuatuoGPT, Lawformer, Lawyer LLaMA, ChatLaw), with few examples for other languages. The abstract and introduction promise coverage of 'diverse language communities'. This selection bias suggests the survey's coverage is not as comprehensive as claimed; the authors should either broaden the selection or explicitly acknowledge and justify the bias.
- [Section 7.2, Table 3] The abstract explicitly promises evaluation of 'reasoning' and Section 7.2 presents a task-specific evaluation taxonomy, but there is no category for multilingual reasoning benchmarks such as MGSM or similar multilingual math/reasoning datasets. This is a gap between the stated scope and the actual coverage, and the taxonomy should be extended or the omission justified.
- [Section 9.3] The description of Orca-Math states it is 'a small language model constructed with 700M parameters, is fine-tuned on the Mistral-7B architecture.' This is internally inconsistent: a model fine-tuned on Mistral-7B would have 7B parameters, not 700M. This factual error weakens the reliability of the application survey and should be corrected.
minor comments (5)
- [Section 8.2] The citation 'K et al. (2020)' is incomplete; the full reference is missing from the bibliography.
- [Section 5.2, Eq. (1)] The loss is written as L = -Σ_t p(x_{t+1}|x_{<t};θ), which is the negative of a probability, not a standard loss. It should likely be the negative log-likelihood, -Σ_t log p(x_{t+1}|x_{<t};θ).
- [Table 3] The 'Language Family' column contains overlapping and imprecise entries (e.g., 'Indo-European' appears in many rows, and some rows such as Jigsaw have an empty family). The table would benefit from careful proofreading.
- [Figure 1] Figure 1 contains the stray label 'MANUAL' which appears to be a leftover placeholder.
- [References] Some references are duplicated (e.g., DeepSeek-AI et al. 2024a and 2024b appear to be the same technical report).
Circularity Check
No significant circularity: the survey's claims are descriptive, and the few self-citations are used as external content rather than as load-bearing premises.
full rationale
This is a survey paper, so its central claims are descriptive and organizational rather than derivational. The six-part taxonomy in Figure 1 is presented as the authors' categorization scheme, not as a quantity derived from fitted inputs; the evaluation section lists external benchmarks; and the future-directions sections are qualitative. I checked the few citations to the authors' own group (e.g., Sun et al. FuxiTranyu in Section 3.1; Zhu et al. 2023b/2024e/f in Section 6.3.1; Cui et al. 2024 in Section 6.4.2). Each is used as an external published result to be surveyed (for example, 'Another MLLM FuxiTranyu (Sun et al., 2024) makes a balance between performance and training efficiency'), not as a premise that establishes the survey's comprehensiveness or taxonomy. There is no fitted parameter later renamed a prediction, no equation whose output equals its input, and no uniqueness or self-citation chain invoked to force a choice. The main vulnerability—absence of documented search or inclusion criteria and possible selection bias in the applications covered—is a coverage and transparency limitation, not an instance of circular reasoning. Accordingly, no circular step can be quoted with the required reduction, and the honest finding is no significant circularity.
Assumptions & free parameters
assumptions (2)
- domain assumption The facts and results reported in cited papers are accurately represented.
- ad hoc to paper The six-domain taxonomy is a complete and non-redundant organization of MLLM research.
Cite this review
Pith. "Pith review of Multilingual Large Language Models: A Systematic Survey." pith.science (2026). https://pith.science/paper/2SJQLOXT
@misc{pith2026241111072,
author = {Pith},
title = {Pith review of: Multilingual Large Language Models: A Systematic Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/2SJQLOXT}},
note = {Machine review of arXiv:2411.11072}
}
read the original abstract
This paper provides a comprehensive survey of the latest research on multilingual large language models (MLLMs). MLLMs not only are able to understand and generate language across linguistic boundaries, but also represent an important advancement in artificial intelligence. We first discuss the architecture and pre-training objectives of MLLMs, highlighting the key components and methodologies that contribute to their multilingual capabilities. We then discuss the construction of multilingual pre-training and alignment datasets, underscoring the importance of data quality and diversity in enhancing MLLM performance. An important focus of this survey is on the evaluation of MLLMs. We present a detailed taxonomy and roadmap covering the assessment of MLLMs' cross-lingual knowledge, reasoning, alignment with human values, safety, interpretability and specialized applications. Specifically, we extensively discuss multilingual evaluation benchmarks and datasets, and explore the use of LLMs themselves as multilingual evaluators. To enhance MLLMs from black to white boxes, we also address the interpretability of multilingual capabilities, cross-lingual transfer and language bias within these models. Finally, we provide a comprehensive review of real-world applications of MLLMs across diverse domains, including biology, medicine, computer science, mathematics and law. We showcase how these models have driven innovation and improvements in these specialized fields while also highlighting the challenges and opportunities in deploying MLLMs within diverse language communities and application scenarios. We listed the paper related in this survey and publicly available at https://github.com/tjunlp-lab/Awesome-Multilingual-LLMs-Papers.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 7 Pith papers
-
VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
A multilingual, multi-page document retrieval benchmark with 35K+ QA pairs shows MLLM retrievers lead but still fail on tables and low-resource languages.
-
A quantitative analysis of semantic information in deep representations of text and images
Using an asymmetric rank-based measure, the paper maps which layers of large language and vision models carry shared semantic information, finding central-layer peaks for translations and cross-modal pairs, and system...
-
HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong
A DeepSeek-based model fine-tuned for Hong Kong outperforms general models on Hong Kong benchmarks, but most of those benchmarks are self-authored and unreleased.
-
Just Go Parallel: Improving the Multilingual Capabilities of Large Language Models
Adding parallel data during continued pretraining improves a 1.1B LLM's translation and multilingual common-sense reasoning, with end-of-training placement performing best.
-
The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks
A large-scale analysis of multilingual benchmarks finds English overrepresentation, weak alignment of translated benchmarks with human preferences, and stronger alignment for localized benchmarks like CMMLU.
-
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
Selectively fine-tuning only 40% of a multimodal LLM's layers and neurons can match or slightly beat full fine-tuning on multilingual image-to-text translation benchmarks, though the measured gains are marginal.
-
FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation
FuxiMT combines a frozen BLOOMz model with sparse mixture-of-experts layers, Chinese-first pretraining, and curriculum learning to translate into Chinese from 65 languages, with claimed low-resource gains that the pap...
Reference graph
Works this paper leans on
-
[6]
doi: 10.48550/ARXIV.2311.04919. URL https://doi.org/10.48550/arXiv.2311. 04919. Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. The flores- 101 evaluation benchmark for low-resource and multilingual machine translation.Trans. Assoc. Comput. Lingu...
-
[7]
URL https://doi.org/10.48550/arXiv.2403
doi: 10.48550/ARXIV.2403.03814. URL https://doi.org/10.48550/arXiv.2403. 03814. Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, et al. Minicpm: Unveiling the potential of small language models with scalable training strategies.arXiv preprint arXiv:2404.06395, 2024. Wenyue Hua, Yuchen Zha...
-
[10]
URL https://doi.org/10.48550/arXiv.2203
doi: 10.48550/ARXIV.2203.07814. URL https://doi.org/10.48550/arXiv.2203. 07814. 79 Yunxiang Li, Zihan Li, Kai Zhang, Ruilong Dan, and You Zhang. Chatdoctor: A medical chat model fine-tuned on llama model using medical domain knowledge.CoRR, abs/2303.14070, 2023h. doi: 10.48550/ARXIV.2303.14070. URLhttps://doi.org/10.48550/arXiv.2303. 14070. Zihao Li, Shao...
-
[11]
URL https://doi.org/10.48550/arXiv.2401
doi: 10.48550/ARXIV.2401.13303. URL https://doi.org/10.48550/arXiv.2401. 13303. 80 Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, ...
-
[13]
URL https://doi.org/10.18653/v1/d18-1206
doi: 10.18653/V1/D18-1206. URL https://doi.org/10.18653/v1/d18-1206. Sina Bagheri Nezhad and Ameeta Agrawal. What drives performance in multilingual language models? arXiv preprint arXiv:2404.19159, 2024. 84 Ha-Thanh Nguyen. A brief report on lawgpt 1.0: A virtual legal assistant based on GPT-3. CoRR, abs/2302.05729, 2023. doi: 10.48550/ARXIV.2302.05729. ...
arXiv 2024
-
[15]
URL https://doi.org/10.18653/ v1/2022.findings-acl.103
doi: 10.18653/V1/2022.FINDINGS-ACL.103. URL https://doi.org/10.18653/ v1/2022.findings-acl.103. Aida Ramezani and Yang Xu. Knowledge of cultural moral norms in large language models. In Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (eds.),Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Paper...
-
[17]
URL https://doi.org/10.22364/bjmc.2022
doi: 10.22364/BJMC.2022.10.3.16. URL https://doi.org/10.22364/bjmc.2022. 10.3.16. Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Siamak Shakeri, Dara Bahri, Tal Schuster, et al. Ul2: Unifying language learning paradigms. arXiv preprint arXiv:2205.05131, 2022. Gemma Team, Thomas Mesnard, Cassidy Hardin, Rober...
arXiv 2022
-
[18]
URL https://doi.org/10.18653/v1/ 2023.findings-acl.21
doi: 10.18653/V1/2023.FINDINGS-ACL.21. URL https://doi.org/10.18653/v1/ 2023.findings-acl.21. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. David Vilares, Miguel A. Alonso, and Carlos Gómez-...
Show all 22 references
-
[20]
URL https://doi.org/10.48550/arXiv.2402
doi: 10.48550/ARXIV.2402.10588. URL https://doi.org/10.48550/arXiv.2402. 10588. Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. Ccnet: Extracting high quality monolingual 99 datasets from web crawl da...
-
[22]
URL https://doi.org/10.48550/arXiv.2305
doi: 10.48550/ARXIV.2305.10425. URL https://doi.org/10.48550/arXiv.2305. 10425. Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi, and Lidong Bing. How do large language models handle multilingualism?CoRR, abs/2402.18815, 2024c. doi: 10.48550/ARXIV.2402.18815. URL https...
- [211]
-
[224]
Nianwen Si, Hao Zhang, and Weiqiang Zhang
URL https://doi.org/10.18653/v1/2022.findings-acl.224. Nianwen Si, Hao Zhang, and Weiqiang Zhang. Mpn: Leveraging multilingual patch neuron for cross-lingual model editing.arXiv preprint arXiv:2401.03190, 2024. Shivalika Singh, Freddie Vargus, Daniel D’souza, Börje F. Karlsson...
2022 arXiv
-
[617]
Fred Philippy, Siwen Guo, and Shohreh Haddadan
URL https://doi.org/10.18653/v1/2020.emnlp-main.617. Fred Philippy, Siwen Guo, and Shohreh Haddadan. Towards a common understanding of contributing factors for cross-lingual transfer in multilingual language models: A review. In Anna Rogers, Jordan L. Boyd-Graber, and Naoaki O...
- [715]
-
[818]
Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel
URL https://doi.org/10.18653/v1/2021.emnlp-main.818. Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel. Few-shot parameter-efficient fine-tuning is better and cheaper than in- context learning. In Sanmi Koyejo, S. Mohamed, A. Aga...
-
[1356]
URLhttps://aclanthology
Association for Computational Linguistics, 2024b. URLhttps://aclanthology. org/2024.findings-eacl.90. Yijie Chen, Yijin Liu, Fandong Meng, Yufeng Chen, Jinan Xu, and Jie Zhou. Improving translation faithfulness of large language models via augmenting instructions.CoRR, abs/230...
-
[2018]
URL https://doi.org/10.18653/v1/d18-1269
doi: 10.18653/V1/D18-1269. URL https://doi.org/10.18653/v1/d18-1269. Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Un- supervised cross-lingual represent...
-
[2020]
Anubha Kabra, Emmy Liu, Simran Khanuja, Alham Fikri Aji, Genta Indra Winata, Samuel Cahyawijaya, Aremu Anuoluwapo, Perez Ogayo, and Graham Neubig
URL https://openreview.net/forum?id=HJeT3yrtDr. Anubha Kabra, Emmy Liu, Simran Khanuja, Alham Fikri Aji, Genta Indra Winata, Samuel Cahyawijaya, Aremu Anuoluwapo, Perez Ogayo, and Graham Neubig. Multi-lingual and multi-cultural figurative language understanding. In Anna Rogers...
-
[2021]
URL https://doi.org/10.18653/v1/ 2021.naacl-main.41
doi: 10.18653/V1/2021.NAACL-MAIN.41. URL https://doi.org/10.18653/v1/ 2021.naacl-main.41. Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. ByT5: Towards a token-free future with pre-trained byte-to-byte models. ...
-
[2022]
URL https://doi.org/10.48550/arXiv.2212
doi: 10.48550/ARXIV.2212.08204. URL https://doi.org/10.48550/arXiv.2212. 08204. Jing Huang and Diyi Yang. Culturally aware natural language inference. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.),Findings of the Association for Computational Lin- guistics: EMNLP 2023, S...
-
[2023]
AlexisConneauandGuillaumeLample
URL https://github.com/togethercomputer/RedPajama-Data. AlexisConneauandGuillaumeLample. Cross-linguallanguagemodelpretraining. InHannaM. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (eds.),Advances in Neural Information Pr...
2019
-
[2024]
URL https://doi.org/10.48550/arXiv.2401
doi: 10.48550/ARXIV.2401.05861. URL https://doi.org/10.48550/arXiv.2401. 05861. Iker García-Ferrero, Rodrigo Agerri, Aitziber Atutxa Salazar, Elena Cabrio, Iker de la Iglesia, Alberto Lavelli, Bernardo Magnini, Benjamin Molinet, Johana Ramirez-Romero, German Rigau, et al. Medi...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.