REVIEW 4 major objections 5 minor 3 cited by
Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study
T0 review · 4 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read The paper claims that a 9B open model, trained parallel-first, matches Google Translate and GPT-4-turbo across 28 languages.
desk verdict Useful data-recipe study and a strong open 9B translation model, but the leakage check is too casual and the recipe was selected on the same FLORES devtest used for final reporting. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
PFMS is the central mechanism. For each of the 28 languages the paper allocates a 2-billion-token continual-pretraining budget: it fills that budget with cleaned English-centric and Chinese-centric parallel sentence pairs from the OPUS collection as far as the available parallel data reaches, and then tops up the remainder with monolingual text from large public corpora. The paper's diagnostic is that high-resource languages already possess the generation ability the model needs, so they mainly need parallel pairs to align representations across languages; low- and mid-resource languages still lack generation capacity, so monolingual text matters for them. PFMS is the compromise that supplies parallel alignment where possible and monolingual mass where needed, and the experiments compare it against monolingual-only, 2:1, 1:1, 1:2, and parallel-only mixtures.
What would settle it
A decisive test would be to re-evaluate on a freshly created, human-translated test set for the same 28 languages, or to run a contamination search for FLORES-200 devtest sentences in the OPUS-derived training data, and check whether GemmaX2-28-9B still outperforms TowerInstruct and stays within the reported distance of Google Translate and GPT-4-turbo.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the ordering and ratio of parallel versus monolingual data in continual pretraining determines multilingual translation quality more than either data type alone. The authors' GemmaX2-28-9B—Gemma2-9B continually pretrained with the PFMS mixture and then instruction-finetuned on about 196,000 hand-filtered translation pairs—consistently outperforms open baselines such as TowerInstruct, X-ALMA, Aya, and LLaMAX on overlapping language directions, and its averaged scores on FLORES-200 and WMT-24 are comparable to Google Translate and GPT-4-turbo. The same recipe improves a 2B variant, which the paper takes as evidence that the strategy transfers across model sizes.
Load-bearing premise
The claim assumes the evaluation benchmarks—FLORES-200 devtest and WMT-24—are uncontaminated and valid proxies for quality, yet the finetuning data is partly drawn from FLORES-200 dev and the paper's leakage check is descriptive rather than a rigorous contamination test.
Editorial extensions
If this is right
- A sub-10B open model can match closed commercial systems on average across 28 languages, making high-quality translation available without closed APIs.
- Parallel data should be prioritized over monolingual data during continual pretraining for multilingual MT, reversing the emphasis of earlier monolingual-only recipes.
- The PFMS advantage appears at both 9B and 2B scale, so the recipe is not tied to one model size.
- Low- and mid-resource languages gain most from the PFMS mixture, suggesting capacity, not alignment, is their main bottleneck.
Reading between the lines
- An editor's inference: PFMS is a candidate general recipe for any multilingual backbone, since the paper shows the same ordering advantage at 2B and 9B scale; re-running the recipe on other base models would test that generality.
- An editor's inference: the per-language budget design suggests that for low-resource languages, monolingual volume is the binding constraint, so adding monolingual data may yield larger gains than collecting more parallel data for those languages.
- An editor's inference: a 9B model at this quality implies that offline, on-premise multilingual translation is feasible for privacy-sensitive or cost-constrained settings, a deployment path the paper does not discuss.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents an empirical study of multilingual machine translation (MT) with open-source LLMs under 10 billion parameters. The authors benchmark six models (Mistral-7B, Qwen2/2.5-7B, LLaMA3/3.1-8B, Gemma2-9B) across 28 languages on FLORES-200 and WMT-24, finding Gemma2-9B to be the strongest open model. They then propose a Parallel-First Monolingual-Second (PFMS) data mixing strategy for continual pretraining, followed by instruction finetuning on a small high-quality parallel dataset, yielding GemmaX2-28-9B and a 2B variant. The central claim is that GemmaX2-28-9B consistently outperforms existing open SOTA models such as TowerInstruct and X-ALMA, and is competitive with Google Translate and GPT-4-turbo.
Significance. If the reported results hold, the paper demonstrates that a 9B open model can approach production-grade multilingual translation across 28 languages, which is practically significant and would be a valuable resource for the community. The systematic comparison of data mixing ratios (monolingual-only, 2:1, 1:1, 1:2, parallel-only, PFMS) is a useful empirical contribution, and the public release of the models enhances reproducibility and independent verification. However, the strength of these contributions is contingent on addressing the evaluation-integrity concerns raised below, particularly the risk of benchmark contamination and the use of the evaluation set for recipe selection.
major comments (4)
- [§4.2, §5.1, §5.2, Tables 1–2] The evaluation-integrity check is not a contamination test. Section 4.2 states 'We do not observe serious data leakage issues... share a similar trend,' but comparing rankings across FLORES-200 and WMT-24 does not rule out memorization or near-duplicate overlap. The pretraining parallel data (Section 5.1) is drawn from the entire OPUS collection up to August 2024, which may include the WMT-24 test sets, and the SFT data (Section 5.2) is sampled from FLORES-200 dev and NTREX-128—the same benchmark family used for evaluation. Without an exact or approximate overlap analysis (e.g., n-gram or embedding similarity) between the training corpora and the FLORES-200 devtest and WMT-24 test sets, the headline numbers in Tables 1 and 2 could be inflated. Please add such an analysis, deduplicate against the test sets if overlaps are found, and report results on a truly held-out test set.
- [§5.4, Figures 4–5, Table 2] The PFMS recipe is selected using the same FLORES-200 devtest split on which the final results are reported. Figures 4 and 5 plot recipe performance on FLORES-200 devtest, and Table 2 then reports the selected model on that same split. This gives the recipe selection access to the test set, and the reported gains of PFMS over the alternatives may be partly due to selection on this benchmark. The paper should use a held-out validation split for recipe selection (e.g., a subset of FLORES-200 dev or a separate multilingual test set) and report final numbers on a test set not used in any design decision.
- [Tables 1, 2, 10–12] All scores are single-run point estimates with no variance or significance information. Many of the central comparisons are small; for example, in Table 2 (23-language row) the WMT-24 en→xx XCOMET difference between GemmaX2-28-9B (82.05) and X-ALMA (81.67) is 0.38 points, and per-direction results in Table 10 show GemmaX2 losing on en→ms (78.29 vs 81.13). Without multiple runs, bootstrap confidence intervals, or significance tests, the claim that GemmaX2 'consistently outperforms' these models is not statistically supported. Please provide uncertainty estimates or temper the claim to 'on average, in these evaluations.'
- [Limitations section (after Conclusion)] The Limitations section identifies only the constraint of compute and model scale. It does not mention the contamination risk, the use of the evaluation set for recipe selection, or the lack of significance testing—all of which are primary threats to the validity of the paper's central claim. The paper should either address these issues experimentally or explicitly state them as limitations; as written, the stated limitations omit the most consequential threats to the findings.
minor comments (5)
- [Abstract] The term 'XALMA' should be written as 'X-ALMA' for consistency with the body text and with the cited work.
- [Footnote 2] The model release URL 'https://huggingface/GemmaX2' appears incomplete; it should link to the actual repository (e.g., a huggingface.co address).
- [Table 8 caption] The caption contains a typo: 'finetuing' should be 'finetuning.'
- [§5.2] The phrase 'either from the NTREX-128 and FLORES-200 dev datasets or the OPUS dataset' is awkward; consider 'from either the NTREX-128/FLORES-200 dev datasets or the OPUS dataset.'
- [§3.3] The WMT-24 evaluation uses reference-free XCOMET and COMETKiwi; the paper should note that these are learned metrics and may not perfectly reflect human judgments, and ideally it should report at least one reference-based metric for a subset of directions.
Circularity Check
No significant circularity: the paper is an empirical benchmark study; the main risks are contamination/model-selection on FLORES family, not circular reasoning.
full rationale
This paper contains no formal derivation whose conclusion is equivalent to its inputs. The central claims are empirical contrasts (GemmaX2-28-9B vs. open and closed baselines) measured on FLORES-200 devtest and WMT-24, with metrics spBLEU/COMET/XCOMET/COMETKiwi. The PFMS data-mixing strategy is not defined in terms of the evaluation scores; it is a data-construction rule ('utilize parallel data as much as possible and supplement it with monolingual data') selected after comparing five recipes on development splits. The only self-citations (Cui et al. 2024a and Gao et al. 2024) appear in related-work context and are not load-bearing for the reported results. The nearest concern is benchmark contamination/selection: Section 5.2 samples finetuning pairs from FLORES-200 dev and NTREX-128, and Figures 4-5 choose the recipe on FLORES-200 devtest, so the same benchmark family influences training decisions and final evaluation; Section 4.2's leakage check is only a trend comparison. This is a validity threat, not a circularity reduction: no fitted parameter is renamed as a prediction, and no equation equates the claimed outcome with the input by construction.
Assumptions & free parameters
free parameters (3)
- Per-language token budget in PFMS =
2 billion tokens
- Parallel data similarity filtering thresholds =
0.75 to 0.99 LaBSE/MuSR similarity
- Minimum parallel tokens before adding monolingual =
up to 2B parallel tokens per language; monolingual used only when parallel is insufficient
assumptions (4)
- domain assumption CulturaX and MADLAD-400 provide sufficiently clean monolingual data for continual pretraining.
- domain assumption The OPUS-derived parallel corpus plus fastText/LaBSE/MuSR filtering yields high-quality parallel pairs.
- domain assumption spBLEU, COMET, XCOMET, and COMETKiwi scores are reliable proxies for translation quality.
- domain assumption The FLORES-200 dev and devtest sets are not contaminated with the pretraining or finetuning data.
Cite this review
Pith. "Pith review of Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study." pith.science (2026). https://pith.science/paper/QPGU7RWX
@misc{pith2026250202481,
author = {Pith},
title = {Pith review of: Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study},
year = {2026},
howpublished = {\url{https://pith.science/paper/QPGU7RWX}},
note = {Machine review of arXiv:2502.02481}
}
read the original abstract
Large language models (LLMs) have shown continuously improving multilingual capabilities, and even small-scale open-source models have demonstrated rapid performance enhancement. In this paper, we systematically explore the abilities of open LLMs with less than ten billion parameters to handle multilingual machine translation (MT) tasks. We conduct comprehensive evaluations on six popular LLMs and find that models like Gemma2-9B exhibit impressive multilingual translation capabilities. We then introduce the Parallel-First Monolingual-Second (PFMS) data mixing strategy in the continual pretraining stage to further enhance the MT performance and present GemmaX2-28, a 9B model achieving top-tier multilingual translation performance across 28 languages. Specifically, GemmaX2-28 consistently outperforms the state-of-the-art (SOTA) models such as TowerInstruct and XALMA and achieves competitive performance with Google Translate and GPT-4-turbo.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 3 Pith papers
-
Hunyuan-MT Technical Report
Hunyuan-MT and Chimera, a 7B open-source translation model and its multi-candidate fusion variant, claim state-of-the-art multilingual translation including Mandarin to minority languages, with open weights.
-
QUST_NLP at SemEval-2025 Task 7: A Three-Stage Retrieval Framework for Monolingual and Crosslingual Fact-Checked Claim Retrieval
A three-stage ensemble of retrieval models, rerankers, and weighted voting achieves strong multilingual fact-checked claim retrieval results at SemEval-2025 Task 7.
-
KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025
KIT combines LLM-based ASR fusion with quality-filtered fine-tuning and post-editing for offline speech translation, and a contrastively pretrained SpeechLLM with post-editing for multilingual instruction following.
Reference graph
Works this paper leans on
-
[1]
Sweta Agrawal, Chunting Zhou, Mike Lewis, Luke Zettlemoyer, and Marjan Ghazvininejad. 2023. https://doi.org/10.18653/v1/2023.findings-acl.564 In-context examples selection for machine translation . In Findings of the Association for Computational Linguistics: ACL 2023, pages 8857--8873, Toronto, Canada. Association for Computational Linguistics
-
[2]
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023. https://arxiv.org/abs/2311.16867 The falcon series of open language models . Preprint, arXiv:2...
arXiv 2023
-
[3]
Duarte Miguel Alves, Jos \'e Pombal, Nuno M Guerreiro, Pedro Henrique Martins, Jo \ a o Alves, Amin Farajian, Ben Peters, Ricardo Rei, Patrick Fernandes, Sweta Agrawal, Pierre Colombo, Jos \'e G. C. de Souza, and Andre Martins. 2024. https://openreview.net/forum?id=EHPns3hVkj Tower: An open multilingual large language model for translation-related tasks ....
2024
-
[4]
Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Z...
arXiv 2023
-
[5]
Akari Asai, Sneha Kudugunta, Xinyan Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, and Hannaneh Hajishirzi. 2024. https://doi.org/10.18653/v1/2024.naacl-long.100 BUFFET : Benchmarking large language models for few-shot cross-lingual transfer . In Proceedings of the 2024 Conference of the North American Chapter of the Associat...
-
[6]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradb...
arXiv 2022
-
[7]
Christos Christodouloupoulos and Mark Steedman. 2015. A massively parallel corpus: the bible in 100 languages. Language resources and evaluation, 49:375--395
work page 2015
-
[8]
Menglong Cui, Jiangcun Du, Shaolin Zhu, and Deyi Xiong. 2024 a . https://aclanthology.org/2024.findings-acl.646 Efficiently exploring large language models for document-level machine translation with in-context learning . In Findings of the Association for Computational Linguistics ACL 2024, pages 10885--10897, Bangkok, Thailand and virtual meeting. Assoc...
work page 2024
Show all 63 references
-
[9]
Yiming Cui, Ziqing Yang, and Xin Yao. 2024 b . https://arxiv.org/abs/2304.08177 Efficient and effective text encoding for chinese llama and alpaca . Preprint, arXiv:2304.08177
2024 arXiv
-
[10]
John Dang, Shivalika Singh, Daniel D'souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, Sandra Kublik, Meor Amer, Viraat Aryabumi, Jon Ander Campos, Yi-Chern Tan, Tom Kocmi, Florian Strub, Nathan Grinsztaj...
2024 arXiv
-
[11]
Bosheng Ding, Chengwei Qin, Ruochen Zhao, Tianze Luo, Xinze Li, Guizhen Chen, Wenhan Xia, Junjie Hu, Anh Tuan Luu, and Shafiq Joty. 2024. https://aclanthology.org/2024.findings-acl.97 Data augmentation using LLM s: Data perspectives, learning paradigms and challenges . In Find...
2024
-
[12]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zha...
2024 arXiv
-
[13]
Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Gim \'e nez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, and Katharina...
2022 doi
-
[14]
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzm \'a n, and Philipp Koehn. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.480 CCA ligned: A massive collection of cross-lingual web-document pairs . In Proceedings of the 2020 Conference on Empirical Methods in Natural Langu...
2020 doi
-
[15]
Christian Federmann, Tom Kocmi, and Ying Xin. 2022. https://aclanthology.org/2022.sumeval-1.4 NTREX -128 -- news test references for MT evaluation of 128 languages . In Proceedings of the First Workshop on Scaling Up Multilingual Evaluation, pages 21--24, Online. Association f...
2022
-
[16]
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2022. https://doi.org/10.18653/v1/2022.acl-long.62 Language-agnostic BERT sentence embedding . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long...
2022 doi
-
[17]
Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster. 2023. https://doi.org/10.18653/v1/2023.wmt-1.51 Results of ...
2023 doi
-
[18]
Pengzhi Gao, Zhongjun He, Hua Wu, and Haifeng Wang. 2024. https://arxiv.org/abs/2401.05861 Towards boosting many-to-many multilingual machine translation with large language models . Preprint, arXiv:2401.05861
2024 arXiv
-
[19]
Pengzhi Gao, Liwen Zhang, Zhongjun He, Hua Wu, and Haifeng Wang. 2023. https://doi.org/10.18653/v1/2023.emnlp-industry.25 Learning multilingual sentence representations with cross-lingual consistency regularization . In Proceedings of the 2023 Conference on Empirical Methods i...
2023 doi
-
[20]
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc ' Aurelio Ranzato, Francisco Guzm \'a n, and Angela Fan. 2022. https://doi.org/10.1162/tacl_a_00474 The F lores-101 evaluation benchmark for low-resource and multilingua...
2022 doi
-
[21]
Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and André F
Nuno M. Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and André F. T. Martins. 2023. https://arxiv.org/abs/2310.10482 xcomet: Transparent machine translation evaluation through fine-grained error detection . Preprint, arXiv:2310.10482
2023 arXiv
-
[22]
Jiaxin Guo, Hao Yang, Zongyao Li, Daimeng Wei, Hengchao Shang, and Xiaoyu Chen. 2024. https://doi.org/10.18653/v1/2024.findings-naacl.42 A novel paradigm boosting translation capabilities of large language models . In Findings of the Association for Computational Linguistics: ...
2024 doi
-
[23]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
2023 arXiv
-
[24]
Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Xing Wang, Shuming Shi, and Zhaopeng Tu. 2023. https://arxiv.org/abs/2301.08745 Is chatgpt a good translator? yes with gpt-4 as the engine . Preprint, arXiv:2301.08745
2023 arXiv
-
[25]
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. https://doi.org/10.18653/v1/2020.acl-main.560 The state and fate of linguistic diversity and inclusion in the NLP world . In Proceedings of the 58th Annual Meeting of the Association for Co...
2020 doi
-
[26]
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. 2016. https://arxiv.org/abs/1612.03651 Fasttext.zip: Compressing text classification models . Preprint, arXiv:1612.03651
2016 arXiv
-
[27]
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017. https://aclanthology.org/E17-2068 Bag of tricks for efficient text classification . In Proceedings of the 15th Conference of the E uropean Chapter of the Association for Computational Linguistics: Volume ...
2017
-
[28]
Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat
Sneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Christopher A. Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat. 2023. https://arxiv.org/abs/2309.04662 Madlad-400: A multilingual and document-level large audited ...
2023 arXiv
-
[29]
Chong Li, Shaonan Wang, Jiajun Zhang, and Chengqing Zong. 2024 a . https://doi.org/10.18653/v1/2024.naacl-long.445 Improving in-context learning of multilingual generative language models with cross-lingual alignment . In Proceedings of the 2024 Conference of the North America...
2024 doi
-
[30]
Jiahuan Li, Hao Zhou, Shujian Huang, Shanbo Cheng, and Jiajun Chen. 2024 b . https://doi.org/10.1162/tacl_a_00655 Eliciting the translation ability of large language models via multilingual finetuning with translation instructions . Transactions of the Association for Computat...
2024 doi
-
[31]
Baohao Liao, Christian Herold, Shahram Khadivi, and Christof Monz. 2024. https://arxiv.org/abs/2408.11512 Ikun for wmt24 general mt task: Llms are here for multilingual machine translation . Preprint, arXiv:2408.11512
2024 arXiv
-
[32]
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O ' Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, M...
2022 doi
-
[33]
Yinquan Lu, Wenhao Zhu, Lei Li, Yu Qiao, and Fei Yuan. 2024. https://arxiv.org/abs/2407.05975 Llamax: Scaling linguistic horizons of llm by enhancing translation capabilities beyond 100 languages . Preprint, arXiv:2407.05975
2024 arXiv
-
[34]
Rossi, and Thien Huu Nguyen
Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. 2024. https://aclanthology.org/2024.lrec-main.377 C ultura X : A cleaned, enormous, and multilingual dataset for large language models in 167 langu...
2024
-
[35]
OpenAI. 2023 a . Chatgpt [large language model]. https://chat.openai.com/
2023
-
[36]
OpenAI. 2023 b . https://arxiv.org/abs/2303.08774 Gpt-4 technical report . Preprint, arXiv:2303.08774
2023 arXiv
-
[37]
Matt Post. 2018. https://doi.org/10.18653/v1/W18-6319 A call for clarity in reporting BLEU scores . In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186--191, Brussels, Belgium. Association for Computational Linguistics
2018 doi
-
[38]
Guerreiro, Jos \ A Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, Jos \'e G
Ricardo Rei, Nuno M. Guerreiro, Jos \ A Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, Jos \'e G. C. de Souza, and Andr \'e Martins. 2023. https://doi.org/10.18653/v1/2023.wmt-1.73 Scaling up C omet K iwi: Unbabel- IST 2023 submission for the quality estimation shared t...
2023 doi
-
[39]
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.213 COMET : A neural framework for MT evaluation . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685--2702, ...
2020 doi
-
[40]
Holger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave, Armand Joulin, and Angela Fan. 2021. https://doi.org/10.18653/v1/2021.acl-long.507 CCM atrix: Mining billions of high-quality parallel sentences on the web . In Proceedings of the 59th Annual Meeting of the Associ...
2021 doi
-
[41]
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan...
2024 arXiv
-
[42]
NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prang...
2022 arXiv
-
[43]
J \"o rg Tiedemann. 2016. https://aclanthology.org/L16-1559 Finding alternative translations in a large corpus of movie subtitle . In Proceedings of the Tenth International Conference on Language Resources and Evaluation ( LREC '16) , pages 3518--3522, Portoro z , Slovenia. Eu...
2016
-
[44]
Jörg Tiedemann. 2012. Parallel data, tools and interfaces in opus. In Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12), Istanbul, Turkey. European Language Resources Association (ELRA)
2012
-
[45]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 a . https://arxiv.org/abs/2302.13971 Lla...
2023 arXiv
-
[46]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...
2023 arXiv
-
[47]
Ahmet \"U st \"u n, Viraat Aryabumi, Zheng Yong, Wei-Yin Ko, Daniel D ' souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker. 2024. ht...
2024 doi
-
[48]
David Vilar, Markus Freitag, Colin Cherry, Jiaming Luo, Viresh Ratnakar, and George Foster. 2023. https://doi.org/10.18653/v1/2023.acl-long.859 Prompting P a LM for translation: Assessing strategies and performance . In Proceedings of the 61st Annual Meeting of the Association...
2023 doi
-
[49]
Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. https://arxiv.org/abs/2206.07682 Emergen...
2022 arXiv
-
[50]
Genta Indra Winata, Andrea Madotto, Zhaojiang Lin, Rosanne Liu, Jason Yosinski, and Pascale Fung. 2021. https://doi.org/10.18653/v1/2021.mrl-1.1 Language models are few-shot multilingual learners . In Proceedings of the 1st Workshop on Multilingual Representation Learning, pag...
2021 doi
-
[51]
BigScience Workshop, :, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanch...
2023 arXiv
-
[52]
Di Wu, Shaomu Tan, Yan Meng, David Stap, and Christof Monz. 2024. https://aclanthology.org/2024.findings-acl.896 How far can 100 samples go? unlocking zero-shot translation with tiny multi-parallel data . In Findings of the Association for Computational Linguistics ACL 2024, p...
2024
-
[53]
Zhenyu Wu, Yaoxiang Wang, Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Jingjing Xu, and Yu Qiao. 2023. https://doi.org/10.18653/v1/2023.acl-demo.47 O pen ICL : An open-source framework for in-context learning . In Proceedings of the 61st Annual Meeting of the Association for Comput...
2023 doi
-
[54]
Haoran Xu, Young Jin Kim, Amr Sharaf, and Hany Hassan Awadalla. 2024 a . https://openreview.net/forum?id=farT6XXntP A paradigm shift in machine translation: Boosting translation performance of large language models . In The Twelfth International Conference on Learning Represen...
2024
-
[55]
Haoran Xu, Kenton Murray, Philipp Koehn, Hieu Hoang, Akiko Eriguchi, and Huda Khayrallah. 2024 b . https://arxiv.org/abs/2410.03115 X-alma: Plug & play modules and adaptive rejection for quality translation at scale . Preprint, arXiv:2410.03115
2024 arXiv
-
[56]
Haoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan, Lingfeng Shen, Benjamin Van Durme, Kenton Murray, and Young Jin Kim. 2024 c . https://openreview.net/forum?id=51iwkioZpn Contrastive preference optimization: Pushing the boundaries of LLM performance in machine translation . In F...
2024
-
[57]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze...
2024 arXiv
-
[58]
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2...
2022 arXiv
-
[59]
Yaowei Zheng, Richong Zhang, Junhao Zhang, YeYanhan YeYanhan, and Zheyan Luo. 2024. https://aclanthology.org/2024.acl-demos.38 L lama F actory: Unified efficient fine-tuning of 100+ language models . In Proceedings of the 62nd Annual Meeting of the Association for Computationa...
2024
-
[60]
Shaolin Zhu, Menglong Cui, and Deyi Xiong. 2024 a . https://aclanthology.org/2024.lrec-main.1444 Towards robust in-context learning for machine translation with large language models . In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Lang...
2024
-
[61]
Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li. 2024 b . https://doi.org/10.18653/v1/2024.findings-naacl.176 Multilingual machine translation with large language models: Empirical results and analysis . In Findings of t...
2024 doi
-
[62]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[63]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.