Pith. sign in

REVIEW 4 major objections 5 minor 3 cited by

Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study

T0 review · 4 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read The paper claims that a 9B open model, trained parallel-first, matches Google Translate and GPT-4-turbo across 28 languages.

desk verdict Useful data-recipe study and a strong open 9B translation model, but the leakage check is too casual and the recipe was selected on the same FLORES devtest used for final reporting. read the letter →

arxiv 2502.02481 v4 pith:QPGU7RWX submitted 2025-02-04 cs.CL

classification cs.CL
keywords multilingualmachinetranslationopenlargelanguagemodelscontinualpretrainingparallel-firstdatamixingmonolingualvsparallellow-resourcelanguagesGemma2many-to-many
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Large language models under ten billion parameters are often treated as too small for production translation, and this paper asks whether that is still true. It benchmarks six open models across 28 languages, identifies Gemma2-9B as the strongest base, and then shows that a particular data-ordering choice—parallel sentence pairs first, monolingual text only as filler—during continued pretraining, followed by a small high-quality finetuning set, yields a 9B model that beats current open translation models and matches Google Translate and GPT-4-turbo. The paper thus claims that a sub-10B open model can reach commercial-grade multilingual translation. If the claim holds, practical translation no longer requires closed APIs or very large proprietary models.

What carries the argument

PFMS is the central mechanism. For each of the 28 languages the paper allocates a 2-billion-token continual-pretraining budget: it fills that budget with cleaned English-centric and Chinese-centric parallel sentence pairs from the OPUS collection as far as the available parallel data reaches, and then tops up the remainder with monolingual text from large public corpora. The paper's diagnostic is that high-resource languages already possess the generation ability the model needs, so they mainly need parallel pairs to align representations across languages; low- and mid-resource languages still lack generation capacity, so monolingual text matters for them. PFMS is the compromise that supplies parallel alignment where possible and monolingual mass where needed, and the experiments compare it against monolingual-only, 2:1, 1:1, 1:2, and parallel-only mixtures.

What would settle it

A decisive test would be to re-evaluate on a freshly created, human-translated test set for the same 28 languages, or to run a contamination search for FLORES-200 devtest sentences in the OPUS-derived training data, and check whether GemmaX2-28-9B still outperforms TowerInstruct and stays within the reported distance of Google Translate and GPT-4-turbo.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that the ordering and ratio of parallel versus monolingual data in continual pretraining determines multilingual translation quality more than either data type alone. The authors' GemmaX2-28-9B—Gemma2-9B continually pretrained with the PFMS mixture and then instruction-finetuned on about 196,000 hand-filtered translation pairs—consistently outperforms open baselines such as TowerInstruct, X-ALMA, Aya, and LLaMAX on overlapping language directions, and its averaged scores on FLORES-200 and WMT-24 are comparable to Google Translate and GPT-4-turbo. The same recipe improves a 2B variant, which the paper takes as evidence that the strategy transfers across model sizes.

Load-bearing premise

The claim assumes the evaluation benchmarks—FLORES-200 devtest and WMT-24—are uncontaminated and valid proxies for quality, yet the finetuning data is partly drawn from FLORES-200 dev and the paper's leakage check is descriptive rather than a rigorous contamination test.

Editorial extensions

If this is right

  • A sub-10B open model can match closed commercial systems on average across 28 languages, making high-quality translation available without closed APIs.
  • Parallel data should be prioritized over monolingual data during continual pretraining for multilingual MT, reversing the emphasis of earlier monolingual-only recipes.
  • The PFMS advantage appears at both 9B and 2B scale, so the recipe is not tied to one model size.
  • Low- and mid-resource languages gain most from the PFMS mixture, suggesting capacity, not alignment, is their main bottleneck.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An editor's inference: PFMS is a candidate general recipe for any multilingual backbone, since the paper shows the same ordering advantage at 2B and 9B scale; re-running the recipe on other base models would test that generality.
  • An editor's inference: the per-language budget design suggests that for low-resource languages, monolingual volume is the binding constraint, so adding monolingual data may yield larger gains than collecting more parallel data for those languages.
  • An editor's inference: a 9B model at this quality implies that offline, on-premise multilingual translation is feasible for privacy-sensitive or cost-constrained settings, a deployment path the paper does not discuss.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper presents an empirical study of multilingual machine translation (MT) with open-source LLMs under 10 billion parameters. The authors benchmark six models (Mistral-7B, Qwen2/2.5-7B, LLaMA3/3.1-8B, Gemma2-9B) across 28 languages on FLORES-200 and WMT-24, finding Gemma2-9B to be the strongest open model. They then propose a Parallel-First Monolingual-Second (PFMS) data mixing strategy for continual pretraining, followed by instruction finetuning on a small high-quality parallel dataset, yielding GemmaX2-28-9B and a 2B variant. The central claim is that GemmaX2-28-9B consistently outperforms existing open SOTA models such as TowerInstruct and X-ALMA, and is competitive with Google Translate and GPT-4-turbo.

Significance. If the reported results hold, the paper demonstrates that a 9B open model can approach production-grade multilingual translation across 28 languages, which is practically significant and would be a valuable resource for the community. The systematic comparison of data mixing ratios (monolingual-only, 2:1, 1:1, 1:2, parallel-only, PFMS) is a useful empirical contribution, and the public release of the models enhances reproducibility and independent verification. However, the strength of these contributions is contingent on addressing the evaluation-integrity concerns raised below, particularly the risk of benchmark contamination and the use of the evaluation set for recipe selection.

major comments (4)
  1. [§4.2, §5.1, §5.2, Tables 1–2] The evaluation-integrity check is not a contamination test. Section 4.2 states 'We do not observe serious data leakage issues... share a similar trend,' but comparing rankings across FLORES-200 and WMT-24 does not rule out memorization or near-duplicate overlap. The pretraining parallel data (Section 5.1) is drawn from the entire OPUS collection up to August 2024, which may include the WMT-24 test sets, and the SFT data (Section 5.2) is sampled from FLORES-200 dev and NTREX-128—the same benchmark family used for evaluation. Without an exact or approximate overlap analysis (e.g., n-gram or embedding similarity) between the training corpora and the FLORES-200 devtest and WMT-24 test sets, the headline numbers in Tables 1 and 2 could be inflated. Please add such an analysis, deduplicate against the test sets if overlaps are found, and report results on a truly held-out test set.
  2. [§5.4, Figures 4–5, Table 2] The PFMS recipe is selected using the same FLORES-200 devtest split on which the final results are reported. Figures 4 and 5 plot recipe performance on FLORES-200 devtest, and Table 2 then reports the selected model on that same split. This gives the recipe selection access to the test set, and the reported gains of PFMS over the alternatives may be partly due to selection on this benchmark. The paper should use a held-out validation split for recipe selection (e.g., a subset of FLORES-200 dev or a separate multilingual test set) and report final numbers on a test set not used in any design decision.
  3. [Tables 1, 2, 10–12] All scores are single-run point estimates with no variance or significance information. Many of the central comparisons are small; for example, in Table 2 (23-language row) the WMT-24 en→xx XCOMET difference between GemmaX2-28-9B (82.05) and X-ALMA (81.67) is 0.38 points, and per-direction results in Table 10 show GemmaX2 losing on en→ms (78.29 vs 81.13). Without multiple runs, bootstrap confidence intervals, or significance tests, the claim that GemmaX2 'consistently outperforms' these models is not statistically supported. Please provide uncertainty estimates or temper the claim to 'on average, in these evaluations.'
  4. [Limitations section (after Conclusion)] The Limitations section identifies only the constraint of compute and model scale. It does not mention the contamination risk, the use of the evaluation set for recipe selection, or the lack of significance testing—all of which are primary threats to the validity of the paper's central claim. The paper should either address these issues experimentally or explicitly state them as limitations; as written, the stated limitations omit the most consequential threats to the findings.
minor comments (5)
  1. [Abstract] The term 'XALMA' should be written as 'X-ALMA' for consistency with the body text and with the cited work.
  2. [Footnote 2] The model release URL 'https://huggingface/GemmaX2' appears incomplete; it should link to the actual repository (e.g., a huggingface.co address).
  3. [Table 8 caption] The caption contains a typo: 'finetuing' should be 'finetuning.'
  4. [§5.2] The phrase 'either from the NTREX-128 and FLORES-200 dev datasets or the OPUS dataset' is awkward; consider 'from either the NTREX-128/FLORES-200 dev datasets or the OPUS dataset.'
  5. [§3.3] The WMT-24 evaluation uses reference-free XCOMET and COMETKiwi; the paper should note that these are learned metrics and may not perfectly reflect human judgments, and ideally it should report at least one reference-based metric for a subset of directions.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is an empirical benchmark study; the main risks are contamination/model-selection on FLORES family, not circular reasoning.

full rationale

This paper contains no formal derivation whose conclusion is equivalent to its inputs. The central claims are empirical contrasts (GemmaX2-28-9B vs. open and closed baselines) measured on FLORES-200 devtest and WMT-24, with metrics spBLEU/COMET/XCOMET/COMETKiwi. The PFMS data-mixing strategy is not defined in terms of the evaluation scores; it is a data-construction rule ('utilize parallel data as much as possible and supplement it with monolingual data') selected after comparing five recipes on development splits. The only self-citations (Cui et al. 2024a and Gao et al. 2024) appear in related-work context and are not load-bearing for the reported results. The nearest concern is benchmark contamination/selection: Section 5.2 samples finetuning pairs from FLORES-200 dev and NTREX-128, and Figures 4-5 choose the recipe on FLORES-200 devtest, so the same benchmark family influences training decisions and final evaluation; Section 4.2's leakage check is only a trend comparison. This is a validity threat, not a circularity reduction: no fitted parameter is renamed as a prediction, and no equation equates the claimed outcome with the input by construction.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on empirical recipe choices, not on a derivation. The main ad hoc choices are the token budget and filtering thresholds, and the main domain assumptions are about data quality and metric validity. No new theoretical entities are introduced.

free parameters (3)
  • Per-language token budget in PFMS = 2 billion tokens
    The choice of 2B tokens per language is arbitrary; no sensitivity analysis is provided, and the central result depends on this budget (Section 5.4/5.5).
  • Parallel data similarity filtering thresholds = 0.75 to 0.99 LaBSE/MuSR similarity
    The filtering thresholds are chosen by hand; the paper does not justify them with experiments (Section 5.1).
  • Minimum parallel tokens before adding monolingual = up to 2B parallel tokens per language; monolingual used only when parallel is insufficient
    The PFMS rule 'use as much parallel as possible' is a heuristic; the cutoff for 'as much as possible' is not systematically optimized (Section 5.5).
assumptions (4)
  • domain assumption CulturaX and MADLAD-400 provide sufficiently clean monolingual data for continual pretraining.
    Stated in Section 5.1; the paper relies on the cleaning and deduplication claimed by the dataset authors.
  • domain assumption The OPUS-derived parallel corpus plus fastText/LaBSE/MuSR filtering yields high-quality parallel pairs.
    Section 5.1; no manual quality audit is reported.
  • domain assumption spBLEU, COMET, XCOMET, and COMETKiwi scores are reliable proxies for translation quality.
    Section 3.3; the paper uses these metrics without validating them on this language set.
  • domain assumption The FLORES-200 dev and devtest sets are not contaminated with the pretraining or finetuning data.
    Section 4.2 argues no serious leakage based on trend similarity with WMT-24, but this is not a rigorous contamination test.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study." pith.science (2026). https://pith.science/paper/QPGU7RWX

@misc{pith2026250202481,
  author       = {Pith},
  title        = {Pith review of: Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QPGU7RWX}},
  note         = {Machine review of arXiv:2502.02481}
}
read the original abstract

Large language models (LLMs) have shown continuously improving multilingual capabilities, and even small-scale open-source models have demonstrated rapid performance enhancement. In this paper, we systematically explore the abilities of open LLMs with less than ten billion parameters to handle multilingual machine translation (MT) tasks. We conduct comprehensive evaluations on six popular LLMs and find that models like Gemma2-9B exhibit impressive multilingual translation capabilities. We then introduce the Parallel-First Monolingual-Second (PFMS) data mixing strategy in the continual pretraining stage to further enhance the MT performance and present GemmaX2-28, a 9B model achieving top-tier multilingual translation performance across 28 languages. Specifically, GemmaX2-28 consistently outperforms the state-of-the-art (SOTA) models such as TowerInstruct and XALMA and achieves competitive performance with Google Translate and GPT-4-turbo.

Figures

Figures reproduced from arXiv: 2502.02481 by the authors.

Figure 1
Figure 1. The tokenizer efficiency of open-source LLMs for each non-English language. The smaller the length [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. MT performance on the FLORES-200 bench￾mark with different numbers of in-context exemplars. 5 GemmaX: Boosting Multilingual Translation with Gemma Models Given its impressive multilingual capabilities dis￾cussed in Section 4, we select Gemma2-9B as our backbone model for learning many-to-many mul￾tilingual machine translation across 28 languages. We continue the pretraining of Gemma2-9B with multilingual corpora and… view at source ↗
Figure 3
Figure 3. Number of sentences in different languages for Chinese-centric and English-centric parallel dataset. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: The translation performance (COMET) of models trained with different data recipes during continual [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: The translation performance (BLEU) of models trained with different data recipes during continual [PITH_FULL_IMAGE:figures/full_fig_p020_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hunyuan-MT Technical Report

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Hunyuan-MT and Chimera, a 7B open-source translation model and its multi-candidate fusion variant, claim state-of-the-art multilingual translation including Mandarin to minority languages, with open weights.

  2. QUST_NLP at SemEval-2025 Task 7: A Three-Stage Retrieval Framework for Monolingual and Crosslingual Fact-Checked Claim Retrieval

    cs.IR 2025-06 conditional novelty 4.0 of 10

    A three-stage ensemble of retrieval models, rerankers, and weighted voting achieves strong multilingual fact-checked claim retrieval results at SemEval-2025 Task 7.

  3. KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025

    cs.CL 2025-05 conditional novelty 4.0 of 10

    KIT combines LLM-based ASR fusion with quality-filtered fine-tuning and post-editing for offline speech translation, and a contrastively pretrained SpeechLLM with post-editing for multilingual instruction following.

Reference graph

Works this paper leans on

63 extracted references · 14 canonical work pages · cited by 3 Pith papers

  1. [1]

    Sweta Agrawal, Chunting Zhou, Mike Lewis, Luke Zettlemoyer, and Marjan Ghazvininejad. 2023. https://doi.org/10.18653/v1/2023.findings-acl.564 In-context examples selection for machine translation . In Findings of the Association for Computational Linguistics: ACL 2023, pages 8857--8873, Toronto, Canada. Association for Computational Linguistics

  2. [2]

    Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023. https://arxiv.org/abs/2311.16867 The falcon series of open language models . Preprint, arXiv:2...

  3. [3]

    Duarte Miguel Alves, Jos \'e Pombal, Nuno M Guerreiro, Pedro Henrique Martins, Jo \ a o Alves, Amin Farajian, Ben Peters, Ricardo Rei, Patrick Fernandes, Sweta Agrawal, Pierre Colombo, Jos \'e G. C. de Souza, and Andre Martins. 2024. https://openreview.net/forum?id=EHPns3hVkj Tower: An open multilingual large language model for translation-related tasks ....

  4. [4]

    Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H

    Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Z...

  5. [5]

    Akari Asai, Sneha Kudugunta, Xinyan Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, and Hannaneh Hajishirzi. 2024. https://doi.org/10.18653/v1/2024.naacl-long.100 BUFFET : Benchmarking large language models for few-shot cross-lingual transfer . In Proceedings of the 2024 Conference of the North American Chapter of the Associat...

  6. [6]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradb...

  7. [7]

    Christos Christodouloupoulos and Mark Steedman. 2015. A massively parallel corpus: the bible in 100 languages. Language resources and evaluation, 49:375--395

  8. [8]

    Menglong Cui, Jiangcun Du, Shaolin Zhu, and Deyi Xiong. 2024 a . https://aclanthology.org/2024.findings-acl.646 Efficiently exploring large language models for document-level machine translation with in-context learning . In Findings of the Association for Computational Linguistics ACL 2024, pages 10885--10897, Bangkok, Thailand and virtual meeting. Assoc...

Show all 63 references
  1. [9]

    Yiming Cui, Ziqing Yang, and Xin Yao. 2024 b . https://arxiv.org/abs/2304.08177 Efficient and effective text encoding for chinese llama and alpaca . Preprint, arXiv:2304.08177

  2. [10]

    John Dang, Shivalika Singh, Daniel D'souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, Sandra Kublik, Meor Amer, Viraat Aryabumi, Jon Ander Campos, Yi-Chern Tan, Tom Kocmi, Florian Strub, Nathan Grinsztaj...

  3. [11]

    Bosheng Ding, Chengwei Qin, Ruochen Zhao, Tianze Luo, Xinze Li, Guizhen Chen, Wenhan Xia, Junjie Hu, Anh Tuan Luu, and Shafiq Joty. 2024. https://aclanthology.org/2024.findings-acl.97 Data augmentation using LLM s: Data perspectives, learning paradigms and challenges . In Find...

  4. [12]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zha...

  5. [13]

    Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Gim \'e nez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, and Katharina...

  6. [14]

    Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzm \'a n, and Philipp Koehn. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.480 CCA ligned: A massive collection of cross-lingual web-document pairs . In Proceedings of the 2020 Conference on Empirical Methods in Natural Langu...

  7. [15]

    Christian Federmann, Tom Kocmi, and Ying Xin. 2022. https://aclanthology.org/2022.sumeval-1.4 NTREX -128 -- news test references for MT evaluation of 128 languages . In Proceedings of the First Workshop on Scaling Up Multilingual Evaluation, pages 21--24, Online. Association f...

  8. [16]

    Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2022. https://doi.org/10.18653/v1/2022.acl-long.62 Language-agnostic BERT sentence embedding . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long...

  9. [17]

    Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster. 2023. https://doi.org/10.18653/v1/2023.wmt-1.51 Results of ...

  10. [18]

    Pengzhi Gao, Zhongjun He, Hua Wu, and Haifeng Wang. 2024. https://arxiv.org/abs/2401.05861 Towards boosting many-to-many multilingual machine translation with large language models . Preprint, arXiv:2401.05861

  11. [19]

    Pengzhi Gao, Liwen Zhang, Zhongjun He, Hua Wu, and Haifeng Wang. 2023. https://doi.org/10.18653/v1/2023.emnlp-industry.25 Learning multilingual sentence representations with cross-lingual consistency regularization . In Proceedings of the 2023 Conference on Empirical Methods i...

  12. [20]

    Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc ' Aurelio Ranzato, Francisco Guzm \'a n, and Angela Fan. 2022. https://doi.org/10.1162/tacl_a_00474 The F lores-101 evaluation benchmark for low-resource and multilingua...

  13. [21]

    Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and André F

    Nuno M. Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and André F. T. Martins. 2023. https://arxiv.org/abs/2310.10482 xcomet: Transparent machine translation evaluation through fine-grained error detection . Preprint, arXiv:2310.10482

  14. [22]

    Jiaxin Guo, Hao Yang, Zongyao Li, Daimeng Wei, Hengchao Shang, and Xiaoyu Chen. 2024. https://doi.org/10.18653/v1/2024.findings-naacl.42 A novel paradigm boosting translation capabilities of large language models . In Findings of the Association for Computational Linguistics: ...

  15. [23]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...

  16. [24]

    Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Xing Wang, Shuming Shi, and Zhaopeng Tu. 2023. https://arxiv.org/abs/2301.08745 Is chatgpt a good translator? yes with gpt-4 as the engine . Preprint, arXiv:2301.08745

  17. [25]

    Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. https://doi.org/10.18653/v1/2020.acl-main.560 The state and fate of linguistic diversity and inclusion in the NLP world . In Proceedings of the 58th Annual Meeting of the Association for Co...

  18. [26]

    Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. 2016. https://arxiv.org/abs/1612.03651 Fasttext.zip: Compressing text classification models . Preprint, arXiv:1612.03651

  19. [27]

    Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017. https://aclanthology.org/E17-2068 Bag of tricks for efficient text classification . In Proceedings of the 15th Conference of the E uropean Chapter of the Association for Computational Linguistics: Volume ...

  20. [28]

    Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat

    Sneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Christopher A. Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat. 2023. https://arxiv.org/abs/2309.04662 Madlad-400: A multilingual and document-level large audited ...

  21. [29]

    Chong Li, Shaonan Wang, Jiajun Zhang, and Chengqing Zong. 2024 a . https://doi.org/10.18653/v1/2024.naacl-long.445 Improving in-context learning of multilingual generative language models with cross-lingual alignment . In Proceedings of the 2024 Conference of the North America...

  22. [30]

    Jiahuan Li, Hao Zhou, Shujian Huang, Shanbo Cheng, and Jiajun Chen. 2024 b . https://doi.org/10.1162/tacl_a_00655 Eliciting the translation ability of large language models via multilingual finetuning with translation instructions . Transactions of the Association for Computat...

  23. [31]

    Baohao Liao, Christian Herold, Shahram Khadivi, and Christof Monz. 2024. https://arxiv.org/abs/2408.11512 Ikun for wmt24 general mt task: Llms are here for multilingual machine translation . Preprint, arXiv:2408.11512

  24. [32]

    Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O ' Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, M...

  25. [33]

    Yinquan Lu, Wenhao Zhu, Lei Li, Yu Qiao, and Fei Yuan. 2024. https://arxiv.org/abs/2407.05975 Llamax: Scaling linguistic horizons of llm by enhancing translation capabilities beyond 100 languages . Preprint, arXiv:2407.05975

  26. [34]

    Rossi, and Thien Huu Nguyen

    Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. 2024. https://aclanthology.org/2024.lrec-main.377 C ultura X : A cleaned, enormous, and multilingual dataset for large language models in 167 langu...

  27. [35]

    OpenAI. 2023 a . Chatgpt [large language model]. https://chat.openai.com/

  28. [36]

    OpenAI. 2023 b . https://arxiv.org/abs/2303.08774 Gpt-4 technical report . Preprint, arXiv:2303.08774

  29. [37]

    Matt Post. 2018. https://doi.org/10.18653/v1/W18-6319 A call for clarity in reporting BLEU scores . In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186--191, Brussels, Belgium. Association for Computational Linguistics

  30. [38]

    Guerreiro, Jos \ A Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, Jos \'e G

    Ricardo Rei, Nuno M. Guerreiro, Jos \ A Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, Jos \'e G. C. de Souza, and Andr \'e Martins. 2023. https://doi.org/10.18653/v1/2023.wmt-1.73 Scaling up C omet K iwi: Unbabel- IST 2023 submission for the quality estimation shared t...

  31. [39]

    Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.213 COMET : A neural framework for MT evaluation . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685--2702, ...

  32. [40]

    Holger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave, Armand Joulin, and Angela Fan. 2021. https://doi.org/10.18653/v1/2021.acl-long.507 CCM atrix: Mining billions of high-quality parallel sentences on the web . In Proceedings of the 59th Annual Meeting of the Associ...

  33. [41]

    Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan...

  34. [42]

    NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prang...

  35. [43]

    J \"o rg Tiedemann. 2016. https://aclanthology.org/L16-1559 Finding alternative translations in a large corpus of movie subtitle . In Proceedings of the Tenth International Conference on Language Resources and Evaluation ( LREC '16) , pages 3518--3522, Portoro z , Slovenia. Eu...

  36. [44]

    Jörg Tiedemann. 2012. Parallel data, tools and interfaces in opus. In Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12), Istanbul, Turkey. European Language Resources Association (ELRA)

  37. [45]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 a . https://arxiv.org/abs/2302.13971 Lla...

  38. [46]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...

  39. [47]

    Ahmet \"U st \"u n, Viraat Aryabumi, Zheng Yong, Wei-Yin Ko, Daniel D ' souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker. 2024. ht...

  40. [48]

    David Vilar, Markus Freitag, Colin Cherry, Jiaming Luo, Viresh Ratnakar, and George Foster. 2023. https://doi.org/10.18653/v1/2023.acl-long.859 Prompting P a LM for translation: Assessing strategies and performance . In Proceedings of the 61st Annual Meeting of the Association...

  41. [49]

    Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

    Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. https://arxiv.org/abs/2206.07682 Emergen...

  42. [50]

    Genta Indra Winata, Andrea Madotto, Zhaojiang Lin, Rosanne Liu, Jason Yosinski, and Pascale Fung. 2021. https://doi.org/10.18653/v1/2021.mrl-1.1 Language models are few-shot multilingual learners . In Proceedings of the 1st Workshop on Multilingual Representation Learning, pag...

  43. [51]

    BigScience Workshop, :, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanch...

  44. [52]

    Di Wu, Shaomu Tan, Yan Meng, David Stap, and Christof Monz. 2024. https://aclanthology.org/2024.findings-acl.896 How far can 100 samples go? unlocking zero-shot translation with tiny multi-parallel data . In Findings of the Association for Computational Linguistics ACL 2024, p...

  45. [53]

    Zhenyu Wu, Yaoxiang Wang, Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Jingjing Xu, and Yu Qiao. 2023. https://doi.org/10.18653/v1/2023.acl-demo.47 O pen ICL : An open-source framework for in-context learning . In Proceedings of the 61st Annual Meeting of the Association for Comput...

  46. [54]

    Haoran Xu, Young Jin Kim, Amr Sharaf, and Hany Hassan Awadalla. 2024 a . https://openreview.net/forum?id=farT6XXntP A paradigm shift in machine translation: Boosting translation performance of large language models . In The Twelfth International Conference on Learning Represen...

  47. [55]

    Haoran Xu, Kenton Murray, Philipp Koehn, Hieu Hoang, Akiko Eriguchi, and Huda Khayrallah. 2024 b . https://arxiv.org/abs/2410.03115 X-alma: Plug & play modules and adaptive rejection for quality translation at scale . Preprint, arXiv:2410.03115

  48. [56]

    Haoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan, Lingfeng Shen, Benjamin Van Durme, Kenton Murray, and Young Jin Kim. 2024 c . https://openreview.net/forum?id=51iwkioZpn Contrastive preference optimization: Pushing the boundaries of LLM performance in machine translation . In F...

  49. [57]

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze...

  50. [58]

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2...

  51. [59]

    Yaowei Zheng, Richong Zhang, Junhao Zhang, YeYanhan YeYanhan, and Zheyan Luo. 2024. https://aclanthology.org/2024.acl-demos.38 L lama F actory: Unified efficient fine-tuning of 100+ language models . In Proceedings of the 62nd Annual Meeting of the Association for Computationa...

  52. [60]

    Shaolin Zhu, Menglong Cui, and Deyi Xiong. 2024 a . https://aclanthology.org/2024.lrec-main.1444 Towards robust in-context learning for machine translation with large language models . In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Lang...

  53. [61]

    Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li. 2024 b . https://doi.org/10.18653/v1/2024.findings-naacl.176 Multilingual machine translation with large language models: Empirical results and analysis . In Findings of t...

  54. [62]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  55. [63]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.