Pith. sign in

REVIEW 3 major objections 5 minor 2 cited by

An open 12B model family, adapted by continued pretraining and metric-filtered instruction tuning, reaches translation quality competitive with proprietary systems across 46 languages.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 23:56 UTC pith:DBDG6XTE

load-bearing objection Systematic four-size × five-budget × five-budget scaling study with real open-model gains; the 'matches Google Translate/Gemini' claim is inflated by training on the same XCOMET/COMETKiwi metrics used for WMT24++ evaluation. the 3 major comments →

arxiv 2602.11961 v3 pith:DBDG6XTE submitted 2026-02-12 cs.CL

Scaling Model and Data for Multilingual Machine Translation with Open Large Language Models

classification cs.CL
keywords multilingual machine translationopen large language modelsmodel scalingdata scalingcontinual pretraininginstruction finetuningMiLMMT-46Gemma3
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks how model size and training data scale interact when turning a general-purpose open LLM into a multilingual translation system. Its central claim is that a Gemma3-based family, MiLMMT-46, after continual pretraining on several billion tokens per language and fine-tuning on roughly 100,000–264,000 distilled, quality-filtered instruction pairs, outperforms other open translation models and matches or approaches commercial systems such as Google Translate and Gemini 3 Pro across 46 languages. The scaling study yields a concrete rule: larger models convert pretraining into translation skill more efficiently, so the instruction-tuning budget can be small. If true, these results matter because a self-hostable 12B model could replace commercial translation APIs for many production uses without sacrificing quality.

Core claim

Adapting an open LLM to many-to-many MT works best with a two-stage recipe: continue pretraining on a parallel-first mix of up to about 3 billion tokens per language, then instruction-finetune on a small set of candidate translations generated by stronger closed models and selected by reference-free quality metrics. Across the Gemma3 family (roughly 270M to 12B), continual-pretraining data gives stable gains with diminishing returns, while instruction-tuning data shows a striking economy: for the 12B model, about 100K high-quality sentence pairs already yields strong performance across all 46 languages. On FLORES+ and WMT24++, MiLMMT-46-12B consistently beats leading open models and is compe

What carries the argument

The load-bearing object is the two-stage adaptation pipeline, not a new architecture. The base is the open Gemma3 model family; stage one is continual pretraining with a Parallel-First Monolingual-Second (PFMS) data mix, flooding the model with parallel sentence pairs (about 4.9 billion cleaned pairs across 46 languages) before topping up with monolingual text; stage two is instruction finetuning on roughly 264K pairs distilled from closed LLM outputs and filtered by the reference-free metrics XCOMET and COMETKiwi. The scaling conclusions hinge on these same metrics being both the filter and the evaluation yardstick.

Load-bearing premise

The load-bearing assumption is that the metrics used to filter the instruction data (XCOMET and COMETKiwi) genuinely track human translation quality; if training on those metrics inflates them without improving real translations, the paper's advantage over commercial systems on WMT24++ could be a measurement artifact.

What would settle it

Run a blind human evaluation, or a reference-based metric not used in training such as COMET-22 or chrF, on a random subset of WMT24++ outputs from MiLMMT-46-12B, Google Translate, and Gemini 3 Pro; if human preference or the held-out metric no longer ranks MiLMMT-46 with or above the commercial systems, the central claim collapses.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If MiLMMT-46's results hold, organizations can deploy a 12B open model for high-quality translation across 46 languages without sending text to a commercial API.
  • The data-efficiency result indicates that for larger models, translation ability is mostly acquired during continued pretraining, so instruction data can be kept to about 100K pairs; this lowers the cost of reproducing the system.
  • The consistent scaling curves suggest that allocating more continued-pretraining compute to lower-resource languages yields measurable gains, but with diminishing returns as the token budget grows.
  • The reported zero-shot behavior on Chinese-centric directions (only about 154 instruction pairs in the 100K setting) implies many-to-many translation can emerge from English- and Chinese-centric supervision plus pretraining.
  • The recipe, if correct, provides a reproducible baseline for future open multilingual MT: start from a strong open LLM, continue pretraining on parallel data, and distill from closed models under quality filtering.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same two-stage recipe could be lifted onto other open multilingual LLM families; if the data-efficiency pattern generalizes, a roughly 100K-pair instruction budget should be a sensible starting point for any large model.
  • The near-zero instruction coverage for Chinese-centric directions in the 100K setting suggests true many-to-many capability may be largely a pretraining phenomenon, which can be tested by probing pairwise directions never seen in supervised finetuning.
  • Because the filtering and evaluation metrics are the same test, a human-judgment benchmark is the natural next experiment; the released models make that check directly runnable by anyone.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper adapts the Gemma3 model family (270M–12B) to multilingual machine translation across 46 languages via continual pretraining on OPUS/DCAD data followed by instruction finetuning on roughly 264K distilled, metric-filtered sentence pairs. It reports WMT24++ (XCOMET/COMETKiwi) and FLORES+ (spBLEU/COMET) results, studies model- and data-scaling trends, and compares MiLMMT-46 models against open baselines and proprietary systems. The central claims are that MiLMMT-46-12B outperforms recent open SOTA models and is competitive with Google Translate, Gemini 3 Pro, and GPT-5.

Significance. If the WMT24++ comparisons were clean, this would be a practical contribution: an open, self-hostable 12B model family with a documented training recipe, broad language coverage, and a quantitative data-efficiency result. The paper's strengths include the public release of models and code, the systematic scaling curves across continual pretraining and SFT, and the breadth of the FLORES+/WMT24++ evaluation. However, the headline claim of competitiveness with proprietary systems rests on WMT24++ scores obtained with the same metric family used to filter the SFT training data, and the independent FLORES+ results only partially support that claim. The open-source gains over TranslateGemma, HY-MT, and NLLB appear robust on both metric families, but the comparison to Google Translate and Gemini needs stronger, metric-independent evidence.

major comments (3)
  1. [§5.2 vs. §3.3, Table 24] The SFT dataset is constructed by generating candidates with Gemini 3.0 Pro and GPT-5 and selecting/filtering them using reference-free quality metrics, namely XCOMET and COMETKiwi (§5.2). The WMT24++ evaluation uses exactly these same instruments, XCOMET-XXL and COMETKiwi-XXL (§3.3, Table 24). This creates a circularity: MiLMMT is trained to imitate outputs that score high on the same metrics used to grade it, whereas Google Translate, Gemini, and GPT-5 were not selected against those metrics. The independent FLORES+ metrics do not support the 'competitive with proprietary systems' wording: MiLMMT-12B is behind Google Translate on en→xx spBLEU (41.24 vs. 42.90) and xx→en spBLEU (45.45 vs. 47.42), and behind Gemini 3 Pro on several FLORES+ COMET aggregates. I recommend either removing or substantially softening the proprietary-comparison claim, or adding a human evaluation and/or metrics
  2. [Table 2, FLORES+ vs. WMT24++ rows] There are large, unremarked inconsistencies between the two evaluation instruments, especially for the HY-MT/Hunyuan baselines. For example, HY-MT1.5-1.8B scores 85.30 XCOMET on WMT24++ en→xx, higher than MiLMMT-1B's 77.40, yet on FLORES+ en→xx spBLEU it scores 24.81, far below MiLMMT-1B's 33.60. Similar reversals occur for HY-MT1.5-7B and Hunyuan-MT-7B. Because the two benchmarks produce conflicting rankings, the claim of 'consistent' outperformance is not established. The paper should analyze this metric disagreement (e.g., correlation per language/direction) and should not rely on WMT24++ reference-free scores alone for comparative claims.
  3. [§5.3, all experiments] All results are from a single greedy-decoding pass, with no repeated decoding, variance estimates, or significance tests. Several headline comparisons in Table 24 are within 1–2 XCOMET points (e.g., en→de and en→nl against Google Translate), which is well within the likely noise of a single 1012-sentence or WMT24++ test set. Without confidence intervals or significance testing, 'consistently outperforms' overstates the evidence. At minimum, report bootstrap intervals or per-direction standard errors, and avoid definite comparative language for differences smaller than about 1 metric point.
minor comments (5)
  1. [Abstract/footnote 2] The release URLs in footnote 2 are malformed ('https://huggingface/MiLMMT' and 'https://github/MiLMMT'); the abstract contains the correct HuggingFace collection URL, but the GitHub link is incomplete.
  2. [§5.2] The dataset statistics say English-centric directions account for 94.5% of the data and simplified Chinese-centric directions approximately 7.4%; these numbers sum to more than 100% and should be clarified (e.g., overlapping directions or approximate rounding).
  3. [Table 2 and Table 24] The table formatting has a few spacing errors (e.g., '33.25/ 88.06' in the Tower-Plus-9B row and similar entries in Table 2), which make some cells hard to read.
  4. [Figures 2 and 3] The legend labels for the four Gemma3 sizes are small and partially overlap in the rendered figure; enlarging them or using distinct markers would improve readability.
  5. [§3.1] The description of WMT24++ says 'adopt the English sentences' and excludes low-quality ones for reference-free evaluation; please specify how many sentences were excluded and whether the reported averages are over the remaining set, for reproducibility.

Circularity Check

1 steps flagged

WMT24++ evaluation reuses the exact XCOMET/COMETKiwi metrics used to filter SFT training data, partially manufacturing the edge over proprietary systems; FLORES+ metrics provide independent but weaker support.

specific steps
  1. fitted input called prediction [§5.2 (Supervised Finetuning Data) and §3.3 (Evaluation), with Table 24]
    "For each source sentence, we generate multiple candidate translations using closed-source large language models, including Gemini 3.0 Pro and GPT-5, and select the best candidate based on reference-free quality metrics, namely XCOMET and COMETKiwi. To further ensure data quality, we filter out samples with scores below a predefined threshold. ... For the WMT24++ benchmark, we adopt two reference-free models, XCOMET and COMETKiwi, each of which has 10B parameters and demonstrates high correlation with human judgments."

    The SFT training targets were selected to maximize XCOMET and COMETKiwi scores, and the WMT24++ benchmark is scored with exactly those same instruments. Thus MiLMMT's high WMT24++ numbers—e.g., en→xx 86.68/82.87 vs Google Translate 84.73/81.48 in Table 24—partly verify that the model imitates outputs selected by these metrics, rather than independently confirming translation quality. The proprietary baselines were not filtered or trained against these metrics, so the reported advantage is inflated by selection. This is partial circularity: the model still must learn to imitate the high-scoring candidates, and FLORES+ uses different metrics (spBLEU and reference-based COMET-22), which provide more independent evidence, though with smaller or reversed gaps for the proprietary comparison.

full rationale

The paper has one concrete, load-bearing circular step: the same reference-free metric pair used to filter SFT training candidates (§5.2) is the WMT24++ evaluation instrument (§3.3, Table 24). Because MiLMMT's finetuning targets were chosen by XCOMET/COMETKiwi, part of its WMT24++ advantage over Google Translate, Gemini, and GPT-5 reflects metric-based selection rather than verified translation quality. This affects the 'competitive with proprietary systems' claim, which leans heavily on Table 24. However, the paper is not globally circular. The scaling conclusions (Figures 2--3) and the headline open-source gains over TranslateGemma, HY-MT, Seed-X, and GemmaX2 are also supported by FLORES+ spBLEU and COMET-22, which were not used for SFT filtering. The self-citation to Cui et al. (2025) for PFMS data mixing is an adopted training recipe, not a uniqueness theorem, and the scaling behavior is re-tested empirically in this paper, so it is not load-bearing circularity. Score 6 reflects partial circularity confined to the WMT24++ comparison, not a collapse of the whole derivation.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 0 invented entities

No invented entities (no new particles, forces, or mechanisms). The paper's free parameters are curation/design constants rather than fitted-to-target parameters; the most impactful is the unreported XCOMET/COMETKiwi quality threshold governing the SFT set. The central claims rest on domain assumptions about automatic-metric trustworthiness, corpus cleanliness, and cross-model comparability — with the dual-use of XCOMET/COMETKiwi as both SFT filter and WMT24++ evaluator being the most consequential.

free parameters (2)
  • SFT quality threshold (XCOMET/COMETKiwi cutoff) = not reported
    §5.2 filters distilled candidates 'below a predefined threshold'; the threshold value is unspecified, yet it determines the composition and size (264K) of the instruction dataset and therefore downstream quality.
  • PFMS monolingual supplement ratio (0.1n) = 0.1
    §5.3: 'we additionally include 0.1n billion tokens of monolingual data for each language' — a hand-chosen design constant inherited from Cui et al. (2025), not derived or swept.
axioms (4)
  • domain assumption XCOMET and COMETKiwi (10B reference-free metrics) track human translation quality well enough to serve as both training-data filters and evaluation instruments
    Invoked implicitly in §3.3 (evaluation) and §5.2 (data curation); cites Freitag et al. (2023) for metric–human correlation. Load-bearing for the WMT24++ comparisons, and the same metrics select the training data.
  • domain assumption The cleaned OPUS/DCAD corpora (heuristic filtering, language ID, semantic-similarity filtering) are of sufficient quality for continual pretraining without per-language human validation
    §5.1: all OPUS corpora are concatenated and machine-cleaned; the ~4.9B-pair corpus is never human-audited per language.
  • domain assumption FLORES+ devtest and WMT24++ are representative, and automatic scores are comparable across models with different tokenizers and orthographies
    §5.2 includes FLORES+ dev and NTREX-128 dev in SFT while §3.1 evaluates on FLORES+ devtest; §3.3/§4.2 relies on cross-model comparability that the HY-MT1.5 FLORES/WMT divergence (Table 2) calls into question.
  • domain assumption Gemma3 is a suitable substrate whose scaling behavior transfers to other open LLM families
    All scaling conclusions are drawn from the Gemma3 family alone (§5); the Limitations section concedes that larger models and other base families are unexplored.

pith-pipeline@v1.3.0-alltime-deepseek · 86192 in / 25223 out tokens · 222992 ms · 2026-08-02T23:56:51.866336+00:00 · methodology

0 comments
read the original abstract

Open large language models (LLMs) have demonstrated improving multilingual capabilities in recent years. In this paper, we present a study of open LLMs for multilingual machine translation (MT) across a range of languages, and investigate the effects of model scaling and data scaling when adapting open LLMs to multilingual MT through continual pretraining and instruction finetuning. Based on the Gemma3 model family, we develop MiLMMT-46, which achieves top-tier multilingual translation performance across 46 languages. Extensive experiments show that MiLMMT-46 consistently outperforms recent state-of-the-art (SOTA) models, including Seed-X, HY-MT-1.5, and TranslateGemma, and achieves competitive performance with strong proprietary systems such as Google Translate and Gemini 3 Pro. Models are released at https://huggingface.co/collections/xiaomi-research/milmmt-46. Codes are released at https://github.com/xiaomi-research/gemmax.

Figures

Figures reproduced from arXiv: 2602.11961 by Jian Luan, Jinsong Su, Pengzhi Gao, Wei Liu, Yuzhe Shang.

Figure 1
Figure 1. Figure 1: The tokenizer efficiency of open-source LLMs for each non-English language. The smaller the length [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The translation performance (COMET) of different models trained with different [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: The translation performance (COMET) of different models trained with varying numbers of sentence [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Number of sentence pairs for simplified Chinese-centric and English-centric parallel datasets. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: The translation performance (spBLEU) of different models trained with different [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: The translation performance (spBLEU) of different models trained with varying numbers of sentence [PITH_FULL_IMAGE:figures/full_fig_p010_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AlignAtt4LLM: Fast AlignAtt for Decoder-Only LLMs at IWSLT 2026 Simultaneous Speech Translation Task

    cs.CL 2026-06 unverdicted novelty 6.0

    AlignAtt4LLM adapts AlignAtt to decoder-only LLMs via prompt layout, head selection, and attention replay, outperforming IWSLT 2026 baselines for En-De and En-It at ~2s and <4s latency.

  2. Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation

    cs.AI 2026-05 unverdicted novelty 5.0

    ESRT achieves SOTA many-to-many S2TT across 45 languages on FLEURS via edge-cloud split inference that compresses features 10x and a multi-task curriculum learning strategy for cross-lingual balance.

Reference graph

Works this paper leans on

44 extracted references · 11 linked inside Pith · cited by 2 Pith papers

  1. [1]

    Farhad Akhbardeh, Arkady Arkhangorodsky, Magdalena Biesialska, Ond r ej Bojar, Rajen Chatterjee, Vishrav Chaudhary, Marta R. Costa-jussa, Cristina Espa \ n a-Bonet, Angela Fan, Christian Federmann, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Leonie Harter, Kenneth Heafield, Christopher Homan, Matthias Huck, Kwabena Amponsah-Kaakyire, ...

  2. [2]

    Alves, José Pombal, Nuno M

    Duarte M. Alves, José Pombal, Nuno M. Guerreiro, Pedro H. Martins, João Alves, Amin Farajian, Ben Peters, Ricardo Rei, Patrick Fernandes, Sweta Agrawal, Pierre Colombo, José G. C. de Souza, and André F. T. Martins. 2024. https://arxiv.org/abs/2402.17733 Tower: An open multilingual large language model for translation-related tasks . Preprint, arXiv:2402.17733

  3. [3]

    Pierre Andrews, Mikel Artetxe, Mariano Coria Meglioli, Marta R. Costa-juss \`a , Joe Chuang, David Dale, Mark Duppenthaler, Nathanial Paul Ekberg, Cynthia Gao, Daniel Edward Licht, Jean Maillard, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Eduardo S \'a nchez, Ioannis Tsiamas, Arina Turkatenko, Albert Ventayol-Boada, and Shireen Yates. 2025. ...

  4. [4]

    Lo \"i c Barrault, Magdalena Biesialska, Ond r ej Bojar, Marta R. Costa-juss \`a , Christian Federmann, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Matthias Huck, Eric Joanis, Tom Kocmi, Philipp Koehn, Chi-kiu Lo, Nikola Ljube s i \'c , Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Santanu Pal, and 2 others. 2020. https://aclant...

  5. [5]

    Lo \"i c Barrault, Ond r ej Bojar, Marta R. Costa-juss \`a , Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias M \"u ller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019. https://doi.org/10.18653/v1/W19-5301 Findings of the 2019 conference on machine translation ( WMT 19...

  6. [6]

    Ond r ej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Raphael Rubino, Lucia Specia, and Marco Turchi. 2017. https://doi.org/10.18653/v1/W17-4717 Findings of the 2017 conference on machine translation ( WMT 17) . In...

  7. [7]

    Ond r ej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aur \'e lie N \'e v \'e ol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, and 2 others. 2016. https://doi.org/10.1865...

  8. [8]

    Ond r ej Bojar, Rajen Chatterjee, Christian Federmann, Barry Haddow, Matthias Huck, Chris Hokamp, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Carolina Scarton, Lucia Specia, and Marco Turchi. 2015. https://doi.org/10.18653/v1/W15-3001 Findings of the 2015 workshop on statistical machine translation . In Proceedings of the Ten...

  9. [9]

    Ond r ej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, and Christof Monz. 2018. https://doi.org/10.18653/v1/W18-6401 Findings of the 2018 conference on machine translation ( WMT 18) . In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pages 272--303, Belgium, Brussels. A...

  10. [10]

    Shanbo Cheng, Yu Bao, Qian Cao, Luyang Huang, Liyan Kang, Zhicheng Liu, Yu Lu, Wenhao Zhu, Jingwen Chen, Zhichao Huang, Tao Li, Yifu Li, Huiying Lin, Sitong Liu, Ningxin Peng, Shuaijie She, Lu Xu, Nuo Xu, Sen Yang, and 7 others. 2025. https://arxiv.org/abs/2507.13618 Seed-x: Building strong multilingual translation llm with 7b parameters . Preprint, arXiv...

  11. [11]

    Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, Luke Marris, Sam Petulla, Colin Gaffney, Asaf Aharoni, Nathan Lintz, Tiago Cardal Pais, Henrik Jacobsson, Idan Szpektor, Nan-Jiang Jiang, and 3416 others. 2025. https://arxiv.org/abs/2507.06261 Gemini 2.5: Pus...

  12. [12]

    Menglong Cui, Pengzhi Gao, Wei Liu, Jian Luan, and Bin Wang. 2025. https://doi.org/10.18653/v1/2025.naacl-long.280 Multilingual machine translation with open large language models at practical scale: An empirical study . In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Lan...

  13. [13]

    DeepMind . 2025. Introducing gemini 3. https://blog.google/products/gemini/gemini-3-collection/. Accessed: 2025-12-31

  14. [14]

    Daniel Deutsch, Eleftheria Briakou, Isaac Rayburn Caswell, Mara Finkelstein, Rebecca Galor, Juraj Juraska, Geza Kovacs, Alison Lui, Ricardo Rei, Jason Riesa, Shruti Rijhwani, Parker Riley, Elizabeth Salesky, Firas Trabelsi, Stephanie Winkler, Biao Zhang, and Markus Freitag. 2025. https://doi.org/10.18653/v1/2025.findings-acl.634 WMT 24++: Expanding the la...

  15. [15]

    Christian Federmann, Tom Kocmi, and Ying Xin. 2022. https://doi.org/10.18653/v1/2022.sumeval-1.4 NTREX -128 -- news test references for MT evaluation of 128 languages . In Proceedings of the First Workshop on Scaling Up Multilingual Evaluation, pages 21--24, Online. Association for Computational Linguistics

  16. [16]

    Mara Finkelstein, Isaac Caswell, Tobias Domhan, Jan-Thorsten Peter, Juraj Juraska, Parker Riley, Daniel Deutsch, Cole Dilanni, Colin Cherry, Eleftheria Briakou, Elizabeth Nielsen, Jiaming Luo, Kat Black, Ryan Mullins, Sweta Agrawal, Wenda Xu, Erin Kats, Stephane Jaskiewicz, Markus Freitag, and David Vilar. 2026. https://arxiv.org/abs/2601.09012 Translateg...

  17. [17]

    Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster. 2023. https://doi.org/10.18653/v1/2023.wmt-1.51 Results of WMT 23 metrics shared task: Metrics might be guilty but references are not innoc...

  18. [18]

    Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc ' Aurelio Ranzato, Francisco Guzm \'a n, and Angela Fan. 2022. https://doi.org/10.1162/tacl_a_00474 The F lores-101 evaluation benchmark for low-resource and multilingual machine translation . Transactions of the Association for Computational Lingui...

  19. [19]

    Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and Andr \'e F

    Nuno M. Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and Andr \'e F. T. Martins. 2024. https://doi.org/10.1162/tacl_a_00683 x COMET : Transparent machine translation evaluation through fine-grained error detection . Transactions of the Association for Computational Linguistics, 12:979--995

  20. [20]

    Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. https://doi.org/10.18653/v1/2020.acl-main.560 The state and fate of linguistic diversity and inclusion in the NLP world . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6282--6293, Online. Association for Computational...

  21. [21]

    Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ond r ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Philipp Koehn, Benjamin Marie, Christof Monz, Makoto Morishita, Kenton Murray, Masaaki Nagata, Toshiaki Nakazawa, Martin Popel, and 3 others. 2023. https://doi.org/10.18653/v1/...

  22. [22]

    Tom Kocmi, Rachel Bawden, Ond r ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Nov \'a k, Martin Popel, and Maja Popovi \'c . 2022. https://aclanthology.org/2022.wmt-1.1/ F...

  23. [23]

    NLLB Team , Marta R. Costa-juss \`a , James Cross, Onur C elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, and 20 others. 2024. https://doi.org/10.1038/s41586-024-073...

  24. [24]

    OpenAI . 2025. Introducing gpt-5. https://openai.com/zh-Hans-CN/index/introducing-gpt-5/. Accessed: 2025-12-31

  25. [25]

    Matt Post. 2018. https://doi.org/10.18653/v1/W18-6319 A call for clarity in reporting BLEU scores . In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186--191, Brussels, Belgium. Association for Computational Linguistics

  26. [26]

    Guerreiro, Jos \'e Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, Jos \'e G

    Ricardo Rei, Nuno M. Guerreiro, Jos \'e Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, Jos \'e G. C. de Souza, and Andr \'e F. T. Martins. 2023. https://doi.org/10.18653/v1/2023.wmt-1.73 Scaling up C omet K iwi: Unbabel- IST 2023 submission for the quality estimation shared task . In Proceedings of the Eighth Conference on Machine Translation, page...

  27. [27]

    Guerreiro, José Pombal, João Alves, Pedro Teixeirinha, Amin Farajian, and André F

    Ricardo Rei, Nuno M. Guerreiro, José Pombal, João Alves, Pedro Teixeirinha, Amin Farajian, and André F. T. Martins. 2025. https://arxiv.org/abs/2506.17080 Tower+: Bridging generality and translation specialization in multilingual llms . Preprint, arXiv:2506.17080

  28. [28]

    Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.213 COMET : A neural framework for MT evaluation . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685--2702, Online. Association for Computational Linguistics

  29. [29]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. https://arxiv.org/abs/2402.03300 Deepseekmath: Pushing the limits of mathematical reasoning in open language models . Preprint, arXiv:2402.03300

  30. [30]

    Yingli Shen, Wen Lai, Shuo Wang, Xueren Zhang, Kangyang Luo, Alexander Fraser, and Maosong Sun. 2025. https://arxiv.org/abs/2502.11546 Dcad-2000: A multilingual dataset across 2000+ languages with data cleaning as anomaly detection . Preprint, arXiv:2502.11546

  31. [31]

    Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, Etienne Pot, Ivo Penchev, and 197 others. 2025. https://arxiv.org/abs/2503.19786...

  32. [32]

    Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan, Sammy Jerome, and 179 others. 2024. https://arxiv.org/abs/2408.00118 Gemma 2: ...

  33. [33]

    NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, and 20 others. 2022. https://arxiv.org/abs/2207.04672 No language...

  34. [34]

    J \"o rg Tiedemann. 2012. https://aclanthology.org/L12-1246/ Parallel data, tools and interfaces in OPUS . In Proceedings of the Eighth International Conference on Language Resources and Evaluation ( LREC '12) , pages 2214--2218, Istanbul, Turkey. European Language Resources Association (ELRA)

  35. [35]

    Zhenyu Wu, Yaoxiang Wang, Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Jingjing Xu, and Yu Qiao. 2023. https://doi.org/10.18653/v1/2023.acl-demo.47 O pen ICL : An open-source framework for in-context learning . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 489--498, Toronto, ...

  36. [36]

    Haoran Xu, Young Jin Kim, Amr Sharaf, and Hany Hassan Awadalla. 2024. https://openreview.net/forum?id=farT6XXntP A paradigm shift in machine translation: Boosting translation performance of large language models . In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net

  37. [37]

    Haoran Xu, Kenton Murray, Philipp Koehn, Hieu Hoang, Akiko Eriguchi, and Huda Khayrallah. 2025. https://openreview.net/forum?id=csbf1p8xUq X-ALMA: plug & play modules and adaptive rejection for quality translation at scale . In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025 . OpenReview.net

  38. [38]

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, and 41 others. 2025 a . https://arxiv.org/abs/2505.09388 Qwen3 technical report . Preprint, arXiv:2505.09388

  39. [39]

    An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, and 23 others. 2025 b . https://arxiv.org/abs/2412.15115 Qwen2.5 technical report . Preprint, arXiv:2412.15115

  40. [40]

    Mao Zheng, Zheng Li, Tao Chen, Mingyang Song, and Di Wang. 2025 a . https://arxiv.org/abs/2512.24092 Hy-mt1.5 technical report . Preprint, arXiv:2512.24092

  41. [41]

    Mao Zheng, Zheng Li, Bingxin Qu, Mingyang Song, Yang Du, Mingrui Sun, and Di Wang. 2025 b . https://arxiv.org/abs/2509.05209 Hunyuan-mt technical report . Preprint, arXiv:2509.05209

  42. [42]

    Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, and Zheyan Luo. 2024. https://doi.org/10.18653/v1/2024.acl-demos.38 L lama F actory: Unified efficient fine-tuning of 100+ language models . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 400--410, Bangkok, Thailand. A...

  43. [43]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  44. [44]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...