{"id":"20a6aace-f881-4a88-a892-70d402e56688","arxiv_id":"2502.02481","paper_version":4,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new data-mixing recipe (Parallel-First Monolingual-Second) and a 9B model, GemmaX2-28, achieve translation quality competitive with Google Translate and GPT-4 across 28 languages.","lead":"This paper develops a 9-billion-parameter open model, GemmaX2-28, that translates across 28 languages and matches or beats Google Translate and GPT-4 on several benchmarks. The result suggests that with the right data-mixing recipe, practical-scale open models can reach commercial translation quality.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Benchmark contamination is the load-bearing risk: SFT samples FLORES-200 dev and OPUS may contain WMT-24 test data, yet the paper's leakage check is only a trend comparison, not a test-set overlap analysis.","rationale":"The reader's weakest assumption correctly identifies evaluation-benchmark contamination as the central risk. My reading agrees: the paper's own leakage check is descriptive rather than rigorous, and the finetuning data explicitly samples from FLORES-200 dev, the same benchmark family as the devtest evaluation. The OPUS pretraining collection also predates WMT-24 by only a few months, making inclusion of WMT-24 test material plausible and requiring an explicit overlap check. The concern is load-bearing because it directly affects the validity of the headline comparisons to Google Translate and GPT-4-turbo, and because the PFMS recipe was selected using the same FLORES-200 devtest on which final numbers are reported. I do not see an internal inconsistency or a fatal flaw in the experimental design; the released model and the systematic recipe comparisons are valuable. The right remedy is a conditional accept with a required contamination audit and, ideally, an external benchmark evaluation. Therefore the reader's CONDITIONAL verdict remains appropriate, with no change needed.","tokens_in":44541,"tokens_out":5214,"duration_ms":52164,"concrete_test":"Perform a systematic overlap audit: (1) Normalize and exact-match or near-duplicate search (e.g., 13-gram containment or BLEU > 0.9) between all training sources (OPUS, CulturaX, MADLAD-400, TowerBlock, NTREX-128, FLORES-200 dev) and the FLORES-200 devtest and WMT-24 test sets used in Tables 1-2. (2) Remove any matched test items and recompute all headline scores for GemmaX2-28-9B and the main baselines. If the deltas versus TowerInstruct/X-ALMA and Google Translate shrink below metric noise (e.g., less than 0.5 COMET), the SOTA claim should be downgraded. Additionally, re-evaluate on a truly external benchmark not released before August 2024, such as WMT-25 or a fresh held-out test set, to confirm the ranking.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that GemmaX2-28-9B 'consistently outperforms' TowerInstruct and X-ALMA and is 'competitive with Google Translate and GPT-4-turbo' rests on FLORES-200 devtest and WMT-24 scores (Tables 1-2). Section 5.2 states that the supervised finetuning data includes sentence pairs sampled from FLORES-200 dev and NTREX-128, i.e., the same benchmark family used for final evaluation. Section 5.1 says pretraining parallel data comes from 'all Chinese-centric and English-centric parallel datasets from the OPUS collection up to August 2024', with no deduplication against WMT-24 test sets; public test sets can appear in such web-scale collections. The paper's leakage check in Section 4.2 ('We do not observe serious data leakage issues... share a similar trend') is not a contamination test: similar rankings across benchmarks do not rule out memorization or near-duplicate overlap. Moreover, the PFMS recipe was selected by inspecting FLORES-200 devtest performance (Figures 4-5), so the reported numbers are partly the result of benchmark-driven model selection. If the training data overlaps with either evaluation set, the headline comparisons could be inflated and may not reflect genuine translation quality on unseen text.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents an empirical study of multilingual machine translation (MT) with open-source LLMs under 10 billion parameters. The authors benchmark six models (Mistral-7B, Qwen2/2.5-7B, LLaMA3/3.1-8B, Gemma2-9B) across 28 languages on FLORES-200 and WMT-24, finding Gemma2-9B to be the strongest open model. They then propose a Parallel-First Monolingual-Second (PFMS) data mixing strategy for continual pretraining, followed by instruction finetuning on a small high-quality parallel dataset, yielding GemmaX2-28-9B and a 2B variant. The central claim is that GemmaX2-28-9B consistently outperforms existing open SOTA models such as TowerInstruct and X-ALMA, and is competitive with Google Translate and GPT-4-turbo.","tokens_in":44829,"tokens_out":8167,"duration_ms":68328,"significance":"If the reported results hold, the paper demonstrates that a 9B open model can approach production-grade multilingual translation across 28 languages, which is practically significant and would be a valuable resource for the community. The systematic comparison of data mixing ratios (monolingual-only, 2:1, 1:1, 1:2, parallel-only, PFMS) is a useful empirical contribution, and the public release of the models enhances reproducibility and independent verification. However, the strength of these contributions is contingent on addressing the evaluation-integrity concerns raised below, particularly the risk of benchmark contamination and the use of the evaluation set for recipe selection.","major_comments":[{"comment":"The evaluation-integrity check is not a contamination test. Section 4.2 states 'We do not observe serious data leakage issues... share a similar trend,' but comparing rankings across FLORES-200 and WMT-24 does not rule out memorization or near-duplicate overlap. The pretraining parallel data (Section 5.1) is drawn from the entire OPUS collection up to August 2024, which may include the WMT-24 test sets, and the SFT data (Section 5.2) is sampled from FLORES-200 dev and NTREX-128—the same benchmark family used for evaluation. Without an exact or approximate overlap analysis (e.g., n-gram or embedding similarity) between the training corpora and the FLORES-200 devtest and WMT-24 test sets, the headline numbers in Tables 1 and 2 could be inflated. Please add such an analysis, deduplicate against the test sets if overlaps are found, and report results on a truly held-out test set.","section":"§4.2, §5.1, §5.2, Tables 1–2"},{"comment":"The PFMS recipe is selected using the same FLORES-200 devtest split on which the final results are reported. Figures 4 and 5 plot recipe performance on FLORES-200 devtest, and Table 2 then reports the selected model on that same split. This gives the recipe selection access to the test set, and the reported gains of PFMS over the alternatives may be partly due to selection on this benchmark. The paper should use a held-out validation split for recipe selection (e.g., a subset of FLORES-200 dev or a separate multilingual test set) and report final numbers on a test set not used in any design decision.","section":"§5.4, Figures 4–5, Table 2"},{"comment":"All scores are single-run point estimates with no variance or significance information. Many of the central comparisons are small; for example, in Table 2 (23-language row) the WMT-24 en→xx XCOMET difference between GemmaX2-28-9B (82.05) and X-ALMA (81.67) is 0.38 points, and per-direction results in Table 10 show GemmaX2 losing on en→ms (78.29 vs 81.13). Without multiple runs, bootstrap confidence intervals, or significance tests, the claim that GemmaX2 'consistently outperforms' these models is not statistically supported. Please provide uncertainty estimates or temper the claim to 'on average, in these evaluations.'","section":"Tables 1, 2, 10–12"},{"comment":"The Limitations section identifies only the constraint of compute and model scale. It does not mention the contamination risk, the use of the evaluation set for recipe selection, or the lack of significance testing—all of which are primary threats to the validity of the paper's central claim. The paper should either address these issues experimentally or explicitly state them as limitations; as written, the stated limitations omit the most consequential threats to the findings.","section":"Limitations section (after Conclusion)"}],"minor_comments":[{"comment":"The term 'XALMA' should be written as 'X-ALMA' for consistency with the body text and with the cited work.","section":"Abstract"},{"comment":"The model release URL 'https://huggingface/GemmaX2' appears incomplete; it should link to the actual repository (e.g., a huggingface.co address).","section":"Footnote 2"},{"comment":"The caption contains a typo: 'finetuing' should be 'finetuning.'","section":"Table 8 caption"},{"comment":"The phrase 'either from the NTREX-128 and FLORES-200 dev datasets or the OPUS dataset' is awkward; consider 'from either the NTREX-128/FLORES-200 dev datasets or the OPUS dataset.'","section":"§5.2"},{"comment":"The WMT-24 evaluation uses reference-free XCOMET and COMETKiwi; the paper should note that these are learned metrics and may not perfectly reflect human judgments, and ideally it should report at least one reference-based metric for a subset of directions.","section":"§3.3"}],"recommendation":"major_revision","confidential_remarks":"This is a well-executed engineering study with a potentially important practical outcome: a 9B open model competitive with much larger systems. The main risk is evaluation integrity: the training data may overlap with the test benchmarks, and the recipe is selected on the same test split used for final reporting. I believe these issues are fixable with additional analysis (overlap detection, held-out validation, and significance testing), and the authors should be given the opportunity to do so before acceptance. I would not recommend rejection at this stage, but I also would not recommend acceptance without the requested revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper delivers a solid empirical contribution with a useful data-mixing finding and a competitively strong open model, but benchmark hygiene is not as rigorous as the headline claims. I would send it out, with a request for a proper contamination check and a held-out validation split.\n\nWhat's new: the paper systematically benchmarks six open 7-9B LLMs on 28-language translation with two evaluation sets, then ablates continual pretraining data recipes (monolingual-only, 2:1, 1:1, 1:2, parallel-only, and their PFMS). The finding that parallel-first helps high-resource languages while monolingual supplementation helps low-resource ones is a useful practical result. The released GemmaX2-28-2B/9B models are a real artifact, and the tokenizer efficiency comparison is a nice addition. Credit where due: the evaluation infrastructure is thorough, with multiple metrics and strong baselines including Google Translate, GPT-4-turbo, and NLLB.\n\nSoft spots: the largest worry is data leakage. Section 5.2 says the SFT data samples from FLORES-200 dev and NTREX-128, and Section 5.1 uses all OPUS up to August 2024 with no deduplication against WMT-24. The leakage check in Section 4.2 is just a trend comparison, not a contamination analysis. Also, the PFMS recipe was selected by inspecting FLORES-200 devtest (Figures 4-5), so the final reported numbers on that same split are partly a result of benchmark-driven model selection. The authors do report WMT-24 too, and the recipe also wins there, which gives some reassurance, but the lack of a proper held-out validation set and significance tests is a genuine weakness. The 'consistently outperforms' phrasing also oversells: Google Translate still wins many high-resource directions per language, and the average comparisons are not tested for significance.\n\nWho this is for: practitioners building open translation systems and researchers working on data recipes for continual pretraining. The model release alone makes it worth a serious look.\n\nRecommendation: worth a serious referee. I would ask the authors to add a real deduplication analysis against both evaluation sets, report a held-out validation split for recipe selection, and provide per-language significance or variance estimates. The core empirical claim is likely to survive, but it needs these fixes.","headline":"Useful data-recipe study and a strong open 9B translation model, but the leakage check is too casual and the recipe was selected on the same FLORES devtest used for final reporting.","tokens_in":45379,"tokens_out":2947,"would_cite":true,"duration_ms":28717,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a 9B open model, trained parallel-first, matches Google Translate and GPT-4-turbo across 28 languages.","keywords":["multilingual machine translation","open large language models","continual pretraining","parallel-first data mixing","monolingual vs parallel data","low-resource languages","Gemma2","many-to-many translation"],"falsifier":"A decisive test would be to re-evaluate on a freshly created, human-translated test set for the same 28 languages, or to run a contamination search for FLORES-200 devtest sentences in the OPUS-derived training data, and check whether GemmaX2-28-9B still outperforms TowerInstruct and stays within the reported distance of Google Translate and GPT-4-turbo.","tokens_in":44370,"feed_emoji":"🌐","tokens_out":8748,"duration_ms":77451,"temperature":0.7,"pith_summary":"Large language models under ten billion parameters are often treated as too small for production translation, and this paper asks whether that is still true. It benchmarks six open models across 28 languages, identifies Gemma2-9B as the strongest base, and then shows that a particular data-ordering choice—parallel sentence pairs first, monolingual text only as filler—during continued pretraining, followed by a small high-quality finetuning set, yields a 9B model that beats current open translation models and matches Google Translate and GPT-4-turbo. The paper thus claims that a sub-10B open model can reach commercial-grade multilingual translation. If the claim holds, practical translation no longer requires closed APIs or very large proprietary models.","feed_headline":"9B open model matches Google Translate on 28 languages","feed_subtitle":"Parallel-first pretraining pushed a Gemma2 base past open rivals and up to commercial closed models.","key_machinery":"PFMS is the central mechanism. For each of the 28 languages the paper allocates a 2-billion-token continual-pretraining budget: it fills that budget with cleaned English-centric and Chinese-centric parallel sentence pairs from the OPUS collection as far as the available parallel data reaches, and then tops up the remainder with monolingual text from large public corpora. The paper's diagnostic is that high-resource languages already possess the generation ability the model needs, so they mainly need parallel pairs to align representations across languages; low- and mid-resource languages still lack generation capacity, so monolingual text matters for them. PFMS is the compromise that supplies parallel alignment where possible and monolingual mass where needed, and the experiments compare it against monolingual-only, 2:1, 1:1, 1:2, and parallel-only mixtures.","core_discovery":"On the paper's own terms, the central discovery is that the ordering and ratio of parallel versus monolingual data in continual pretraining determines multilingual translation quality more than either data type alone. The authors' GemmaX2-28-9B—Gemma2-9B continually pretrained with the PFMS mixture and then instruction-finetuned on about 196,000 hand-filtered translation pairs—consistently outperforms open baselines such as TowerInstruct, X-ALMA, Aya, and LLaMAX on overlapping language directions, and its averaged scores on FLORES-200 and WMT-24 are comparable to Google Translate and GPT-4-turbo. The same recipe improves a 2B variant, which the paper takes as evidence that the strategy transfers across model sizes.","pith_inferences":["An editor's inference: PFMS is a candidate general recipe for any multilingual backbone, since the paper shows the same ordering advantage at 2B and 9B scale; re-running the recipe on other base models would test that generality.","An editor's inference: the per-language budget design suggests that for low-resource languages, monolingual volume is the binding constraint, so adding monolingual data may yield larger gains than collecting more parallel data for those languages.","An editor's inference: a 9B model at this quality implies that offline, on-premise multilingual translation is feasible for privacy-sensitive or cost-constrained settings, a deployment path the paper does not discuss."],"forward_implications":["A sub-10B open model can match closed commercial systems on average across 28 languages, making high-quality translation available without closed APIs.","Parallel data should be prioritized over monolingual data during continual pretraining for multilingual MT, reversing the emphasis of earlier monolingual-only recipes.","The PFMS advantage appears at both 9B and 2B scale, so the recipe is not tied to one model size.","Low- and mid-resource languages gain most from the PFMS mixture, suggesting capacity, not alignment, is their main bottleneck."],"supporting_citations":[{"why":"Provides the Gemma2-9B backbone whose multilingual capacity the paper measures and then continual-pretrains.","marker":"Team et al. 2024"},{"why":"Supplies TowerInstruct, the strongest open translation-specific baseline that GemmaX2 must beat, and the 2:1 monolingual-to-parallel mixture recipe the paper tests.","marker":"Alves et al. 2024"},{"why":"Supplies X-ALMA, the open multilingual MT baseline that GemmaX2-28-9B outperforms on overlapping directions.","marker":"Xu et al. 2024b"},{"why":"Provides the FLORES-200 benchmark used for evaluation and dev-set exemplars, plus the NLLB-54.5B supervised baseline.","marker":"Team et al. 2022"},{"why":"Provides the OPUS parallel corpora that, after cleaning, yield about 3.4 billion English- and Chinese-centric sentence pairs used in PFMS.","marker":"Tiedemann, 2012"},{"why":"Establishes the monolingual-continual-pretraining-then-small-finetuning paradigm and the claim that high-quality finetuning data can dramatically boost translation, which the paper adopts.","marker":"Xu et al. 2024a"},{"why":"Shows that parallel data matters during continual pretraining, the finding PFMS extends by varying the monolingual-to-parallel order and ratio.","marker":"Guo et al. 2024"},{"why":"Supplies the CulturaX monolingual corpus used to fill the monolingual portion of the pretraining data.","marker":"Nguyen et al. 2024"}],"fun_headline_variants":["9B open LLM matches Google Translate across 28 languages","PFMS data order lifts small open LLM to top MT","Gemma2-based 9B beats open rivals, ties closed giants","Parallel-first pretraining makes 9B model translation leader"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim assumes the evaluation benchmarks—FLORES-200 devtest and WMT-24—are uncontaminated and valid proxies for quality, yet the finetuning data is partly drawn from FLORES-200 dev and the paper's leakage check is descriptive rather than a rigorous contamination test.","fun_headline_variants_meta":{"raw":{"variants":["9B open LLM matches Google Translate across 28 languages","PFMS data order lifts small open LLM to top MT","Gemma2-based 9B beats open rivals, ties closed giants","Parallel-first pretraining makes 9B model translation leader"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000996,"raw_usage":{"total_tokens":4175,"prompt_tokens":859,"completion_tokens":3316,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":475,"completion_tokens_details":{"reasoning_tokens":3256}},"tokens_in":475,"tokens_out":3316,"duration_ms":24824,"temperature":1.0,"reasoning_tokens":3256,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T11:58:48.456709+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test would be to re-evaluate on a freshly created, human-translated test set for the same 28 languages, or to run a contamination search for FLORES-200 devtest sentences in the OPUS-derived training data, and check whether GemmaX2-28-9B still outperforms TowerInstruct and stays within the reported distance of Google Translate and GPT-4-turbo.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the OPUS parallel corpora that, after cleaning, yield about 3.4 billion English- and Chinese-centric sentence pairs used in PFMS."}],"review_version":1}