{"id":"075874e6-0f02-46cc-a158-71f296f11fb1","arxiv_id":"2506.04079","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"EuroLLM-9B is a new open 9B multilingual model covering 35 languages, reported by its authors as the leading open European-made model of its size on multilingual benchmarks and machine translation.","lead":"This report describes EuroLLM-9B, a 9-billion-parameter language model trained on 4 trillion tokens to cover all 24 official EU languages plus 11 others. It also releases EuroFilter, an AI-based data filter, and a synthetic instruction dataset, with evaluations positioning the model as the leading open European-made LLM of its size.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline MT ranking rests on COMET-22 scores, a metric from the same family used to filter the training data; an independent re-scoring is needed before the 3+ point lead can be accepted.","rationale":"The reader already identified evaluation bias as the weakest assumption, but bundled the COMET metric loop together with the Tower v2 translations of MMLU-Pro and MUSR. I find the COMET loop to be the single most load-bearing concern because the headline quantitative claim is expressed entirely in COMET-22 units, and the pre-training pipeline explicitly uses the sibling metric COMETKIWI-22 to filter parallel data. That creates a concrete mechanism by which COMET-22 scores could be inflated for EuroLLM specifically. The Tower v2 issue is real but less central: MMLU-Pro and MUSR are secondary to the MT headline, and their translations also affect all baselines. I looked for stronger internal problems—contamination via WMT21/22 examples in EuroBlocks, missing languages, or errors in the reported tables—but those are either disclosed or not decisive. The report is structurally honest: data sources, hyperparameters, ablations, and artifacts are released, and the models, EuroFilter, and EuroBlocks-Synthetic provide independent value. The proposed re-scoring is feasible because the evaluation code and outputs are public. This does not justify rejection; it does justify keeping the verdict conditional until the COMET-22 dependence is tested.","tokens_in":58915,"tokens_out":4955,"duration_ms":53125,"concrete_test":"Use the released evaluation code and model outputs to re-score the WMT24++ predictions of EuroLLM-9B-IT and Gemma-2-9B-IT with an independent reference-based metric, for example MetricX-23 from Google, and with chrF. If EuroLLM-9B-IT's average margin over Gemma-2-9B-IT falls below 3 points in either direction, or if the Borda ranking changes, then the headline MT claim should be rephrased as specific to COMET-22 and the conditional verdict should be maintained pending this re-scoring.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The report's strongest quantitative claim (§4.3, Table 5) is that EuroLLM-9B-IT beats Gemma-2-9B-IT by more than 3 COMET-22 points on WMT24++ in both translation directions. The load-bearing assumption is that this COMET-22 margin reflects general translation quality. That assumption is insecure. The parallel training data was filtered with COMETKIWI-22 (§2.4.1) at a threshold of 0.7, and the evaluation uses COMET-22 (§4.1); these are sibling learned metrics from the same group, so training on data selected by COMET-family scores can preferentially improve the same metric family without equivalent gains on other quality measures. In addition, WMT24++ was itself created with contributions from the same group, further reducing the independence of the evaluation. The report reports no human evaluation, no alternative learned or lexical metric, and no confidence intervals. Thus the 3+ point margin in Table 5 is not yet shown to be robust to metric choice; it may partly reflect optimization toward the evaluator. The 'leading open European-made LLM of its size' claim inherits this fragility.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents EuroLLM-9B, a 9B-parameter multilingual large language model trained from scratch on roughly 4 trillion tokens in three phases, covering all 24 official EU languages plus 11 additional languages. The report describes the tokenizer, architecture, data collection and filtering (including the new EuroFilter classifier and COMETKIWI-22 filtering at threshold 0.7), the three-phase pre-training schedule, and the post-training procedure that produces EuroLLM-9B-Instruct using the EuroBlocks dataset. Evaluation on external EU20 and Okapi multilingual benchmarks and on WMT24++ machine translation shows that EuroLLM-9B is competitive among European-made models, with the instruction-tuned model claiming a COMET-22 margin of more than three points over Gemma-2-9B-IT in both MT directions. The paper also releases the models, EuroFilter, EuroBlocks-Synthetic, and evaluation code.","tokens_in":59041,"tokens_out":6083,"duration_ms":56957,"significance":"The work is a substantial engineering contribution: it openly releases two 9B models, a multilingual data filter, a synthetic post-training dataset, and evaluation code, which will be useful for future European-language LLM research. The use of externally sourced EU20 and Okapi benchmarks for the general multilingual evaluation, and of WMT24++ with post-edited references for MT, gives the evaluation meaningful breadth. However, the headline machine-translation claim rests entirely on COMET-22, a learned metric from the same group that built the COMETKIWI-22 filter applied to the training data, and no independent metric or human evaluation is reported. If the MT margin is confirmed by an independent evaluation, the paper's claim of a leading open European-made LLM in its size class becomes well supported.","major_comments":[{"comment":"The headline MT claim is scored exclusively with COMET-22, a learned metric developed by the same group (Unbabel/IST) as COMETKIWI-22, which was used to filter the parallel training data at a threshold of 0.7 in §2.4.1. Since the model was trained on data selected by a sibling metric, the 3+ COMET-22 point margin over Gemma-2-9B-IT in Table 5 may partly reflect optimization toward the metric family rather than general translation quality. An independent evaluation (chrF/BLEU, an alternative learned metric such as BLEURT, or a human evaluation on a subset) is needed before the abstract's 'best results' claim can be accepted.","section":"§2.4.1, §4.1, Table 5"},{"comment":"All reported scores are single-point estimates without confidence intervals or significance tests. The per-language tables show that the MT margin over the second-best model varies widely (e.g., 1.07 points for Bulgarian xx→en versus 4.91 for en→xx in Table 7), so the averaged 'more than three points' difference in §4.3 is not established as a stable, statistically reliable effect. A bootstrap or paired significance test should be reported for the WMT24++ COMET-22 averages.","section":"§4.3, Table 5"},{"comment":"The MMLU-Pro and MUSR multilingual evaluations use translations produced by Tower v2, the same translation system used to generate parts of the post-training data, including the translated Cosmopedia documents and translated prompt-answer pairs in §3.1. This creates a potential evaluator-alignment risk for those two benchmarks, analogous to the COMET issue. Since these benchmarks contribute to the Borda counts supporting the 'leading open European-made LLM' claim, the paper should either use an independent translation source for these benchmarks or demonstrate that the ranking is robust to the translation tool.","section":"§3.1, §4.1"}],"minor_comments":[{"comment":"The captions for the Chinese (ZH) pre-trained and post-trained tables (Tables 54 and 55) say 'Arabic benchmarks'; they should say 'Chinese benchmarks'.","section":"Appendix A.2.4"},{"comment":"The phrase 'TOWER V2-supported languages' should use the consistent spelling 'Tower v2' to match the rest of the paper and the cited reference.","section":"§2.4.1"},{"comment":"The benchmark name is written inconsistently as 'MMLU-PRO', 'MMLU-PRO', and 'MMLU-Pro'; please unify the capitalization.","section":"§4.1"},{"comment":"The header of Table 5 uses 'WMT24++ WMT24++' without naming the metric; adding 'COMET-22' to the header would make the table self-contained.","section":"Table 5"},{"comment":"The phase-2 and phase-3 mixture ablations are conducted on reduced-scale settings (80B tokens on the 1.7B model for phase 2), and the paper should state explicitly that these are proxy experiments rather than full-scale validations of the final training setup.","section":"§5.1.1, §5.1.2"}],"recommendation":"major_revision","confidential_remarks":"The central concern is the self-referential evaluation: the paper's strongest quantitative claim uses a metric from the same family as the training-data filter and from the same research group. This is not a reason to reject, but an independent MT evaluation should be requested. The authors may be able to provide chrF/BLEU scores or an external metric quite easily, which would materially strengthen the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the kind of technical report that should exist. It trains and releases a 9B model covering 24 EU languages, and it gives away the useful pieces: EuroFilter, the synthetic EuroBlocks-SFT data, the models, and evaluation code. The genuinely new content is EuroFilter—transferring FineWeb-Edu quality scores to non-English data via translation and an mDeBERTa classifier—plus the RAG-style synthetic instruction generation and the annealing-to-zero third phase. That is real engineering value, and the report is transparent about data sources, thresholds, hyperparameters, and ablations.\n\nThe evaluation is broader than most: EU20 and Okapi are external translated benchmarks, WMT24++ uses post-edited references, and Borda counts are a sensible way to rank across noisy single-run numbers. The authors also dated the \"leading open European-made LLM\" claim to release date, which is the right way to phrase it.\n\nNow the soft spots, in order of weight. The headline MT claim—more than 3 COMET-22 points ahead of Gemma-2-9B-IT—rests entirely on COMET-22, and the parallel training data was filtered with COMETKIWI-22, a sibling metric from the same group. Training on data selected by one learned metric family can preferentially improve that same family. WMT24++ also had contributions from this group, which weakens the independence further. I would not call the result fabricated; the model is also strong on non-MT benchmarks, so it is clearly competitive. But the specific 3-point margin should not be treated as metric-independent until someone re-scores outputs with at least one other metric (chrF, BLEURT, or a different learned metric) or a human sample, and reports variance. That is an addressable flaw, not a fatal one.\n\nSecond, MMLU-Pro and MUSR translations were produced with Tower v2, the same tool used to synthesize parts of the post-training data. That is a bias risk for those two benchmarks, though the direction and size are unknown and the effect is likely smaller than the MT issue. Third, all scores are single-run point estimates; no significance tests. Given the scale, that is common in this literature, but it matters because the paper does ranking claims.\n\nFinally, the coverage claim (24 EU languages) exceeds evaluation coverage for several languages; the EU20 benchmark itself skips Irish, Maltese, and Croatian, and the per-language tables do not include all languages for every benchmark. The authors should state this gap explicitly.\n\nVerdict: worth a serious referee. I would ask for alternative-metric MT rescoring and a clearer limitation paragraph before acceptance, but the artifacts and the transparency justify the review cycle. I would also cite it if I work on multilingual data filtering.","headline":"A solid, artifact-rich systems report whose headline MT lead is plausible but not yet metric-independent; worth peer review with a demand for alternative-metric rescoring.","tokens_in":59802,"tokens_out":2631,"would_cite":true,"duration_ms":28561,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"EuroLLM-9B claims to be the leading open European-made 9B model, with its instruction-tuned variant topping WMT24++ machine translation by more than three COMET points over Gemma-2-9B-IT in both directions.","keywords":["EuroLLM-9B","European languages","multilingual language model","machine translation","data filtering","synthetic instruction data","BPE tokenizer","WMT24++"],"falsifier":"Re-score all WMT24++ outputs with human post-edited judgments or an independent metric family such as chrF or BLEURT, and re-translate the MMLU-Pro and MUSR test sets with a third-party translation provider instead of the authors' own translation model; if EuroLLM-9B-IT's lead over Gemma-2-9B-IT falls below the reported three-plus COMET points, the headline claim would not hold.","tokens_in":58574,"feed_emoji":"🌍","tokens_out":10295,"duration_ms":91719,"temperature":0.7,"pith_summary":"This technical report presents EuroLLM-9B, a large language model trained from scratch on roughly 4 trillion tokens across all 24 official European Union languages and 11 additional languages. Its central claim is that both the base model and its instruction-tuned variant are the leading open European-made LLMs of their size, and specifically that EuroLLM-9B-Instruct achieves the best WMT24++ machine translation results among all compared European and non-European models, in both translation directions, by more than three COMET points over the second-best model, Gemma-2-9B-IT. The paper also introduces and releases two reusable pieces: EuroFilter, an AI-based multilingual data filter, and EuroBlocks-Synthetic, a synthetic post-training dataset that strengthens instruction-following in European languages. A sympathetic reader would take this as evidence that an open, European-language-first model at the 9B scale can compete with or beat general-purpose open models on multilingual benchmarks and translation.","feed_headline":"Open 9B model tops EU-language machine translation","feed_subtitle":"EuroLLM-9B-Instruct beats all compared open models on WMT24++ in both directions, by more than 3 COMET points.","key_machinery":"The argument is carried by an integrated data-and-training pipeline. EuroFilter is a multilingual classifier that transfers educational-quality scores from English web text to other languages, trained on translated FineWeb-edu annotations, so lower-resource European languages get quality-filtered web data. EuroBlocks-Synthetic is a post-training dataset built by prompting a strong model with a monolingual document to create an instruction, then using the same document as context to produce an answer in the target language, which expands instruction coverage to less-resourced EU languages. The tokenizer is a byte-fallback BPE with 128,000 pieces, giving token fertility close to that of 256,000-token models while using half the embedding parameters. The three-phase pre-training schedule starts with 50 percent English, reduces English to 32.5 percent while boosting multilingual data, and ends with code and mathematics raised to 23 percent during an annealing-to-zero learning-rate phase, a configuration the paper ties to late-training reasoning gains. Parallel training data is also filtered with COMETKIWI-22 at a threshold of 0.7, and machine translation is scored with COMET-22, the same metric family.","core_discovery":"The discovery, stated on the paper's own terms, is that a 9B-parameter open model trained from scratch with a carefully staged data pipeline can become the most capable open European-made LLM of its size at the time of release. EuroLLM-9B-Instruct scores 84.19 COMET-22 on en-to-xx and 83.94 on xx-to-en translation over the WMT24++ test set, while Gemma-2-9B-IT scores 80.47 and 80.39, a gap of more than three points in both directions. On multilingual general benchmarks averaged across EU languages, the base model posts the best Borda count among European-made pre-trained models and performs comparably to Gemma-2-9B, while the instruct model repeats that pattern and also outperforms all European models on nearly every language-pair translation direction, with Greek-to-English as the sole exception.","pith_inferences":["The reported three-point translation lead may partly reflect a feedback loop: the same COMET metric family used to filter parallel training data is used to score the test translations, and the paper does not quantify how much of the margin would survive a switch to human judgments or an independent metric.","If EuroFilter generalizes beyond the 35 languages tested, the transfer-by-translation recipe could be applied to other low-resource language families; that extension is implicit in the method but not run here.","The sharp late-training gains from the code-and-math-heavy annealing phase suggest the final-phase data mixture is a promising lever for further scaling, yet the paper tests only three candidate mixtures.","The consistent fourth-place TruthfulQA result across languages hints that multilingual instruction tuning may trade away some truthfulness, but the report does not analyze why that happens."],"forward_implications":["EuroLLM-9B-Instruct, at 9B parameters and about 4 trillion training tokens, is the best open model in its size class for European-language machine translation, ahead of Gemma-2-9B-IT by more than three COMET-22 points in both translation directions.","The base model leads European-made models of similar size on the averaged multilingual benchmarks, with the best Borda count among pre-trained European models and performance comparable to Gemma-2-9B.","The 128k-tokenizer reaches fertility close to 256k-token models while saving half the embedding parameters, which lowers the memory cost of broad language coverage.","The public release of EuroFilter, EuroBlocks-Synthetic, the base model, and the instruction-tuned model lets other teams reproduce or modify the full pipeline instead of treating the 9B model as a black box."],"supporting_citations":[{"why":"Supplies WMT24++, the 55-language test set with post-edited references on which the claimed translation lead is measured.","marker":"Deutsch et al., 2025"},{"why":"Provides COMET-22, the metric used to score all machine translation results.","marker":"Rei et al., 2022a"},{"why":"Provides COMETKIWI-22, the quality-estimation metric used to filter parallel training data at threshold 0.7.","marker":"Rei et al., 2022b"},{"why":"Defines Gemma-2-9B and Gemma-2-9B-IT, the strongest non-European baseline that EuroLLM-9B-IT is claimed to beat by more than three points.","marker":"Gemma 2 Team et al., 2024"},{"why":"Provides the authors' translation model used to create MMLU-Pro and MUSR multilingual evaluations and to generate part of the post-training data.","marker":"Rei et al., 2024"},{"why":"Supplies EU20-Benchmarks, the translated multilingual benchmark set used for European-language evaluation.","marker":"Thellmann et al., 2024"},{"why":"Describes EuroLLM-1.7B, the predecessor whose data-mixture choices and scaling-law analysis guide the 9B training design.","marker":"Martins et al., 2024"},{"why":"Supplies FineWeb-edu and its educational-quality scores, which EuroFilter transfers to non-English languages.","marker":"Lozhkov et al., 2024"}],"fun_headline_variants":["EuroLLM-9B: open 9B model tops EU translation by 3+ points","From-scratch Euro 9B model beats Gemma-2 by 3+ COMET points","Open Euro 9B model outshines all similar-size rivals in EU MT","EuroLLM-9B-Instruct leads all open models in EU-language MT","9B Euro model from scratch beats all open rivals on EU translation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ranking depends on the evaluation being neutral: the metric family used to filter the parallel training data also scores the test translations, and the translation tool used to build parts of the multilingual evaluations also generated post-training data.","fun_headline_variants_meta":{"raw":{"variants":["EuroLLM-9B: open 9B model tops EU translation by 3+ points","From-scratch Euro 9B model beats Gemma-2 by 3+ COMET points","Open Euro 9B model outshines all similar-size rivals in EU MT","EuroLLM-9B-Instruct leads all open models in EU-language MT","9B Euro model from scratch beats all open rivals on EU translation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.002236,"raw_usage":{"total_tokens":8630,"prompt_tokens":915,"completion_tokens":7715,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":7603}},"tokens_in":531,"tokens_out":7715,"duration_ms":48227,"temperature":1.0,"reasoning_tokens":7603,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:48:41.217473+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-score all WMT24++ outputs with human post-edited judgments or an independent metric family such as chrF or BLEURT, and re-translate the MMLU-Pro and MUSR test sets with a third-party translation provider instead of the authors' own translation model; if EuroLLM-9B-IT's lead over Gemma-2-9B-IT falls below the reported three-plus COMET points, the headline claim would not hold.","supporting_citations":[],"review_version":1}