Pith. sign in

REVIEW 11 cited by

EuroLLM: Multilingual Language Models for Europe

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.16235 v1 pith:54CDTT2D submitted 2024-09-24 cs.CL

classification cs.CL
keywords multilingualdataeurollmeurollm-1languagesllmsmodelsopen-weight
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The quality of open-weight LLMs has seen significant improvement, yet they remain predominantly focused on English. In this paper, we introduce the EuroLLM project, aimed at developing a suite of open-weight multilingual LLMs capable of understanding and generating text in all official European Union languages, as well as several additional relevant languages. We outline the progress made to date, detailing our data collection and filtering process, the development of scaling laws, the creation of our multilingual tokenizer, and the data mix and modeling configurations. Additionally, we release our initial models: EuroLLM-1.7B and EuroLLM-1.7B-Instruct and report their performance on multilingual general benchmarks and machine translation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EsBBQ and CaBBQ: The Spanish and Catalan Bias Benchmarks for Question Answering

    cs.CL 2025-07 conditional novelty 7.0 of 10

    EsBBQ and CaBBQ are new Spanish and Catalan bias benchmarks for multiple-choice QA, built with survey-validated stereotypes from Spain and evaluated on 17 language models.

  2. Comparing and Modeling Argumentation in German Political Communication across Arenas

    cs.CL 2026-07 conditional novelty 6.0 of 10

    In a new 17k-sentence corpus of German COVID-19 political discourse, evidence-based expert justifications are proportionally most frequent in press conferences, not in health committee meetings.

  3. KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    KletterMix is a translated German corpus from English pretraining data that yields measurable gains on German downstream tasks in controlled pretraining experiments.

  4. How Important is `Perfect' English for Machine Translation Prompts?

    cs.CL 2025-07 accept novelty 6.0 of 10

    For LLM machine translation, prompt choice affects output quality more than realistic user errors, with spelling errors hurting most and phrase-level errors often harmless.

  5. Training-free LLM Merging for Multi-task Learning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Hi-Merging merges two task-specialized LLMs by pruning and scaling delta vectors at model and layer level, reporting gains over prior merging and multi-task fine-tuning on English and Chinese MCQA and QA tasks.

  6. NAVER LABS Europe Submission to the Instruction-following Track

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A separately-trained speech projector and text LoRA can be merged with only 1K steps of multimodal fine-tuning to yield a competitive ASR, ST, and spoken QA system.

  7. A Sovereign, Open-Source Foundation Model for German and English

    cs.CL 2026-07 conditional novelty 5.5 of 10

    Soofi S 30B-A3B, a hybrid Mamba-MoE model pretrained on ~27T tokens with deliberately up-weighted German, reports the highest English and German aggregate scores among fully open base models in its comparison while ma...

  8. Assessing the Role of Data Quality in Training Bilingual Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A quality filter trained only on English labels can select better French, German, and Chinese pretraining data, improving bilingual model performance and cutting the monolingual-bilingual gap to about 1%.

  9. Beyond Text Compression: Evaluating Tokenizers Across Scales

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Tokenizer choice matters mostly for multilingual tasks, and 350M-parameter models can predict 2.7B model ranking on translation but not on English benchmarks.

  10. Don't Erase, Inform! Detecting and Contextualizing Harmful Language in Cultural Heritage Collections

    cs.CL 2025-05 conditional novelty 5.0 of 10

    The authors present a multilingual vocabulary and a detection tool using string matching, NER and LLM disambiguation that flags and contextualizes harmful language in cultural heritage metadata, reporting 87 percent p...

  11. Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning

    cs.CL 2025-06 conditional novelty 4.0 of 10

    The IT-IST IWSLT 2025 submission shows that a 1.5B language model with a speech encoder can do reasonable ASR after alignment, but struggles with ST and SQA.

Pith tools