Pith. sign in

REVIEW 4 cited by

Poro 34B and the Blessing of Multilinguality

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01856 v3 pith:6S3OIHA3 submitted 2024-04-02 cs.CL

classification cs.CL
keywords languageslanguagemodelmodelstrainingdatamultilingualityblessing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The pretraining of state-of-the-art large language models now requires trillions of words of text, which is orders of magnitude more than available for the vast majority of languages. While including text in more than one language is an obvious way to acquire more pretraining data, multilinguality is often seen as a curse, and most model training efforts continue to focus near-exclusively on individual large languages. We believe that multilinguality can be a blessing: when the lack of training data is a constraint for effectively training larger models for a target language, augmenting the dataset with other languages can offer a way to improve over the capabilities of monolingual models for that language. In this study, we introduce Poro 34B, a 34 billion parameter model trained for 1 trillion tokens of Finnish, English, and programming languages, and demonstrate that a multilingual training approach can produce a model that substantially advances over the capabilities of existing models for Finnish and excels in translation, while also achieving competitive performance in its class for English and programming languages. We release the model parameters, scripts, and data under open licenses at https://huggingface.co/LumiOpen/Poro-34B.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BenCzechMark : A Czech-centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring Mechanism

    cs.CL 2024-12 conditional novelty 7.0 of 10

    BenCzechMark is a new 50-task Czech benchmark with a duel scoring system, a 320GB Czech corpus, and a leaderboard of 50 open-weight models.

  2. Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    JQL trains small multilingual quality scorers from LLM judgments and human annotations, and filtering pretraining data with them improves downstream multilingual model performance over heuristic baselines.

  3. Small Languages, Big Models: A Study of Continual Training on Languages of Norway

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A three-stage continual training recipe (tokenizer change, embedding alignment, full retraining) produces NorMistral-11B, an open Norwegian and Northern Sámi language model that improves on most Norwegian benchmarks a...

  4. OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches

    cs.CL 2025-02 conditional novelty 4.0 of 10

    Open-weight LLMs reduce character error rates in historical English OCR by 7 to 39 percent, but none reach practical quality for historical Finnish.

Pith tools