Pith. sign in

Small Languages, Big Models: A Study of Continual Training on Languages of Norway

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages like Northern S\'ami. To address this issue, we present a novel three-stage continual training approach that substantially improves the downstream performance together with the inference efficiency for the target languages. Based on our findings, we train, evaluate, and openly release a new generative language model for Norwegian Bokm\r{a}l, Nynorsk, and Northern S\'ami with 11.4 billion parameters: NorMistral-11B.

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Multi-label Scandinavian Language Identification (SLIDE)

cs.CL · 2025-02-10 · conditional · novelty 6.0

SLIDE introduces a manually labeled multi-label evaluation set and BERT/FastText models for Scandinavian language identification, using machine translation identity as a silver-labeling signal.

citing papers explorer

Showing 1 of 1 citing paper.

  • Multi-label Scandinavian Language Identification (SLIDE) cs.CL · 2025-02-10 · conditional · none · ref 28 · internal anchor

    SLIDE introduces a manually labeled multi-label evaluation set and BERT/FastText models for Scandinavian language identification, using machine translation identity as a silver-labeling signal.