Danish Dynaword packages 4.8B tokens of openly licensed Danish text into a continuously versioned, test-gated corpus that improves language-model perplexity compared with Danish Gigaword.
Danish Foundation Models
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Large language models, sometimes referred to as foundation models, have transformed multiple fields of research. However, smaller languages risk falling behind due to high training costs and small incentives for large companies to train these models. To combat this, the Danish Foundation Models project seeks to provide and maintain open, well-documented, and high-quality foundation models for the Danish language. This is achieved through broad cooperation with public and private institutions, to ensure high data quality and applicability of the trained models. We present the motivation of the project, the current status, and future perspectives.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Dynaword: From One-shot to Continuously Developed Datasets
Danish Dynaword packages 4.8B tokens of openly licensed Danish text into a continuously versioned, test-gated corpus that improves language-model perplexity compared with Danish Gigaword.