Danish Dynaword packages 4.8B tokens of openly licensed Danish text into a continuously versioned, test-gated corpus that improves language-model perplexity compared with Danish Gigaword.
In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 5843–5862, Miami, Florida, USA
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Dynaword: From One-shot to Continuously Developed Datasets
Danish Dynaword packages 4.8B tokens of openly licensed Danish text into a continuously versioned, test-gated corpus that improves language-model perplexity compared with Danish Gigaword.