Pith. sign in

REVIEW 2 major objections 6 minor 1 cited by

Dynaword: From One-shot to Continuously Developed Datasets

T0 review · 2 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Danish Dynaword offers an openly licensed, continuously updated Danish corpus that beats Danish Gigaword for language modeling.

desk verdict A genuinely useful Danish dataset and a real community-maintenance framework, but the 'exclusively openly licensed' claim is contradicted by the paper's own Table 4 and needs fixing before the headline claim is credible. read the letter →

arxiv 2508.02271 v2 pith:U2UVQ32K submitted 2025-08-04 cs.CL cs.AI

classification cs.CLcs.AI
keywords dynawordDanishlanguagecorpusopenlylicenseddatadatasetlicensingcontinuousdevelopmentmodelpre-trainingperplexitycommunitycontributions
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that pre-training corpora should be living, community-maintained resources rather than one-shot releases, and it offers a four-principle recipe for building them: traceable and open licensing, reproducible collection, documentation, and extensibility. It then presents Danish Dynaword as a working implementation of that recipe. The corpus contains roughly 4.8 billion tokens from openly licensed Danish sources, over four times the public segment of Danish Gigaword, and it grew through contributions from industry, government, and research. The paper's performance claim is that models trained on Danish Dynaword beat models trained on Danish Gigaword in language modeling perplexity, by 5.9% with continual pre-training and 26% when training from scratch. If this holds, Danish Dynaword is both the largest openly licensed Danish corpus and evidence that community-maintained data can be a practical foundation for language modeling.

What carries the argument

The carrying mechanism is the dynaword approach itself: a four-principle framework (traceable and open licensing, reproducibility, documentation, and extensibility) enforced by per-source datasheets, reproducible collection scripts, versioned releases, lightweight tests for format, quality, and documentation, and a maintenance pipeline for accepting community contributions. The framework turns dataset curation from a one-shot release into a continuous process, and Danish Dynaword is the testbed demonstrating that the process produces a corpus with the claimed size and license properties while improving language modeling.

What would settle it

An independent audit of the license documents for the 'Copyright Law' rows and OpenSubtitles that finds one of them restricts resharing or modification would falsify the 'exclusively openly licensed' claim; alternatively, re-running the reported training regime on the released v1.2.7 data and failing to reproduce the 5.9% and 26% perplexity improvements would falsify the performance claim.

Watch

Extended reading notes

Core claim

The paper's central discovery is that a pre-training corpus built on the dynaword principles can be both large and legally clean: Danish Dynaword v1.2.7 contains roughly 4.8 billion tokens, more than four times the public segment of Danish Gigaword, and is assembled exclusively from openly licensed sources with traceable license documentation. As evidence that the corpus is not only bigger but better, the paper reports that 1-billion-parameter language models continually pre-trained on Danish Dynaword improve perplexity over Danish Gigaword by 5.9% on average, and models trained from scratch improve by 26%; even a size-matched Dynaword subset beats Gigaword by 2.6% in continual pre-training and 18% when training from scratch. The paper also reports gains on seven of nine Danish downstream tasks after continual pre-training. The claim is therefore that a continuously developed, openly licensed dataset can serve as a sustainable foundation for language modeling.

Load-bearing premise

The whole 'exclusively openly licensed' claim rests on the license review being correct for every source in Table 4, including the rows marked 'Copyright Law' and OpenSubtitles; if any of those labels is wrong, the corpus is not exclusively open.

Editorial extensions

If this is right

  • Danish Dynaword is currently the largest openly licensed Danish corpus, at roughly 4.8 billion tokens, more than four times the size of comparable releases.
  • Models trained on it achieve lower perplexity than on Danish Gigaword: 5.9% average improvement in continual pre-training and 26% when training from scratch, with gains persisting even when dataset size is matched.
  • Continual pre-training on Danish Dynaword improves performance on seven of nine Danish downstream tasks, a sign the corpus transfers beyond language modeling.
  • The dynaword framework gives other languages and domains a documented blueprint for building openly licensed, continually updated training data.
  • Versioned releases with a changelog make license and content changes transparent, which lowers the legal and ethical risk of downstream models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the 'exclusively openly licensed' guarantee is only as strong as the legal review behind each source, and because license interpretations vary across jurisdictions, the guarantee should be treated as a claim about documentation and intent rather than absolute legal certainty.
  • Editorial inference: if the corpus grows at the pace shown in the paper's timeline, it may gradually close part of the gap with Common-Crawl-based Danish data; a natural test is whether continued growth yields continued perplexity gains or diminishing returns.
  • Editorial inference: the same pipeline could transfer to other low- and mid-resource languages, but the likely bottleneck will be finding enough contributing institutions and individuals, not the technical tooling.
  • Editorial inference: because the authors mark and exclude evaluation data, downstream users can train models with reduced risk of accidental test-set contamination, a feature that could become more valuable as benchmark overlap grows.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This paper introduces Dynaword, a framework for continuously maintained, openly licensed language corpora, and presents Danish Dynaword, a 4.8B-token Danish corpus assembled from legal, social-media, spoken, web, medical, encyclopedic, literary, news, and dialect sources. The authors report that Danish Dynaword contains over four times as many tokens as comparable Danish releases, is exclusively openly licensed, and has received contributions from industry, government, and research. They evaluate the corpus by continually pre-training and training from scratch a Gemma-3-1B model on Danish Dynaword and Danish Gigaword, reporting average relative perplexity improvements of 5.9% for continual pre-training and 26% for from-scratch training on the full corpus, with additional downstream EuroEval results.

Significance. If the licensing and reproducibility properties are as claimed, this is a valuable community resource and a useful template for future low-resource language corpora. The work is unusually concrete in addressing legal and ethical concerns around web-scale training data: the corpus is versioned, collection scripts are published, datasheets are promised for each source, and the repository includes lightweight tests. The language-modeling comparison is also well designed in several respects: held-out validation sources, external 2025 texts, a size-matched control, and publicly released training code. The perplexity improvements are consistent enough to support the qualitative conclusion that Danish Dynaword is at least as good as Danish Gigaword for language modeling. The main weakness is that the central 'exclusively openly licensed' claim is not fully documented in the manuscript, and the license table itself contains entries that do not appear to be open licenses.

major comments (2)
  1. [Abstract, §2.2, Table 4] The claim that Danish Dynaword is 'exclusively openly licensed' is not established by the evidence in Table 4. The rows for retsinformation.dk (818.25M tokens) and Domsdatabasen.dk (86.35M tokens) list 'Copyright Law' as the license; this is the name of a legal regime, not a license that grants resharing, reuse, or modification, and no Danish-law analysis is provided to show that these official texts can be freely redistributed. These two rows represent roughly 19% of the 4.80B-token total, so the headline claim cannot be read literally. Please replace these labels with a precise legal basis (for example, a specific public-sector-information exception) and cite the relevant legal provisions, or revise the claim. The same concern applies to the labels 'Gutenberg' and 'DanNet 1.0', which are not standard open-license identifiers and are not traceable under the paper's own licensing principle.
  2. [§3, Table 4, Ethical considerations] There is a direct inconsistency between Section 3, which states that 'copyrighted samples from OpenSubtitles (<1M tokens)' were excluded, and Table 4, which lists OpenSubtitles as a 271.60M-token source with license CC-0. The Ethical considerations section further says that a 'notable instance' of copyrighted content occurred with the initial release of OpenSubtitles as part of Danish Gigaword. The paper must clarify exactly which OpenSubtitles content is included in Danish Dynaword, how the CC-0 status of the included subset was verified, and whether all included subtitle text is openly licensed. Without this clarification, the reader cannot determine whether the 'exclusively openly licensed' property holds.
minor comments (6)
  1. [§4.1, Appendix B] The sentence 'which is we eloborate on in the following section' contains a typo, and several other grammatical slips appear (e.g., 'It is by no mean uncommon' and 'it is therefore encourages that model developers exclude evaluation data'). These should be corrected.
  2. [Table 4, Figure 2] The unit 'Llama 3 tokens' is used without specifying the exact tokenizer version or counting script; please provide a pointer to the tokenizer and the script used for token counting.
  3. [Tables 6 and 7] The downstream EuroEval results are mostly within one standard error of the baseline; the statement that Danish Dynaword 'yields gains on 7 out of 9 tasks' should be phrased as directional improvements rather than confirmed gains.
  4. [Table 4] There is a typo in 'OCR'ed Newwspapers from NCC', and the license column would be easier to audit if every license name included a URL or a stable identifier pointing to the full license text.
  5. [§3.2] The mechanism for marking benchmark-containing datasets is described only briefly; please add a pointer to the repository documentation that shows how the marking is implemented and maintained.
  6. [Introduction] The sentence stating that the repository 'includes light-weight tests to ensure data formatting, quality, and documentation' would be more useful if it named the specific tests and the CI system that runs them.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the perplexity comparisons use held-out and contemporary external sources, and the licensing claims are empirical verification claims, not definitional reductions.

full rationale

This is a dataset construction and empirical evaluation paper, not a derivation paper. The central quantitative claims (5.9% relative perplexity improvement for continual pre-training and 26% for training from scratch) come from training Gemma-3-1B models on Danish Dynaword versus Danish Gigaword and evaluating on four datasets (DDT, JVJ, Synnejysk.dk, Nordjyllands News) that the paper explicitly states were excluded from training, plus 2025 DR news articles and Danish Wikipedia articles published after January 1, 2025. That is a held-out, partly external evaluation, not a fitted parameter renamed as a prediction, so it does not reduce to an input by construction. The 4.8B token count and the 'exclusively openly licensed' characterization are empirical descriptions of the released corpus, not conclusions entailed by the definition of Dynaword. The Table 4 license labels for retsinformation.dk and Domsdatabasen.dk ('Copyright Law') and the CC-0 label for OpenSubtitles do conflict with Section 3's statement that copyrighted OpenSubtitles samples were excluded, and the Ethical consideration concedes that seemingly open datasets may contain copyrighted content. These are license-verification and internal-consistency problems that bear on the correctness of the 'openly licensed' claim, but they are not circularity: the claim is not defined in terms of itself, and no derivation is forced by a prior self-citation. Citations to Danish Foundation Models and related work are used for context or source identification and are not load-bearing for the evaluation or for any uniqueness argument. The appended Limitations and Ethical consideration passages explicitly acknowledge the size gap, domain bias, review-quality risk, and residual copyright risk; acknowledging these limitations is consistent with a non-circular empirical contribution.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters were fitted in the paper; the central claims rest on legal and licensing correctness plus data quality assumptions. The axioms listed are domain assumptions, not mathematical axioms.

assumptions (3)
  • domain assumption All included sources are openly licensed or free of copyright, and the documented license labels are correct.
    Section 3's license review is the sole evidence for this; Table 4 includes ambiguous "Copyright Law" labels and OpenSubtitles as CC-0, making this assumption load-bearing and not independently verified.
  • domain assumption Minimal quality checks (Danish, coherent, readable) are sufficient to keep pre-training data useful.
    Section 3 says checks are intentionally minimal and allow downstream filtering; this assumes no harmful low-quality text enters the corpus.
  • domain assumption Held-out evaluation sets and contemporary sources are genuinely disjoint from the training data.
    Section 4.1 states validation sets were excluded, but the matching and deduplication mechanism is not described in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Dynaword: From One-shot to Continuously Developed Datasets." pith.science (2026). https://pith.science/paper/U2UVQ32K

@misc{pith2026250802271,
  author       = {Pith},
  title        = {Pith review of: Dynaword: From One-shot to Continuously Developed Datasets},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/U2UVQ32K}},
  note         = {Machine review of arXiv:2508.02271}
}
read the original abstract

Large-scale datasets are foundational for research and development in natural language processing. However, current approaches face three key challenges: (1) reliance on ambiguously licensed sources restricting use, sharing, and derivative works; (2) static dataset releases that prevent community contributions and diminish longevity; and (3) quality assurance processes restricted to publishing teams rather than leveraging community expertise. To address these limitations, we introduce two contributions: the Dynaword approach and Danish Dynaword. The Dynaword approach is a framework for creating large-scale, open datasets that can be continuously updated through community collaboration. Danish Dynaword is a concrete implementation that validates this approach and demonstrates its potential. Danish Dynaword contains over four times as many tokens as comparable releases, is exclusively openly licensed, and has received multiple contributions across industry and research. The repository includes light-weight tests to ensure data formatting, quality, and documentation, establishing a sustainable framework for ongoing community contributions and dataset evolution.

Figures

Figures reproduced from arXiv: 2508.02271 by the authors.

Figure 1
Figure 1. Overview of the guiding principles for Dyna [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Number of tokens in Danish Dynaword over [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Content by domain. The inner circle shows [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A World in Print: Introducing a Danish-Norwegian corpus of historical newspapers

    cs.DL 2025-09 conditional novelty 6.0 of 10

    A new 474-million-word corpus makes two centuries of Danish and Norwegian newspapers searchable for the first time using neural text recognition.

Reference graph

Works this paper leans on

5 extracted references · 5 canonical work pages · cited by 1 Pith paper

  1. [4]

    Danish Foundation Models

    Danish foundation models. arXiv preprint arXiv:2311.07264. Kenneth Enevoldsen, Márton Kardos, Niklas Muen- nighoff, and Kristoffer Laigaard Nielbo. 2024. The scandinavian embedding benchmarks: Comprehen- sive assessment of multilingual and monolingual text embedding. Neurips. ArXiv: 2406.02396 [cs.CL]. European Union. 2024. Regulation (eu) 2024/1689 of th...

  2. [5]

    The Nordic Pile: A 1.2TB Nordic Dataset for Language Modeling

    The nordic pile: A 1.2 tb nordic dataset for lan- guage modeling. arXiv preprint arXiv:2303.17183. Guilherme Penedo, Hynek Kydlí ˇcek, Loubna Ben al- lal, Anton Lozhkov, Margaret Mitchell, Colin Raf- fel, Leandro V on Werra, and Thomas Wolf. 2024. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. arXiv preprint . ArXiv:2406.17557 ...

  3. [2020]

    Corpora Compared: The Case of the Swedish Gigaword & Wikipedia Corpora

    Corpora compared: The case of the swedish gigaword & wikipedia corpora. arXiv preprint arXiv:2011.03281. Nikolay Arefyev, Mikko Aulamo, Pinzhen Chen, Ona De Gibert Bonet, Barry Haddow, Jind ˇrich Helcl, Bhavitvya Malik, Gema Ramírez-Sánchez, Pavel Stepachev, Jörg Tiedemann, Dušan Variš, and Jaume Zaragoza-Bernabeu. 2024. HPLT‘s first release of data and m...

  4. [2023]

    Towards Best Practices for Open Datasets for LLM Training

    Hplt: High performance language technolo- gies. In Proceedings of the 24th Annual Conference of the European Association for Machine Translation, pages 517–518. Stefan Baack, Stella Biderman, Kasia Odrozek, Aviya Skowron, Ayah Bdeir, Jillian Bommarito, Jennifer Ding, Maximilian Gahntz, Paul Keller, Pierre-Carl Langlais, Greg Lindahl, Sebastian Majstorovic...

  5. [2024]

    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 5843–5862, Miami, Florida, USA

    Pretraining language models using transla- tionese. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 5843–5862, Miami, Florida, USA. Association for Computational Linguistics. Kenneth Enevoldsen, Isaac Chung, Ashwin Mathur, Imene Kerboua, Márton Kardos, David Stap, Jay Gala, Wissam Siblini, Saba Sturua, Sait...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.