REVIEW 3 major objections 5 minor 2 cited by
Towards Best Practices for Open Datasets for LLM Training
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Fully open LLM training data is achievable if metadata improves
desk verdict A solid consensus/roadmap paper on open LLM training datasets, honest about its own limits; the feasibility claim is asserted rather than demonstrated, but the recommendations stand on their own. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the three-tier taxonomy of dataset openness combined with machine-readable metadata and preference signals. The taxonomy separates what is openly licensed from what is merely downloadable from what is replicable, so that claims about openness can be checked against the licensing of both the dataset arrangement and each component. The proposed pipeline treats metadata as the crucial building block: standards such as SPDX license identifiers, content identifiers, and machine-readable opt-out signals are what make downstream governance, removal, and audit possible. The seven guiding principles—competition, reproducibility, harm minimization, diversity, reciprocity, collaboration, and long-term preservation—supply the normative frame that connects these technical practices to the public-good argument.
What would settle it
Try to assemble a complete, machine-readable list of US books published between 1929 and 1964 whose copyrights were not renewed, then check how many have full-text digitizations accessible in bulk. If most of the roughly 480,000 estimated public domain books cannot be obtained in usable form, the paper's core feasibility assumption fails even if transparency remains desirable.
Extended reading notes
Core claim
The paper's central claim is that open, transparent, responsibly governed training datasets for LLMs are both necessary and achievable, but only through deliberate collective infrastructure. It distinguishes three tiers of openness—openly licensed, downloadable/open-access, and replicable—and stresses that a dataset's license does not change the licenses of its constituent parts. The principal obstacles are not legal in principle but practical: license metadata is incomplete, public domain status is hard to verify across jurisdictions, many public domain texts have never been digitized or are locked in gated repositories, and volunteer-driven projects lack the legal and financial buffers to absorb litigation risk. The paper's proposed best practices—preserve metadata, use machine-readable preference signals, document every filtering step, plan for opt-outs, and avoid restrictive terms on public domain data—are offered as the foundation for a shared open-data ecosystem.
Load-bearing premise
That enough openly licensed and public domain text can actually be assembled, digitized, cleaned, and licensed to train models that compete with those trained on copyrighted data, a task the paper's own appendices describe as unfinished work with components still being obtained and the first release deliberately limited to pre-1884 content.
Editorial extensions
If this is right
- Dataset builders who adopt the practices can produce replicable pipelines, so an independent party can rebuild a substantially similar dataset and audit every filtering decision.
- Machine-readable opt-out and preference standards give rights holders a practical way to declare how their content may be used, which the paper connects to the EU AI Act's August 2025 deadline.
- A widely reused, stable open dataset would make model comparisons fairer, because models trained on identical data can be evaluated against one another.
- Public sector and smaller actors gain a competitive route to LLMs without relying on APIs or exclusive licensing deals, reducing market concentration.
- Investments in digitization and metadata reconciliation could unlock hundreds of thousands of public domain books currently unusable, enlarging the open corpus.
Reading between the lines
- Beyond the paper's claims, the first model trained entirely on verified open data at meaningful scale would become a public benchmark for what transparency costs in capability; the paper does not predict its quality, only that it is possible.
- The taxonomy suggests a testable extension: a 'replicable but not openly licensed' tier could be a pragmatic middle path for corpora whose components cannot be relicensed, and the paper's own examples leave room for such graded openness.
- The emphasis on metadata implies that automated license detection tools, not additional legal opinions, may be the highest-leverage investment; one can test this by measuring how much of the open web's CC-licensed content is correctly identifiable with current parsing pipelines.
- If mass opt-outs shrink the open web, the paper's proposal may face a race condition: the metadata infrastructure must be built while the data is still crawlable.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports the outcomes of a June 2024 convening on open datasets for LLM training. It argues that litigation pressure has driven AI labs to reduce training-data transparency, proposes training on openly licensed and public-domain data as a mitigation, and documents technical, legal, and organizational challenges encountered by dataset builders. The paper offers seven guiding principles, a set of best practices for metadata, sourcing, processing, governance, and terms of use, and a series of policy and technology recommendations. Two case studies—EleutherAI's Common Pile and Pleias' Common Corpus/YouTube-Commons—are presented in appendices as emerging efforts. The overarching motivating claim is that no LLMs have yet been trained at a meaningful scale on fully open data, and that assembling such corpora requires cross-disciplinary collaboration and investment.
Significance. If the motivating claim and the feasibility assumption are accepted, the paper provides a useful early consolidation of normative and technical guidance for an emerging community, with concrete examples and a clear taxonomy of openness. Its strengths include explicit definitions, a willingness to acknowledge tensions (e.g., opt-outs vs. reproducibility), a broad set of practical recommendations, and an honest statement that the paper is not a comprehensive analysis. The paper also gives credit to existing community resources and open questions. However, the central claim that no meaningful-scale open-data models exist is imprecise and internally inconsistent, and the case studies do not yet supply quantitative evidence that open-only corpora can meet current training-scale requirements. The best-practices content remains valuable, but the paper needs to clarify what it claims about feasibility and to distinguish established facts from aspirations.
major comments (3)
- [Abstract; §4.2.2; Appendix B] The abstract states that 'at the time of writing, there are no such models (trained at a meaningful scale)' on open access and public domain data, but Appendix B reports that Pleias 'later expanded Common Corpus to include permissibly licenced text, using the updated dataset to train and fine-tune models.' This is at minimum an internal inconsistency, and it is aggravated by the fact that 'meaningful scale' is never defined. Please either define the scope of the claim (e.g., pretrained models above a specified parameter or token budget), cite the specific models that exist and explain why they fall short, or soften the claim to reflect that no models at the scale of contemporary frontier systems are publicly documented. As written, the central motivating claim is not precise enough to support the paper's argument.
- [Appendix A; Appendix B] The feasibility assumption underlying the paper—that sufficient openly licensed and public-domain text can be assembled to train competitive LLMs—is not supported by the manuscript's own evidence. In Appendix A, the only fully processed Common Crawl subset yields 259,728,610 pages comprising 221,715,271,483 words (roughly 0.3 trillion tokens), which is below multi-trillion-token budgets used in contemporary large-model training, and several other components are described as 'still being obtained' or as requiring further validation. Appendix B states that the first Common Corpus release deliberately limited content to pre-1884 works and that YouTube-Commons relies on unvalidated uploader-specified licenses. If the paper intends to claim that open-only corpora can reach the scale required for competitive models, it should provide a quantitative token-availability audit or cite existing audits. If it does not intend to make that claim, the abstract and conclusion should explicitly frame feasibility as an open question rather than a near-term prospect.
- [§2; §4.2.2] The paper's own terminology in §2 distinguishes an 'openly licensed dataset' from a 'downloadable/open-access dataset,' where the latter makes no claim about license compliance. However, §4.2.2 describes Common Pile as 'composed exclusively of public domain and open access data,' which blurs that distinction and could mislead adopters into believing that open-access status is sufficient for legal reuse. Given that the paper's core purpose is to provide legal and practical clarity for dataset builders, this terminological slippage is load-bearing and should be corrected throughout the document.
minor comments (5)
- [§2] Figure 1 is referenced in the text but does not appear in the provided manuscript; ensure that the figure is included in the published version and that its tiers match the definitions in the text.
- [Appendix A] The sentence 'This dataset is similar to the recently released Pleias YouTube Commons dataset' appears in the Common Pile YouTube Transcripts subsection and should be a proper cross-reference with a citation or URL rather than a passing mention.
- [§3.1] The claims about EU and Japanese copyright law in the introduction are stated without citations; a footnote to primary legal sources (e.g., the DSM Directive 2019/790 for the EU) would strengthen the legal framing.
- [§4.1.1] The deadline 'by August 2025' for machine-readable opt-out standards is stated without a citation to the specific EU AI Act provision; please provide a reference or clarify that this is the authors' interpretation.
- [Appendix A] The calculation for the 0.09% false-classification rate assumes independence between errors in titles and registration numbers; the text should note that this is an independence assumption and not a proven upper bound.
Circularity Check
No circularity: the paper's recommendations are normative, its case studies are explicitly work-in-progress, and no fitted quantity or self-citation chain is used to derive its central claims.
full rationale
This paper is a convening report rather than a mathematical derivation, so the usual circularity patterns based on fitted parameters or imported uniqueness theorems do not apply. The central claim in the abstract is that no LLMs trained at a meaningful scale on open access and public domain data exist yet, and that building them requires collaboration, metadata standards, digitization, and a culture of openness. That claim is not derived from a fitted quantity; it is a stated observation and a normative recommendation. The only potentially self-referential evidence is in Appendices A and B, where the authors describe their own projects: EleutherAI's Common Pile and Pleias' Common Corpus and YouTube-Commons. This is transparent rather than circular. Appendix A explicitly says the components 'range from still being obtained to being fully processed' and that license identification is still being validated; Appendix B says Pleias 'deliberately only includes older public domain content in its first release.' The paper never claims these projects have already produced a competitive open-data model, and it concedes in the abstract that no such model exists at meaningful scale. The case studies are therefore candid illustrations of work in progress, not fabricated evidence for a prediction. The recommendations about metadata, sourcing, processing, governance, and release do not reduce by construction to the existence of these datasets; they are independent normative advice. The skeptical concern that the paper lacks a quantitative token-availability audit is a correctness or evidence weakness, not circularity. Under the hard rules, an unsupported feasibility assumption is not the same as a derivation that is equivalent to its inputs by definition. No specific circular step can be exhibited, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Openly licensed and public domain sources can be assembled into corpora large enough to train competitive LLMs at meaningful scale.
- domain assumption Litigation threat is a major driver of reduced training data disclosure.
- domain assumption Openness of training data is necessary for accountability, competition, and innovation.
- domain assumption The EU AI Act and Copyright Directive require machine-readable opt-outs by August 2025.
Cite this review
Pith. "Pith review of Towards Best Practices for Open Datasets for LLM Training." pith.science (2026). https://pith.science/paper/HG5OIUBE
@misc{pith2026250108365,
author = {Pith},
title = {Pith review of: Towards Best Practices for Open Datasets for LLM Training},
year = {2026},
howpublished = {\url{https://pith.science/paper/HG5OIUBE}},
note = {Machine review of arXiv:2501.08365}
}
read the original abstract
Many AI companies are training their large language models (LLMs) on data without the permission of the copyright owners. The permissibility of doing so varies by jurisdiction: in countries like the EU and Japan, this is allowed under certain restrictions, while in the United States, the legal landscape is more ambiguous. Regardless of the legal status, concerns from creative producers have led to several high-profile copyright lawsuits, and the threat of litigation is commonly cited as a reason for the recent trend towards minimizing the information shared about training datasets by both corporate and public interest actors. This trend in limiting data information causes harm by hindering transparency, accountability, and innovation in the broader ecosystem by denying researchers, auditors, and impacted individuals access to the information needed to understand AI models. While this could be mitigated by training language models on open access and public domain data, at the time of writing, there are no such models (trained at a meaningful scale) due to the substantial technical and sociological challenges in assembling the necessary corpus. These challenges include incomplete and unreliable metadata, the cost and complexity of digitizing physical records, and the diverse set of legal and technical skills required to ensure relevance and responsibility in a quickly changing landscape. Building towards a future where AI systems can be trained on openly licensed data that is responsibly curated and governed requires collaboration across legal, technical, and policy domains, along with investments in metadata standards, digitization, and fostering a culture of openness.
Forward citations
Cited by 2 Pith papers
-
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
A new 8TB openly-licensed text corpus trains 7B LLMs that are competitive with Llama 1/2, showing that performant models need not depend on unlicensed web data.
-
Dynaword: From One-shot to Continuously Developed Datasets
Danish Dynaword packages 4.8B tokens of openly licensed Danish text into a continuously versioned, test-gated corpus that improves language-model perplexity compared with Danish Gigaword.
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.