{"id":"c6b4c35f-4c37-4b67-8cf6-9f2c9f060296","arxiv_id":"2501.08365","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A consensus report from Mozilla and EleutherAI's Dataset Convening proposing seven principles and best practices for creating open datasets for LLM training.","lead":"Thirty researchers and practitioners met in June 2024 to draw up principles and practical guidance for building openly licensed datasets for training large language models. The report argues that such datasets are essential for transparency and competition, but face legal, technical, and funding obstacles.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central feasibility claim is unsupported by the paper's own evidence: the case studies are work-in-progress, and a quantitative token-availability audit against competitive model training budgets would determine whether open-only corpora can reach the required scale.","rationale":"The reader identified the feasibility assumption as the weakest link, and I agree. The paper is a consensus document that proposes best practices, but its central claim that open datasets can support competitive LLMs is not supported by empirical evidence in the paper. The case studies from the authors' own organizations are work-in-progress, and the token counts reported fall short of what is typically used to train competitive models today. The paper explicitly acknowledges that no such models exist at meaningful scale, which reinforces that the path is unproven. My concern is not that the recommendations are wrong, but that the load-bearing assumption of corpus sufficiency may fail. A token-availability audit is a concrete, third-party checkable way to estimate the upper bound of open-only data. If the audit shows the total is an order of magnitude below frontier training budgets, the central claim's strongest interpretation is false, though the normative guidance might still be valuable as a partial or long-term agenda. This does not change the reader's conditional verdict; it sharpens the reason for it. I credit the paper for being transparent about the limitations in the appendices, and the legal analysis around licensing metadata is a useful contribution independent of the feasibility question.","tokens_in":28877,"tokens_out":3800,"duration_ms":42154,"concrete_test":"Perform a token-availability audit: sum the deduplicated token counts of all content obtainable from the open repositories cited in the paper — arXiv permissive subset, PubMed Central, Project Gutenberg, Library of Congress Selected Digitized Books, pre-1929 and post-1929 unrenewed books with accessible digitizations, CC-licensed Common Crawl subset, Stack Exchange, US government documents, and similar sources — after applying realistic quality filters and deduplication. Compare the total to the training token budgets of competitive open models (e.g., Llama 3 reported ~15T tokens, Qwen-2.5 ~18T tokens). If the open-only total is less than roughly 10% of those budgets, the central feasibility claim fails under current scaling behavior.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that open and public domain data can support competitive LLMs rests on the feasibility assumption that enough such text can be assembled, cleaned, and deduplicated to train models at a meaningful scale. The paper provides no direct evidence for this assumption; its case studies are explicitly incomplete. In Appendix A, Common Pile components are described as 'still being obtained' or 'in the process of carrying out a detailed study' of license identification accuracy; the only fully processed Common Crawl subset yields 221.7 billion words (roughly 300–400 billion tokens), which is below the trillion-token budgets commonly used for competitive models today. Appendix B notes Pleias deliberately limited the first Common Corpus release to pre-1884 content, and YouTube-Commons transcripts have unvalidated uploader-specified licenses. The paper itself states that 'there are no such models (trained at a meaningful scale),' which is a concession that the path from open data to competitive models has not been demonstrated. If the total volume of deduplicated, quality-filtered open-licensed and public domain text is orders of magnitude smaller than the multi-trillion-token corpora used by frontier models, the recommendations lose their practical purpose and the central claim is false in its strongest form.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports the outcomes of a June 2024 convening on open datasets for LLM training. It argues that litigation pressure has driven AI labs to reduce training-data transparency, proposes training on openly licensed and public-domain data as a mitigation, and documents technical, legal, and organizational challenges encountered by dataset builders. The paper offers seven guiding principles, a set of best practices for metadata, sourcing, processing, governance, and terms of use, and a series of policy and technology recommendations. Two case studies—EleutherAI's Common Pile and Pleias' Common Corpus/YouTube-Commons—are presented in appendices as emerging efforts. The overarching motivating claim is that no LLMs have yet been trained at a meaningful scale on fully open data, and that assembling such corpora requires cross-disciplinary collaboration and investment.","tokens_in":29085,"tokens_out":5109,"duration_ms":56081,"significance":"If the motivating claim and the feasibility assumption are accepted, the paper provides a useful early consolidation of normative and technical guidance for an emerging community, with concrete examples and a clear taxonomy of openness. Its strengths include explicit definitions, a willingness to acknowledge tensions (e.g., opt-outs vs. reproducibility), a broad set of practical recommendations, and an honest statement that the paper is not a comprehensive analysis. The paper also gives credit to existing community resources and open questions. However, the central claim that no meaningful-scale open-data models exist is imprecise and internally inconsistent, and the case studies do not yet supply quantitative evidence that open-only corpora can meet current training-scale requirements. The best-practices content remains valuable, but the paper needs to clarify what it claims about feasibility and to distinguish established facts from aspirations.","major_comments":[{"comment":"The abstract states that 'at the time of writing, there are no such models (trained at a meaningful scale)' on open access and public domain data, but Appendix B reports that Pleias 'later expanded Common Corpus to include permissibly licenced text, using the updated dataset to train and fine-tune models.' This is at minimum an internal inconsistency, and it is aggravated by the fact that 'meaningful scale' is never defined. Please either define the scope of the claim (e.g., pretrained models above a specified parameter or token budget), cite the specific models that exist and explain why they fall short, or soften the claim to reflect that no models at the scale of contemporary frontier systems are publicly documented. As written, the central motivating claim is not precise enough to support the paper's argument.","section":"Abstract; §4.2.2; Appendix B"},{"comment":"The feasibility assumption underlying the paper—that sufficient openly licensed and public-domain text can be assembled to train competitive LLMs—is not supported by the manuscript's own evidence. In Appendix A, the only fully processed Common Crawl subset yields 259,728,610 pages comprising 221,715,271,483 words (roughly 0.3 trillion tokens), which is below multi-trillion-token budgets used in contemporary large-model training, and several other components are described as 'still being obtained' or as requiring further validation. Appendix B states that the first Common Corpus release deliberately limited content to pre-1884 works and that YouTube-Commons relies on unvalidated uploader-specified licenses. If the paper intends to claim that open-only corpora can reach the scale required for competitive models, it should provide a quantitative token-availability audit or cite existing audits. If it does not intend to make that claim, the abstract and conclusion should explicitly frame feasibility as an open question rather than a near-term prospect.","section":"Appendix A; Appendix B"},{"comment":"The paper's own terminology in §2 distinguishes an 'openly licensed dataset' from a 'downloadable/open-access dataset,' where the latter makes no claim about license compliance. However, §4.2.2 describes Common Pile as 'composed exclusively of public domain and open access data,' which blurs that distinction and could mislead adopters into believing that open-access status is sufficient for legal reuse. Given that the paper's core purpose is to provide legal and practical clarity for dataset builders, this terminological slippage is load-bearing and should be corrected throughout the document.","section":"§2; §4.2.2"}],"minor_comments":[{"comment":"Figure 1 is referenced in the text but does not appear in the provided manuscript; ensure that the figure is included in the published version and that its tiers match the definitions in the text.","section":"§2"},{"comment":"The sentence 'This dataset is similar to the recently released Pleias YouTube Commons dataset' appears in the Common Pile YouTube Transcripts subsection and should be a proper cross-reference with a citation or URL rather than a passing mention.","section":"Appendix A"},{"comment":"The claims about EU and Japanese copyright law in the introduction are stated without citations; a footnote to primary legal sources (e.g., the DSM Directive 2019/790 for the EU) would strengthen the legal framing.","section":"§3.1"},{"comment":"The deadline 'by August 2025' for machine-readable opt-out standards is stated without a citation to the specific EU AI Act provision; please provide a reference or clarify that this is the authors' interpretation.","section":"§4.1.1"},{"comment":"The calculation for the 0.09% false-classification rate assumes independence between errors in titles and registration numbers; the text should note that this is an independence assumption and not a proven upper bound.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"This manuscript reads as a workshop proceedings report rather than an empirical research article. Its main value is as a field-building document that consolidates current practice and open questions. The appendices are authored largely by the builders of the datasets being showcased, so a conflict-of-interest or funding statement would be appropriate. The 'no meaningful-scale models' claim is empirically checkable and could become outdated quickly; the authors should date the claim and be explicit about the evidence base. The paper's scope is appropriate for a venue like cs.CY, but it needs the revisions described in the major comments to make its central claims precise."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful consensus document, not a research result. If you want to know where the open-data-for-LLM community currently stands, read it. If you want evidence that open-only corpora can compete with models trained on the open web plus copyrighted books, you won't find it here—and the paper is honest about that.\n\nWhat it does well: the seven principles are sensible and the best-practice lists are concrete enough to act on. The terminology section, especially the distinction between the license on the dataset as an arrangement and the licenses on the constituent works, is genuinely clarifying. It cites prior work (Data Ethics Canvas, Model Openness Framework, AI Act proposals) rather than ignoring it. The case studies are transparent about being work in progress; Appendix A says components are 'still being obtained' and Appendix B says Pleias limited the first Common Corpus release to pre-1884 content. That transparency deserves credit.\n\nThe soft spots are real but not fatal to the paper's stated purpose. The central feasibility claim—that open and public domain data can support competitive LLMs—is asserted, not demonstrated. The stress-test is right: the only fully processed Common Crawl subset in Common Pile is about 221 billion words, below the trillion-token scale common for frontier models, and the paper itself says no models at meaningful scale exist. 'Meaningful scale' is never defined. The sentence in 4.2.2 about Pleias later expanding Common Corpus and training models sits awkwardly with that claim. And the case studies come from the authors' own projects, so there is an element of self-reported feasibility. These issues matter if you read the paper as proof-of-viability. But read it as what it is—an attempt to create shared principles and practices for a community still forming—and the recommendations stand mostly on their own. The paper does not overclaim; the conclusion is appropriately modest.\n\nWho it's for: people working on open datasets, data governance, AI transparency policy, or anyone needing a compact summary of the current open-data agenda. It is not a technical breakthrough and does not need to be. I would send it to peer review as a position/consensus paper; a good referee would push for a precise definition of 'meaningful scale' and perhaps a rough token-availability audit, but the paper deserves referee time. I would bring it to a reading group focused on AI governance, and I would cite it as evidence of emerging consensus.","headline":"A solid consensus/roadmap paper on open LLM training datasets, honest about its own limits; the feasibility claim is asserted rather than demonstrated, but the recommendations stand on their own.","tokens_in":29793,"tokens_out":2538,"would_cite":true,"duration_ms":26750,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fully open LLM training data is achievable if metadata improves","keywords":["open datasets","LLM training data","public domain","copyright","data governance","metadata standards","AI transparency","training data provenance"],"falsifier":"Try to assemble a complete, machine-readable list of US books published between 1929 and 1964 whose copyrights were not renewed, then check how many have full-text digitizations accessible in bulk. If most of the roughly 480,000 estimated public domain books cannot be obtained in usable form, the paper's core feasibility assumption fails even if transparency remains desirable.","tokens_in":28678,"feed_emoji":"📊","tokens_out":5145,"duration_ms":50979,"temperature":0.7,"pith_summary":"This paper argues that the recent trend toward hiding LLM training data harms transparency, accountability, and innovation, and that training on open access and public domain data is the workable alternative. It claims that at the time of writing no model has been trained at meaningful scale on such data, because assembling the corpus faces incomplete and unreliable metadata, costly digitization, jurisdiction-dependent copyright law, and the need to combine legal, technical, and policy skills. Based on a convening of 30 scholars and practitioners and on case studies of three open dataset projects, it proposes seven guiding principles and concrete best practices for sourcing, processing, governing, and releasing open training datasets. If followed, these practices are meant to make fully transparent LLM training data a shared public good rather than a competitive secret.","feed_headline":"Fully open LLM training data is achievable if metadata improves","feed_subtitle":"A convening-backed guide says the path runs through metadata standards, digitization, and governance.","key_machinery":"The load-bearing mechanism is the three-tier taxonomy of dataset openness combined with machine-readable metadata and preference signals. The taxonomy separates what is openly licensed from what is merely downloadable from what is replicable, so that claims about openness can be checked against the licensing of both the dataset arrangement and each component. The proposed pipeline treats metadata as the crucial building block: standards such as SPDX license identifiers, content identifiers, and machine-readable opt-out signals are what make downstream governance, removal, and audit possible. The seven guiding principles—competition, reproducibility, harm minimization, diversity, reciprocity, collaboration, and long-term preservation—supply the normative frame that connects these technical practices to the public-good argument.","core_discovery":"The paper's central claim is that open, transparent, responsibly governed training datasets for LLMs are both necessary and achievable, but only through deliberate collective infrastructure. It distinguishes three tiers of openness—openly licensed, downloadable/open-access, and replicable—and stresses that a dataset's license does not change the licenses of its constituent parts. The principal obstacles are not legal in principle but practical: license metadata is incomplete, public domain status is hard to verify across jurisdictions, many public domain texts have never been digitized or are locked in gated repositories, and volunteer-driven projects lack the legal and financial buffers to absorb litigation risk. The paper's proposed best practices—preserve metadata, use machine-readable preference signals, document every filtering step, plan for opt-outs, and avoid restrictive terms on public domain data—are offered as the foundation for a shared open-data ecosystem.","pith_inferences":["Beyond the paper's claims, the first model trained entirely on verified open data at meaningful scale would become a public benchmark for what transparency costs in capability; the paper does not predict its quality, only that it is possible.","The taxonomy suggests a testable extension: a 'replicable but not openly licensed' tier could be a pragmatic middle path for corpora whose components cannot be relicensed, and the paper's own examples leave room for such graded openness.","The emphasis on metadata implies that automated license detection tools, not additional legal opinions, may be the highest-leverage investment; one can test this by measuring how much of the open web's CC-licensed content is correctly identifiable with current parsing pipelines.","If mass opt-outs shrink the open web, the paper's proposal may face a race condition: the metadata infrastructure must be built while the data is still crawlable."],"forward_implications":["Dataset builders who adopt the practices can produce replicable pipelines, so an independent party can rebuild a substantially similar dataset and audit every filtering decision.","Machine-readable opt-out and preference standards give rights holders a practical way to declare how their content may be used, which the paper connects to the EU AI Act's August 2025 deadline.","A widely reused, stable open dataset would make model comparisons fairer, because models trained on identical data can be evaluated against one another.","Public sector and smaller actors gain a competitive route to LLMs without relying on APIs or exclusive licensing deals, reducing market concentration.","Investments in digitization and metadata reconciliation could unlock hundreds of thousands of public domain books currently unusable, enlarging the open corpus."],"supporting_citations":[],"fun_headline_variants":["Open LLM data needs better metadata, not just open licenses","Unlock open training data: fix metadata and digitize","Open datasets for LLMs: the real fix is metadata hygiene","Achievable open LLM data: just improve metadata and digitize","Open LLM data isn't a legal problem, it's a metadata problem"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"That enough openly licensed and public domain text can actually be assembled, digitized, cleaned, and licensed to train models that compete with those trained on copyrighted data, a task the paper's own appendices describe as unfinished work with components still being obtained and the first release deliberately limited to pre-1884 content.","fun_headline_variants_meta":{"raw":{"variants":["Open LLM data needs better metadata, not just open licenses","Unlock open training data: fix metadata and digitize","Open datasets for LLMs: the real fix is metadata hygiene","Achievable open LLM data: just improve metadata and digitize","Open LLM data isn't a legal problem, it's a metadata problem"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000424,"raw_usage":{"total_tokens":2185,"prompt_tokens":968,"completion_tokens":1217,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":584,"completion_tokens_details":{"reasoning_tokens":1128}},"tokens_in":584,"tokens_out":1217,"duration_ms":9484,"temperature":1.0,"reasoning_tokens":1128,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:29:26.018215+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Try to assemble a complete, machine-readable list of US books published between 1929 and 1964 whose copyrights were not renewed, then check how many have full-text digitizations accessible in bulk. If most of the roughly 480,000 estimated public domain books cannot be obtained in usable form, the paper's core feasibility assumption fails even if transparency remains desirable.","supporting_citations":[],"review_version":1}