{"id":"5846e081-51cd-4eba-a166-706ad4cace5f","arxiv_id":"2412.17847","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with no relative diversity gains since 2013.","lead":"Researchers manually cataloged nearly 4,000 public AI training datasets for text, speech, and video, tracking their sources, licenses, languages, and creator countries from 1990 to 2024. The audit shows most AI training data now comes from web crawls and YouTube, that a large share of that content carries non-commercial restrictions even when dataset licenses look permissive, and that geographic and language diversity has not improved in over a decade.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Manual source-restriction coding is unvalidated; headline percentages hinge on it, and the abstract's 'over 80%' already fails for speech (78.6% in Table 3), so the central claim is not yet airtight.","rationale":"I agree with the reader's weakest assumption that the manual source-term classification is unvalidated and load-bearing; the headline percentages in Table 3 derive directly from that coding. However, I would elevate the internal inconsistency between the abstract's 'over 80%' and Table 3's 78.6% for speech as a concrete, already-demonstrated flaw that needs correcting, rather than a purely hypothetical annotator bias. Both point to the same conclusion: the paper's central numerical claim has not been independently verified and, for speech, is stated incorrectly. The proposed inter-annotator reliability study with sensitivity analysis would settle whether the concern lands by quantifying whether the percentages shift enough to change the qualitative conclusion. If the numbers survive such a test, the claim could be upheld (with the speech figure corrected to 78.6%). If they do not, the claim would need substantial revision. I recommend keeping the reader's CONDITIONAL verdict because the issues are fixable through validation and a revised abstract, not fundamental to the audit's value.","tokens_in":64350,"tokens_out":6242,"duration_ms":50105,"concrete_test":"Have two independent annotators, blind to the original labels, recode the source-terms category (Restricted/Unspecified/Unrestricted) for a random sample of 100 dataset sources per modality that overlaps with the original 10% of records; compute Cohen's kappa and report the disagreement distribution. Then recompute the headline restricted-content percentages from Table 3 under two extreme scenarios: (a) treating all Unspecified as Restricted, and (b) treating all Unspecified as Unrestricted. If kappa falls below 0.6, or if the speech percentage drops below 80% under scenario (b), the central claim is not robust to annotation variation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that over 80% of source content in widely used text, speech, and video datasets carries non-commercial restrictions—depends entirely on the manual classification of each dataset's sources into Unrestricted, Unspecified, Source Closed, or Model Closed (Table 2). The paper reports no inter-annotator agreement, no gold-standard validation, and no robustness checks for these labels. A systematic annotator bias (for example, conservatively coding ambiguous terms as Restricted rather than Unspecified) would directly change the headline percentages in Table 3: 99.8% for text, 78.6% for speech, and 99.1% for video. The abstract's 'over 80%' claim is already internally inconsistent with the paper's own Table 3, where only 78.6% of speech content is restricted. Because the audit code and full annotation data are promised but no link or commit hash is provided, independent verification is currently impossible. The paper itself states that 'all annotations and analysis code will be made publicly available on release,' which is a limitation, not evidence. Without a reliability study, the quantitative claim is not robust enough to support the strong conclusion that the AI data commons is far more legally constrained than dataset licenses indicate.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a large-scale manual audit of 3,916 public text, speech, and video datasets released between 1990 and 2024, tracing sourcing trends, license and source-term restrictions, and geographic and linguistic representation. The authors report three headline findings: (1) multimodal training data increasingly comes from web-crawled, social media, and synthetic sources; (2) although fewer than one-third of datasets carry restrictive licenses, over 80% of the underlying source content carries non-commercial restrictions; and (3) absolute language and country counts have risen since 2013 but relative measures of geographic and multilingual inequality have not significantly improved. The paper includes extensive appendix tables, a detailed taxonomy, and a promised public release of the audit data and code. The central quantitative claims are descriptive measurements rather than fitted-model outputs, but the main restriction percentages depend entirely on manually assigned source-level labels whose reliability is not reported.","tokens_in":64575,"tokens_out":4802,"duration_ms":40420,"significance":"If the headline measurement is correct, the paper provides an important and policy-relevant result: the effective legal constraints on widely used training data are much stronger than dataset licenses alone suggest. The work is the first multimodal provenance audit at this scale, and the appendix tables and attribution cards are a substantial community resource. The paper is appropriately cautious in describing licenses and terms as signals rather than enforceable legal determinations, and it avoids overfitting by presenting descriptive statistics rather than model-based claims. The empirical contribution would be strengthened by the promised release of annotations and code, which is an explicit strength if delivered. However, the central 'over 80%' claim is not yet fully supported: it is internally inconsistent with the paper's own Table 3 for speech, and the underlying manual classifications have no reported reliability validation.","major_comments":[{"comment":"The central quantitative claims—99.8% of text tokens, 78.6% of speech hours, and 99.1% of video hours carrying source restrictions—are direct tabulations of the manually assigned labels in Table 2, yet the paper reports no inter-annotator agreement, no gold-standard validation, and no sensitivity analysis for ambiguous cases. A systematic bias in coding (for example, treating 'no information found' as Restricted rather than Unspecified) would directly change the headline percentages and the abstract's 'over 80%' conclusion. The manuscript states that 'All annotations and analysis code will be made publicly available on release,' but no link or commit hash is given, so independent verification is not currently possible. Please add a reliability study, such as dual annotation with agreement statistics on a representative subsample or an audit of ambiguous cases, and provide the data and code artifact at submission time.","section":"Section 2, 'Annotation Features & Methodology'; Tables 3 and 4"},{"comment":"The abstract claims that 'over 80% of the source content in widely-used text, speech, and video datasets carry non-commercial restrictions,' and the introduction states 'over 80% of content from each modality.' Table 3 reports 78.6% for speech hours, and Section 3.2 itself says '78%' for speech. The claim is therefore internally inconsistent with the paper's own appendix data. The wording should be corrected to 'roughly 80%,' 'over 78%,' or the computation should be adjusted so that the abstract matches the reported tables.","section":"Abstract and Section 3.2 (Table 3)"},{"comment":"The claim that geographical and linguistic representation 'has not significantly improved' rests on Gini coefficients with 95% confidence intervals and statements about significance at the p = 0.05 level, but the method for computing these intervals and tests is not described anywhere in the paper or appendix, and no code is provided. Without knowing whether the intervals come from a bootstrap, jackknife, or analytic approach, the 'not significant' conclusion is not verifiable. Please specify the estimator and test procedure, or downgrade the claim to a descriptive statement about observed trends.","section":"Section 3.3, Figure 4"}],"minor_comments":[{"comment":"Section 2.1 reports 3,713 text datasets from 108 collections, while Table 1 reports 3,717 text datasets; these counts should be reconciled.","section":"Section 2.1 vs Table 1"},{"comment":"References [11] and [12] are the same paper by Buolamwini and Gebru, and the Common Voice citation appears twice, once as [13] and again as [20]; these duplicates should be merged.","section":"References"},{"comment":"The caption says the table is a breakdown 'across datasets' but the cells are percentages of total tokens or hours; clarify that the units are shares of content, not dataset counts, to avoid confusion with Table 4.","section":"Table 3 caption"},{"comment":"The phrase 'bare restrictions' appears in the Figure 2 caption and in the body text; this should be 'bear restrictions'.","section":"Figure 2 caption and Section 3.2"},{"comment":"The phrase 'undocumented restrictions in the dataset's sources' is imprecise: the source restrictions are often documented in terms of service, but not in the dataset license; consider wording such as 'source-level restrictions not reflected in dataset licenses.'","section":"Introduction, finding 2"},{"comment":"The comparison of average synthetic and natural dataset lengths (1,756 vs 1,065 tokens) would benefit from sample sizes and a measure of dispersion to support the word 'notably.'","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"I see no circularity problem: the audit is descriptive and the methodology is inherited transparently from prior work. The main gatekeeping issue is validation of the manual source-restriction labels and the internal inconsistency between the abstract and Table 3 for speech. If the authors add a reliability subsample, provide the data/code artifact, and correct the abstract's overstatement, I would support publication; without those changes the strongest claim is not yet robust enough for the standard abstract-level claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful part: this is the first audit I know that traces provenance, license terms, and geographic/linguistic representation across text, speech, and video from the same methodology. Roughly 4,000 datasets, 1990–2024, with per-modality source and restriction tables and longitudinal Gini coefficients. That is a real reference contribution, and the appendix tables are extensive. If you work on data governance or dataset documentation, you will want this on hand.\n\nThe paper is honest about its assumptions—scope choice, creator-country mapping, the fact that 'relative representation hasn't improved' is measured by Gini, not absolute counts. Those parts hold up.\n\nThe soft spots, in order. First, the abstract says 'over 80% of the source content' is restricted, but their own Table 3 gives speech 78.6%. The body says 78%. That is a mismatch, a public-facing overstatement, and it should be fixed.\n\nSecond, the central restriction numbers rest on manual coding of each source as Unrestricted, Unspecified, Source Closed, or Model Closed. No inter-annotator agreement, no gold-standard check, no robustness test is reported. For a headline like 99.8% of text tokens restricted, systematic annotator bias would move the number materially. This is the main reason I would not call the quantitative claim airtight. It is not fatal—the taxonomy is inherited from their prior published audit, and the direction of bias is not obvious—but a reliability study or a sensitivity analysis is needed before I'd treat 99.8% as fact rather than estimate.\n\nThird, Figure 4 shows 95% CIs for Gini but does not say how they were computed. And the audit data and code is promised 'on release' but no link or commit hash is provided. These are minor at this stage, but for a descriptive audit whose value is the data, the release should be part of the submission.\n\nThe citation pattern is fine; self-citations are to their own earlier audits and the taxonomy, which is legitimate.\n\nBottom line: the central argument—that the public data commons is much more legally constrained than dataset licenses suggest—survives reading, with the speech caveat. The paper deserves a serious referee. I'd send it out; I'd ask for the reliability check and a corrected abstract.\n\nWho is this for? Regulators, dataset builders, responsible-AI researchers, and anyone doing licensing analysis. I would bring it to a reading group and would cite it in my own work.","headline":"First modality-spanning provenance audit; useful reference, but 'over 80%' overstates speech (78.6%) and the manual source coding needs validation before the headline numbers are taken as fact.","tokens_in":65332,"tokens_out":2752,"would_cite":true,"duration_ms":24698,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A multimodal audit of nearly 4,000 datasets argues that dataset licenses drastically understate the legal constraints on AI training data, because over 80% of the content in widely used text, speech, and video datasets carries…","keywords":["data provenance","dataset licensing","training data audit","multimodal datasets","source restrictions","non-commercial restrictions","geographical representation","linguistic representation"],"falsifier":"Take a random sample of a few hundred datasets from the released audit, have a second independent team re-annotate source restrictions from the same lineage, and compute agreement; low agreement on the Source Closed and Model Closed categories would undercut the 80%-plus mismatch claim. A separate check: recompute the 99.8% text, 78% speech, and 99% video figures with the single largest collection per modality removed (for example the roughly 370k-hour YouTube-sourced speech corpus), since one re-coded giant could shift the ecosystem totals.","tokens_in":1640,"feed_emoji":"⚖️","tokens_out":3212,"duration_ms":65186,"temperature":0.7,"pith_summary":"This paper argues that the public record of how AI training data may be used is systematically misleading. Auditing nearly 4,000 text, speech, and video datasets released between 1990 and 2024, the authors find that while fewer than a third of datasets carry restrictive licenses, the underlying source content is far more constrained: 99.8% of text tokens, 78% of speech hours, and 99% of video hours carry non-commercial restrictions that come from the sources — websites, social media platforms, or generative models — the data was drawn from. Because those source-level restrictions are usually dropped when datasets are re-packaged, a permissive dataset license does not tell practitioners what they may actually use. A sympathetic reader would care because the finding puts a measurable legal shadow over the common practice of assembling large training corpora from web crawls, YouTube, and synthetic model outputs, and the released audit lets individual developers trace a dataset's chain of provenance for themselves.","feed_headline":"Over 80% of AI training content carries hidden limits","feed_subtitle":"A 3,916-dataset audit traces restrictions to their sources — and finds licenses tell only part of the story.","key_machinery":"The load-bearing instrument is a manual provenance-tracing protocol that codes each dataset on two axes at once. The dataset license is categorized as Commercial, Non-commercial/Academic, or Unspecified, while each underlying source — every website, platform, or model the content came from — is coded as Unrestricted, Unspecified, Source Closed, or Model Closed, following a defined taxonomy that covers terms of service, acceptable-use policies, and anti-crawling clauses. A dataset's overall terms status is set to the strictest of its sources, and the crosstab of license against terms is then reported both by dataset count and weighted by tokens or hours, which is what produces the headline mismatch figures.","core_discovery":"The paper's central claim is a mismatch between two layers of restriction: dataset licenses and source terms. Counting datasets, only 25% of text, 33% of speech, and 32% of video datasets are licensed non-commercially; weighted by content, those figures are 21%, 26%, and 33%. But when the authors trace each dataset back to its original sources and classify the sources' licenses or terms of service, 99.8% of text tokens, 78% of speech hours, and 99% of video hours carry some non-commercial restriction at the source level, and the license-versus-terms mismatch covers 79% of text tokens, 55% of speech hours, and 65% of video hours. The paper further claims that since 2019 multimodal training has overwhelmingly moved to web-crawled, synthetic, and social-media sources (YouTube alone supplies roughly 71% of video data and 69% of speech data), and that although the absolute number of languages and countries represented keeps rising, the relative dispersion measured by Gini coefficients has not significantly improved since 2013 — Western concentration persists despite diversification at the margins.","pith_inferences":["One implication the paper leaves implicit: if source terms bind, then a 'clean' multimodal training run under current terms would have to be assembled almost entirely from a short list of explicitly permissive sources, and the paper's released tables make that set enumerable.","A testable extension: apply the same two-axis coding to datasets released after April 2024 to see whether the license-versus-source mismatch grows as synthetic outputs and short-video platforms enter the training mix.","The mismatch framing points to a fork the paper deliberately does not resolve: either source terms are largely unenforceable against training (shifting risk to copyright law itself), or a large fraction of existing multimodal training is already in breach of terms — future litigation will pick the branch.","Because the volume-weighted percentages are driven by a handful of giant collections, a sensitivity analysis that recomputes the headline figures with the largest collection per modality removed would show how robust the ecosystem-level claims actually are."],"forward_implications":["Developers who filter training data by dataset license alone will systematically misjudge the restrictions on the content, since source-level terms bind more than 80% of content in each modality.","Because the strictest-source rule intensifies at the collection level, large re-packaged text collections hide commercially usable subsets that practitioners cannot easily extract.","The concentration of speech and video on a single video platform means platform terms, not dataset licenses, are the effective gatekeepers of most audio-visual training data.","If the source coding is correct, the permissive commons for multimodal training reduces to the small residual of content that is both commercially licensed and sourced from unrestricted sources — under 1% of text and video content and about 5% of speech hours.","Relative geographical and linguistic representation has been flat for a decade, so adding more languages and countries at the margins does not by itself reduce concentration."],"supporting_citations":[{"why":"Supplies the license annotation taxonomy and manual audit methodology that this paper extends from text to speech and video.","marker":"[123]"},{"why":"The prior derivation-tracing of pretraining data whose method of following re-packaged datasets back to original sources is carried over here.","marker":"[124]"},{"why":"The YouTube-sourced YODAS corpus that dominates speech hours and anchors the finding that internet video overtook curated sources.","marker":"[89]"},{"why":"The early YouTube-based video benchmark that establishes YouTube as a standard source for large-scale video datasets.","marker":"[45]"},{"why":"The crowdsourced multilingual speech corpus whose release drives part of the measured rise in represented languages.","marker":"[13]"},{"why":"The multilingual instruction-tuning collection used as a reference point for diversity efforts and as a source for the dataset list.","marker":"[134]"},{"why":"The video-generation model release whose undocumented training sources motivate the audit's video scope.","marker":"[115]"},{"why":"A survey used to compile the text dataset list, grounding the selection of popular post-training collections.","marker":"[114]"}],"fun_headline_variants":["AI dataset licenses miss 80% of source-level restrictions","3,916-dataset audit: licenses understate AI training restrictions","AI data Western bias persists despite more languages, countries","Web-crawled and social media now dominate AI training data","Trace AI data: licenses hide 80% non-commercial source terms"],"cache_read_input_tokens":67328,"weakest_assumption_plain":"The central percentages depend on the manual classification of each dataset's sources as Unrestricted, Unspecified, Source Closed, or Model Closed, and the paper reports no inter-annotator agreement or independent validation of that coding, so a systematic bias in it would move every headline number.","fun_headline_variants_meta":{"raw":{"variants":["AI dataset licenses miss 80% of source-level restrictions","3,916-dataset audit: licenses understate AI training restrictions","AI data Western bias persists despite more languages, countries","Web-crawled and social media now dominate AI training data","Trace AI data: licenses hide 80% non-commercial source terms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001082,"raw_usage":{"total_tokens":4591,"prompt_tokens":1076,"completion_tokens":3515,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":692,"completion_tokens_details":{"reasoning_tokens":3429}},"tokens_in":692,"tokens_out":3515,"duration_ms":23175,"temperature":1.0,"reasoning_tokens":3429,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:14:38.666361+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of a few hundred datasets from the released audit, have a second independent team re-annotate source restrictions from the same lineage, and compute agreement; low agreement on the Source Closed and Model Closed categories would undercut the 80%-plus mismatch claim. A separate check: recompute the 99.8% text, 78% speech, and 99% video figures with the single largest collection per modality removed (for example the roughly 370k-hour YouTube-sourced speech corpus), since one re-coded giant could shift the ecosystem totals.","supporting_citations":[],"review_version":1}