Pith. sign in

REVIEW 14 cited by

The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.16787 v3 pith:6DLC3FX5 submitted 2023-10-25 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords datadatasetsauditdatasetlegallicenseprovenancetrace
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The race to train language models on vast, diverse, and inconsistently documented datasets has raised pressing concerns about the legal and ethical risks for practitioners. To remedy these practices threatening data transparency and understanding, we convene a multi-disciplinary effort between legal and machine learning experts to systematically audit and trace 1800+ text datasets. We develop tools and standards to trace the lineage of these datasets, from their source, creators, series of license conditions, properties, and subsequent use. Our landscape analysis highlights the sharp divides in composition and focus of commercially open vs closed datasets, with closed datasets monopolizing important categories: lower resource languages, more creative tasks, richer topic variety, newer and more synthetic training data. This points to a deepening divide in the types of data that are made available under different license conditions, and heightened implications for jurisdictional legal interpretations of copyright and fair use. We also observe frequent miscategorization of licenses on widely used dataset hosting sites, with license omission of 70%+ and error rates of 50%+. This points to a crisis in misattribution and informed use of the most popular datasets driving many recent breakthroughs. As a contribution to ongoing improvements in dataset transparency and responsible use, we release our entire audit, with an interactive UI, the Data Provenance Explorer, which allows practitioners to trace and filter on data provenance for the most popular open source finetuning data collections: www.dataprovenance.org.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Domain-Aware Scaling Laws Uncover Data Synergy

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Domain-aware scaling laws with fitted γ and σ synergy terms recover stable code-math interactions from observational LLM mixtures and correctly predict mixture rankings in controlled small-scale trainings.

  2. FlexOlmo: Open Language Models for Flexible Data Use

    cs.CL 2025-07 conditional novelty 7.0 of 10

    FlexOlmo merges independently trained language-model experts, trained on private data, into a single mixture-of-experts model without joint training.

  3. The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

    cs.CL 2025-05 conditional novelty 7.0 of 10

    NaijaVoices is a 1,838-hour, 5,455-speaker speech-text corpus for Igbo, Hausa, and Yoruba whose use in fine-tuning cuts Word Error Rates by 42-76% relative to unadapted baselines.

  4. The Leaderboard Illusion

    cs.AI 2025-04 conditional novelty 7.0 of 10

    Chatbot Arena's rankings are systematically distorted by undisclosed private testing, selective score reporting, and data access asymmetries that favor large proprietary providers.

  5. Do LLMs Provide Links to Code Similar to what they Generate? A Study with Gemini and Bing CoPilot

    cs.SE 2025-01 conditional novelty 7.0 of 10

    In an empirical study of 437 code snippets, Bing CoPilot usually provided at least one link with code similar to its output, while Gemini rarely did, exposing a real provenance gap.

  6. KatotohananQA: Evaluating Truthfulness of Large Language Models in Filipino

    cs.CL 2025-09 conditional novelty 6.0 of 10

    KatotohananQA is a Filipino translation of TruthfulQA; seven LLMs scored 94.72% in English versus 83.87% in Filipino, with GPT-5 and GPT-5 mini showing the smallest gap.

  7. TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation

    cs.CY 2025-05 conditional novelty 6.0 of 10

    A new 143-indicator rubric applied to 114 human-voice datasets shows that documentation of consent, privacy, and harmful content is rare, and that scraping yields scale at the cost of documented ethical practices.

  8. The AI Agent Index

    cs.SE 2025-02 accept novelty 6.0 of 10

    The AI Agent Index catalogs 67 deployed agentic AI systems and shows that most developers publicly disclose little about safety policies and evaluations.

  9. The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models

    cs.CL 2025-01 conditional novelty 6.0 of 10

    The Heap is a new multilingual code dataset of 32.7M copyleft-licensed files with exact and near-duplicate flags relative to The Stack, Red Pajama, GitHub Code, and CodeParrot.

  10. Bridging the Data Provenance Gap Across Text, Speech and Video

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with...

  11. Uhura: A Benchmark for Evaluating Scientific Question Answering and Truthfulness in Low-Resource African Languages

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A six-language translated benchmark shows large language models are significantly less accurate on science and truthfulness questions in African languages than in English, and closed models beat open models.

  12. Red Teaming AI Policy: A Taxonomy of Avoision and the EU AI Act

    cs.CY 2025-06 accept novelty 5.0 of 10

    A taxonomy of avoision under the EU AI Act, with strategies to escape scope, exploit exemptions, and manipulate risk or operator categories.

  13. The Reality of AI and Biorisk

    cs.AI 2024-12 conditional novelty 4.0 of 10

    A review of the evidence finds no current, statistically significant uplift in biorisk from LLMs or AI biological tools, but the underlying studies are too nascent to justify strong conclusions.

  14. The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)

    cs.SE 2025-05 conditional novelty 3.0 of 10

    The paper catalogs the lifecycle stages and production-readiness challenges of software built around foundation models (FMware) and proposes an action plan of engineering practices and research directions.

Pith tools