Pith. sign in

REVIEW 5 cited by

Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on Hugging Face

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.13822 v1 pith:TR2D6APZ submitted 2024-01-24 cs.LG cs.AI

Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on Hugging Face

classification cs.LG cs.AI
keywords datasetdocumentationsectiondatafacehugginganalyzingcard
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Advances in machine learning are closely tied to the creation of datasets. While data documentation is widely recognized as essential to the reliability, reproducibility, and transparency of ML, we lack a systematic empirical understanding of current dataset documentation practices. To shed light on this question, here we take Hugging Face -- one of the largest platforms for sharing and collaborating on ML models and datasets -- as a prominent case study. By analyzing all 7,433 dataset documentation on Hugging Face, our investigation provides an overview of the Hugging Face dataset ecosystem and insights into dataset documentation practices, yielding 5 main findings: (1) The dataset card completion rate shows marked heterogeneity correlated with dataset popularity. (2) A granular examination of each section within the dataset card reveals that the practitioners seem to prioritize Dataset Description and Dataset Structure sections, while the Considerations for Using the Data section receives the lowest proportion of content. (3) By analyzing the subsections within each section and utilizing topic modeling to identify key topics, we uncover what is discussed in each section, and underscore significant themes encompassing both technical and social impacts, as well as limitations within the Considerations for Using the Data section. (4) Our findings also highlight the need for improved accessibility and reproducibility of datasets in the Usage sections. (5) In addition, our human annotation evaluation emphasizes the pivotal role of comprehensive dataset content in shaping individuals' perceptions of a dataset card's overall quality. Overall, our study offers a unique perspective on analyzing dataset documentation through large-scale data science analysis and underlines the need for more thorough dataset documentation in machine learning research.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Ethical and Technical Limits of Deepfake Speech Datasets

    cs.SD 2026-06 unverdicted novelty 6.0

    Audit of 39 deepfake speech datasets shows most lack demographic metadata making fairness checks infeasible and reveals substantial overlap in bona fide sources that undermines cross-dataset generalization claims.

  2. ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery

    cs.LG 2026-05 unverdicted novelty 6.0

    ArtifactLinker frames SOTA discovery as missing-link prediction on an artifact graph of models and datasets, with a two-stage ranking-plus-verification pipeline and a new benchmark of 14k artifacts.

  3. AdaQE-CG: Adaptive Query Expansion for Web-Scale Generative AI Model and Data Card Generation

    cs.AI 2026-03 unverdicted novelty 6.0

    AdaQE-CG uses context-aware adaptive query expansion and inter-card knowledge transfer from a MetaGAI Pool to generate higher-quality model and data cards than prior methods, validated on the new expert-annotated Meta...

  4. FAIR^2 Drones: An AI-Ready Standard for Cross-Domain Wildlife Drone Datasets

    cs.RO 2026-05 unverdicted novelty 5.0

    FAIR^2 Drones is a proposed standard that adds platform metadata and annotation specifications to existing FAIR and AI-ready frameworks so wildlife drone datasets can support ecological analysis, robotics development,...

  5. How Hyper-Datafication Impacts the Sustainability Costs in Frontier AI

    cs.CY 2026-01 unverdicted novelty 5.0

    Hyper-datafication in frontier AI increases resource consumption and redistributes environmental burdens, labor risks, and representational harms toward the Global South, data workers, and under-represented cultures, ...