Pith. sign in

REVIEW 2 cited by

Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.01716 v1 pith:6UZB74ZX submitted 2021-12-03 cs.LG cs.CLcs.CVcs.CYstat.ML

classification cs.LGcs.CLcs.CVcs.CYstat.ML
keywords acrossdatasetslearningmachinedatasetfieldresearchbeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Benchmark datasets play a central role in the organization of machine learning research. They coordinate researchers around shared research problems and serve as a measure of progress towards shared goals. Despite the foundational role of benchmarking practices in this field, relatively little attention has been paid to the dynamics of benchmark dataset use and reuse, within or across machine learning subcommunities. In this paper, we dig into these dynamics. We study how dataset usage patterns differ across machine learning subcommunities and across time from 2015-2020. We find increasing concentration on fewer and fewer datasets within task communities, significant adoption of datasets from other tasks, and concentration across the field on datasets that have been introduced by researchers situated within a small number of elite institutions. Our results have implications for scientific evaluation, AI ethics, and equity/access within the field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM Performance for Code Generation on Noisy Tasks

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LLMs solve heavily obfuscated benchmark tasks, and performance decay under obfuscation differs sharply between old and new datasets, which the authors interpret as a signature of training-data contamination.

  2. Rethinking Indic AI from a Lens of Cultural Heritage Preservation

    cs.AI 2026-07 conditional novelty 4.0 of 10

    The paper surveys Indic NLP evolution and proposes 'Culture Sensing' to integrate indigenous oral knowledge into foundation models for cultural preservation.

Pith tools