Pith. sign in

REVIEW 1 cited by

A Framework for Deprecating Datasets: Standardizing Documentation, Identification, and Communication

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.04424 v2 pith:O6342L5M submitted 2021-10-18 cs.CY cs.AI

classification cs.CYcs.AI
keywords datasetsdatasetbeencommunitycycledeprecatingdeprecationdocumentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Datasets are central to training machine learning (ML) models. The ML community has recently made significant improvements to data stewardship and documentation practices across the model development life cycle. However, the act of deprecating, or deleting, datasets has been largely overlooked, and there are currently no standardized approaches for structuring this stage of the dataset life cycle. In this paper, we study the practice of dataset deprecation in ML, identify several cases of datasets that continued to circulate despite having been deprecated, and describe the different technical, legal, ethical, and organizational issues raised by such continuations. We then propose a Dataset Deprecation Framework that includes considerations of risk, mitigation of impact, appeal mechanisms, timeline, post-deprecation protocols, and publication checks that can be adapted and implemented by the ML community. Finally, we propose creating a centralized, sustainable repository system for archiving datasets, tracking dataset modifications or deprecations, and facilitating practices of care and stewardship that can be integrated into research and publication processes.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Completeness of Datasets Documentation on ML/AI repositories: an Empirical Investigation

    cs.DL 2025-02 conditional novelty 6.0 of 10

    Most popular ML/AI datasets are poorly documented, especially regarding collection, processing, and maintenance, according to a manual audit of 100 datasets across four repositories.

Pith tools