Pith. sign in

REVIEW 3 major objections 4 minor 25 references

Completeness of Datasets Documentation on ML/AI repositories: an Empirical Investigation

T0 review · 3 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read Popular ML/AI datasets are poorly documented on data collection and processing, an audit of 100 datasets finds.

desk verdict A useful, honest measurement of how little popular ML datasets document collection and processing; the numbers are plausible but the coding was done by an unvalidated single-coder process, so treat the precision with caution. read the letter →

arxiv 2503.13463 v1 pith:6TZWRJYX submitted 2025-02-10 cs.DL cs.AIcs.HCcs.LG

classification cs.DLcs.AIcs.HCcs.LG
keywords datasetdocumentationdatatransparencyAIaccountabilityreproducibilityDatasheetsforDatasetsprovenanceML/AIrepositoriescompleteness
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Machine-learning datasets underpin many AI systems, yet the documentation that accompanies them is rarely scrutinized. This paper measures how complete that documentation is for the most popular datasets on Hugging Face, Kaggle, OpenML, and UC Irvine MLR. Using a checklist called the Documentation Test Sheet, built from existing documentation proposals, the authors examined 100 datasets and found that information about data collection and processing is often absent. Their conclusion is that even the most influential datasets show a paucity of transparency, and that repository design strongly influences which details get documented.

What carries the argument

The Documentation Test Sheet (DTS) is the measurement instrument. It translates recommendations from Datasheets for Datasets and related schemas into a list of Test Fields, each phrased so a reader can answer whether that information is present (1), absent (0), or not applicable (NA). Presence Check Values are then averaged to produce Presence Averages across datasets, sections, and individual fields, turning qualitative documentation into numeric completeness scores. The DTS is deliberately repository-agnostic and designed to be applied to the documentation page where a dataset is actually downloaded.

What would settle it

Have a panel of diverse dataset users independently apply a user-validated version of the DTS to the same 100 datasets and measure inter-rater agreement; if users deem the documentation adequate on fields the DTS counts as missing, or if different readers disagree sharply on what is present, the claim of a transparency paucity would not hold as stated.

Watch

Extended reading notes

Core claim

The paper's central claim is that public repositories fail to document the contexts in which popular ML/AI datasets are produced. By applying the Documentation Test Sheet to 100 datasets, the authors show that documentation completeness is generally below half, with the largest gaps in the Collection processes and Data processing procedures sections. Only information needed for immediate use, such as task descriptions and instance counts, is consistently present; details about ethics review, potential impacts, and maintenance (e.g., DOI, erratum, version management) are almost entirely missing. The authors conclude that documentation practice is utilisation-oriented and remains a bottleneck for accountability and informed reuse.

Load-bearing premise

The central conclusion depends on the Documentation Test Sheet being the right checklist of what a dataset should always disclose; if the checklist is too strict or misaligned with what users actually need, the measured transparency gap would be an artifact of the schema rather than a real deficit.

Editorial extensions

If this is right

  • Dataset users who rely on repository pages cannot currently verify how data were collected or what ethical review took place, so the risk of downstream harm is hard to assess at the point of selection.
  • Repository platforms are in a position to raise documentation quality: fields reinforced by metadata schemas were present far more often, so adding structured prompts could close the largest gaps.
  • The absence of maintenance information such as DOIs, errata, and version management means popular datasets can continue to be used after they become stale or known to contain errors.
  • The DTS can be used directly by dataset creators and hosts as a pre-publication checklist to reduce missing documentation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A replication on newly released datasets, not just popular legacy ones, would show whether documentation practice is improving as data curation awareness grows.
  • A companion study measuring the accuracy of the information that is present would extend the transparency argument beyond completeness, since a complete but incorrect description still misleads users.
  • The finding that repository metadata structure drives documentation suggests an experiment: randomly assigning a set of dataset upload pages to a DTS-aligned form and comparing completeness against a control group.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper constructs a Documentation Test Sheet (DTS), a checklist of information fields derived mainly from Datasheets for Datasets and other documentation proposals, and applies it manually to 100 popular datasets hosted on Hugging Face, Kaggle, OpenML, and UCI. Each field is scored 1, 0, or NA, and the paper reports presence averages at the dataset, section, and field level. The main empirical finding is that documentation is generally sparse, with particularly low scores for collection processes, data processing procedures, ethical review, and maintenance-related information, and that repository metadata structure strongly influences which fields are present. The authors conclude that there is a paucity of transparency in popular ML/AI dataset documentation.

Significance. If the measurements are reliable, the study provides a useful, descriptive evidence base about documentation practices in influential ML/AI repositories, with practical implications for dataset creators and repository maintainers. The paper is transparent about its design: the DTS fields are traced to prior literature, the repository selection is justified, the reading principles are documented in an appendix, and the limitations are stated candidly. The qualitative direction of the findings—particularly the scarcity of information about data collection and processing—is plausible and consistent with existing concerns about dataset stewardship. The paper does not fit models or make predictive claims; its contribution is primarily the measurement protocol and the resulting dataset of presence scores, both of which are valuable for follow-up work.

major comments (3)
  1. [§3.2, §4.3, Appendix D/E] The core measurements in Table 2 are produced by manual reading of repository documentation, but the paper reports no number of coders, no independent coding of any subset, and no inter-rater agreement statistic. Since every field average and every cross-repository comparison inherits these coding decisions, the precision of the reported averages (e.g., 3.04 ethical review = 0.01, 3.07 potential impacts = 0.00) cannot be independently checked. I request a reliability assessment on a random subset of datasets with at least two coders and an agreement measure such as Cohen's kappa or Krippendorff's alpha, or, failing that, an explicit re-framing of the results as a single-coder descriptive screening whose numeric precision is not established. The raw coding should also be made available on a stable repository rather than a temporary link.
  2. [§2.1, §5] The DTS operationalizes what information 'should always be attached' to a dataset, and the conclusion of a 'paucity of transparency' follows directly from the completeness scores produced by this schema. Section 5 states that the DTS was not tested for consistency and validation with target users. This matters because fields that are not relevant to a particular dataset or that are present only through repository metadata (e.g., 2.11 Statistics, 5.04, 6.05) can move the section and field averages. I ask for an explicit discussion of the consequences of this construct-validity limitation for the central conclusion, and for a sensitivity check that reports section averages after excluding fields whose presence is structurally determined by the repository metadata schema. Alternatively, the conclusion should be strictly conditional on the DTS operationalization.
  3. [§2.2, Table 2] The definition of Presence Average in Section 2.2 does not specify how NA values are treated in the averaging. Fields such as 2.08, 2.09, 2.10, 3.05, and 3.06 are explicitly conditional on the dataset being people-related, and the reported overall averages (0.43, 0.15, 0.03, 0.05, 0.05) are inconsistent with a denominator of all 100 datasets, suggesting that NA entries are excluded. This should be stated precisely: give the formula, define the denominator for each aggregation level, and report the number of applicable datasets for each conditional field. Without this, the reader cannot recompute the averages or compare repositories.
minor comments (4)
  1. [Table 2 and text] The comma is used as the decimal separator in tables and many numbers in the text; using decimal points consistently would improve readability for an English-language audience.
  2. [§2, online appendix] The online appendix link is a temporary mega.nz folder; a permanent archival location for the DTS, reading principles, and raw results should be provided.
  3. [§3.2] The paper states that duplicates were eliminated within and across repositories, but it does not report how many duplicates were removed; please give this number and describe how the final 100 datasets were distributed after deduplication.
  4. [§3.2, footnote 5] The parenthetical in footnote 5 reads as an incomplete sentence and should be expanded or reworded.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: the Documentation Test Sheet is an externally sourced normative checklist, presence averages are direct observations, and no fitted parameter or self-cited uniqueness claim is load-bearing.

full rationale

The paper's derivation chain is empirical rather than formal: (1) the Documentation Test Sheet (DTS) is constructed from prior published documentation schemas, explicitly 'largely based on Datasheets for Datasets [8]' with insights from Data Statements and Nutrition Labels; (2) the authors manually read 100 repository documentation pages and assign 1/0/NA per Test Field; (3) Presence Averages are computed by averaging these direct observations; (4) the conclusion of a 'paucity of transparency' follows from the low observed averages. At no point is the DTS fitted to the 100 datasets, nor is any parameter tuned to reproduce the observed completeness values. The only concerns raised by a skeptical reader are construct validity of the DTS and inter-rater reliability of the manual coding. These are genuine threats to validity and are partially acknowledged in Section 5: 'the DTS was not tested for consistency and validation with target users' and the study is 'inherently prone to the risk of interpretation errors.' But these are not circularity: the measurements could in principle have been high, and the low averages are contingent observations rather than consequences of the definition of the DTS. No load-bearing self-citation, uniqueness theorem, or renamed result is present. Accordingly, the circularity score is minimal.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The DTS is a new measurement instrument, but it is assembled from existing literature rather than invented ad hoc. No free parameters are fitted. The main assumptions are the normative validity of the DTS and the reliability of manual inspection.

assumptions (3)
  • domain assumption The Documentation Test Sheet fields, derived from Datasheets for Datasets, Data Statements, and the Dataset Nutrition Label, represent the information that should always accompany a dataset.
    Section 2.1 states the DTS was built from relevant studies in the literature to measure completeness for transparency, accountability, and reproducibility. This normative assumption is load-bearing: the low presence averages and the 'paucity of transparency' conclusion depend on the DTS being a valid measure.
  • domain assumption The popularity proxies (downloads, views, upvotes) identify the most influential datasets for the field.
    Section 3.2 selects 25 datasets per repository using these proxies. If popularity is not a good proxy for influence, the sample would not represent the most used datasets.
  • domain assumption Manual reading of documentation by the researchers, guided by the reading principles in Appendix D, produces reliable and consistent presence values.
    Section 5 acknowledges the inherent risk of interpretation errors and that the DTS was not tested for consistency with target users; no inter-rater reliability is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Completeness of Datasets Documentation on ML/AI repositories: an Empirical Investigation." pith.science (2026). https://pith.science/paper/6TZWRJYX

@misc{pith2026250313463,
  author       = {Pith},
  title        = {Pith review of: Completeness of Datasets Documentation on ML/AI repositories: an Empirical Investigation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6TZWRJYX}},
  note         = {Machine review of arXiv:2503.13463}
}
read the original abstract

ML/AI is the field of computer science and computer engineering that arguably received the most attention and funding over the last decade. Data is the key element of ML/AI, so it is becoming increasingly important to ensure that users are fully aware of the quality of the datasets that they use, and of the process generating them, so that possible negative impacts on downstream effects can be tracked, analysed, and, where possible, mitigated. One of the tools that can be useful in this perspective is dataset documentation. The aim of this work is to investigate the state of dataset documentation practices, measuring the completeness of the documentation of several popular datasets in ML/AI repositories. We created a dataset documentation schema -- the Documentation Test Sheet (DTS) -- that identifies the information that should always be attached to a dataset (to ensure proper dataset choice and informed use), according to relevant studies in the literature. We verified 100 popular datasets from four different repositories with the DTS to investigate which information was present. Overall, we observed a lack of relevant documentation, especially about the context of data collection and data processing, highlighting a paucity of transparency.

Figures

Figures reproduced from arXiv: 2503.13463 by the authors.

Figure 1
Figure 1. Distribution of Dataset Presence Averages grouped by repository. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Section Presence Averages. that structurally expose this information in the metadata schema of the repos￾itory. Conversely, it was almost completely absent in repositories that do not include such information in their metadata schema. Some examples are: 2.11 Statistics (hug 0,00; kag 1,00; oml 1,00; uci 1,00), 5.04 Repository that links to papers or system that use the datasets (hug 0,92; kag 0,00; oml 0,00; uci 1,0… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

25 extracted references · 9 canonical work pages

  1. [1]

    In: 2021 IEEE International Conference on Smart Data Services (SMDS)

    Afzal, S., Rajmohan, C., Kesarwani, M., Mehta, S., Patel, H.: Data Readiness Report. In: 2021 IEEE International Conference on Smart Data Services (SMDS). pp. 42–51 (Sep 2021). https://doi.org/10.1109/SMDS53860.2021.00016

  2. [2]

    arXiv:1808.07261 [cs] p

    Arnold, M., Bellamy, R.K.E., Hind, M., Houde, S., Mehta, S., Mojsilovic, A., Nair, R., Ramamurthy, K.N., Reimer, D., Olteanu, A., Piorkowski, D., Tsay, J., Varsh- ney, K.R.: FactSheets: Increasing Trust in AI Services through Supplier’s Declara- tions of Conformity. arXiv:1808.07261 [cs] p. 31 (Feb 2019)

  3. [3]

    Transactions of the Association for Computational Linguistics 6, 587–604 (2018)

    Bender, E.M., Friedman, B.: Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science. Transactions of the Association for Computational Linguistics 6, 587–604 (2018). https://doi.org/10. 1162/tacl a 00041

  4. [4]

    Bender, E.M., Gebru, T., McMillan-Major, A., Shmitchell, S.: On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. pp. 610–

  5. [5]

    Proceedings of the ACM on Human-Computer Interaction 5(CSCW2), 438:1–438:27 (Oct 2021)

    Boyd, K.L.: Datasheets for Datasets help ML Engineers Notice and Understand Ethical Issues in Training Data. Proceedings of the ACM on Human-Computer Interaction 5(CSCW2), 438:1–438:27 (Oct 2021). https://doi.org/10.1145/3479582

  6. [6]

    A Framework for Deprecating Datasets: Standardizing Documentation, Identification, and Communication

    Corry, F., Sridharan, H., Luccioni, A.S., Ananny, M., Schultz, J., Crawford, K.: The Problem of Zombie Datasets:A Framework For Deprecating Datasets. arXiv:2111.04424 [cs] p. 18 (Oct 2021)

  7. [7]

    Fabris, A., Messina, S., Silvello, G., Susto, G.A.: Algorithmic Fairness Datasets: The Story so Far (May 2022)

  8. [8]

    arXiv:1803.09010 [cs] p

    Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J.W., Wallach, H., Daum´ e III, H., Crawford, K.: Datasheets for Datasets. arXiv:1803.09010 [cs] p. 18 (Dec 2021)

Show all 25 references
  1. [10]

    arXiv:1805.03677 [cs] p

    Holland, S., Hosny, A., Newman, S., Joseph, J., Chmielinski, K.: The Dataset Nutrition Label: A Framework To Drive Higher Data Quality Standards. arXiv:1805.03677 [cs] p. 21 (May 2018)

  2. [11]

    In: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency

    Hutchinson, B., Smart, A., Hanna, A., Denton, E., Greer, C., Kjartansson, O., Barnes, P., Mitchell, M.: Towards Accountability for Machine Learning Datasets: Practices from Software Engineering and Infrastructure. In: Proceedings of the 2021 ACM Conference on Fairness, Account...

  3. [12]

    Proceedings of the 2020 Conference on Fairness, Ac- countability, and Transparency pp

    Jo, E.S., Gebru, T.: Lessons from Archives: Strategies for Collecting Sociocultural Data in Machine Learning. Proceedings of the 2020 Conference on Fairness, Ac- countability, and Transparency pp. 306–316 (Jan 2020). https://doi.org/10.1145/ 3351095.3372829

  4. [13]

    https://doi.org/ 10.48550/arXiv.2112.01716

    Koch, B., Denton, E., Hanna, A., Foster, J.G.: Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research (Dec 2021). https://doi.org/ 10.48550/arXiv.2112.01716

  5. [14]

    Digital Policy, Regulation and Governance 23(5), 475–488 (Jan 2021)

    K¨ onigstorfer, F., Thalmann, S.: Software documentation is not enough! Require- ments for the documentation of AI. Digital Policy, Regulation and Governance 23(5), 475–488 (Jan 2021). https://doi.org/10.1108/DPRG-03-2021-0047

  6. [15]

    Proceedings of the Conference on Fairness, Accountability, and Transparency pp

    Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I.D., Gebru, T.: Model Cards for Model Reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency pp. 220–229 (Jan 2019). https://doi.org/10.1145/32875...

  7. [16]

    arXiv:2108.02922 [cs] p

    Peng, K., Mathur, A., Narayanan, A.: Mitigating Dataset Harms Requires Stew- ardship: Lessons from 1000 Papers. arXiv:2108.02922 [cs] p. 28 (Nov 2021)

  8. [17]

    Journal of Statistical Software 90, 1–38 (Jul 2019)

    Petersen, A.H., Ekstrøm, C.T.: dataMaid: Your Assistant for Documenting Super- vised Data Quality Screening in R. Journal of Statistical Software 90, 1–38 (Jul 2019). https://doi.org/10.18637/jss.v090.i06

  9. [18]

    arXiv:2006.13796 [cs] p

    Richards, J., Piorkowski, D., Hind, M., Houde, S., Mojsilovi´ c, A.: A Methodology for Creating AI FactSheets. arXiv:2006.13796 [cs] p. 18 (Jun 2020)

  10. [19]

    Everyone wants to do the model work, not the data work

    Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P., Aroyo, L.M.: “Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI. In: Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. pp. 1–15. CHI ’21, Ass...

  11. [20]

    Proceedings of the ACM on Human-Computer Interaction 5(CSCW2), 1–37 (Oct 2021)

    Scheuerman, M.K., Denton, E., Hanna, A.: Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset Development. Proceedings of the ACM on Human-Computer Interaction 5(CSCW2), 1–37 (Oct 2021). https://doi.org/10. 1145/3476058

  12. [21]

    https://doi.org/10.48550/arXiv.2202.06675

    Schramowski, P., Tauchmann, C., Kersting, K.: Can Machines Help Us Answering Question 16 in Datasheets, and In Turn Reflecting on Inappropriate Content? (Jul 2022). https://doi.org/10.48550/arXiv.2202.06675

  13. [22]

    In: Proceedings of the 28th ACM International Conference on Information and Knowledge Manage- ment

    Sun, C., Asudeh, A., Jagadish, H.V., Howe, B., Stoyanovich, J.: MithraLabel: Flex- ible Dataset Nutritional Labels for Responsible Data Science. In: Proceedings of the 28th ACM International Conference on Information and Knowledge Manage- ment. pp. 2893–2896. CIKM ’19, Associa...

  14. [23]

    Media, Culture & Society p

    Thylstrup, N.B.: The ethics and politics of data sets in the age of machine learning: Deleting traces and encountering remains. Media, Culture & Society p. 01634437211060226 (Apr 2022). https://doi.org/10.1177/01634437211060226

  15. [24]

    Proceedings of the 2018 International Conference on Management of Data pp

    Yang, K., Stoyanovich, J., Asudeh, A., Howe, B., Jagadish, H.V., Miklau, G.: A Nu- tritional Label for Rankings. Proceedings of the 2018 International Conference on Management of Data pp. 1773–1776 (May 2018). https://doi.org/10.1145/3183713. 3193568

  16. [25]

    https://doi.org/10.48550/arXiv.2103.14000

    Zehlike, M., Yang, K., Stoyanovich, J.: Fairness in Ranking: A Survey (May 2021). https://doi.org/10.48550/arXiv.2103.14000

  17. [575]

    https://doi.org/10.1145/3442188.3445918

    F AccT ’21, Association for Computing Machinery, New York, NY, USA (Mar 2021). https://doi.org/10.1145/3442188.3445918

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.