Pith. sign in

REVIEW 3 cited by

Addressing "Documentation Debt" in Machine Learning Research: A Retrospective Datasheet for BookCorpus

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.05241 v1 pith:FIRA672C submitted 2021-05-11 cs.CL cs.CYcs.LG

classification cs.CLcs.CYcs.LG
keywords bookcorpusdocumentationdatasetdatasheetdebtlearningmachinemodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent literature has underscored the importance of dataset documentation work for machine learning, and part of this work involves addressing "documentation debt" for datasets that have been used widely but documented sparsely. This paper aims to help address documentation debt for BookCorpus, a popular text dataset for training large language models. Notably, researchers have used BookCorpus to train OpenAI's GPT-N models and Google's BERT models, even though little to no documentation exists about the dataset's motivation, composition, collection process, etc. We offer a preliminary datasheet that provides key context and information about BookCorpus, highlighting several notable deficiencies. In particular, we find evidence that (1) BookCorpus likely violates copyright restrictions for many books, (2) BookCorpus contains thousands of duplicated books, and (3) BookCorpus exhibits significant skews in genre representation. We also find hints of other potential deficiencies that call for future research, including problematic content, potential skews in religious representation, and lopsided author contributions. While more work remains, this initial effort to provide a datasheet for BookCorpus adds to growing literature that urges more careful and systematic documentation for machine learning datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Certified Mitigation of Worst-Case LLM Copyright Infringement

    cs.CL 2025-04 conditional novelty 6.0 of 10

    BloomScrub detects long verbatim quotes from a protected corpus with a Bloom filter, rewrites them iteratively, and abstains when needed, certifying that no quote longer than the threshold is emitted.

  2. Bridging the Data Provenance Gap Across Text, Speech and Video

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with...

  3. One Documentation Does Not Fit All: Case Study of TensorFlow Documentation

    cs.SE 2025-05 conditional novelty 4.0 of 10

    TensorFlow's documentation does not effectively support its target users: 64.3% of documentation-related Stack Overflow questions are triggered by inadequate or non-generalizable examples, and user profiles do not dif...

Pith tools