Pith. sign in

REVIEW 4 cited by

Datasheet for the Pile

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.07311 v1 pith:5T53UCRZ submitted 2022-01-13 cs.CL

classification cs.CL
keywords piletextavailabledatadatasheetscrapescompiledcomprised
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This datasheet describes the Pile, a 825 GiB dataset of human-authored text compiled by EleutherAI for use in large-scale language modeling. The Pile is comprised of 22 different text sources, ranging from original scrapes done for this project, to text data made available by the data owners, to third-party scrapes available online.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Causal Estimation of Tokenisation Bias

    cs.CL 2025-06 conditional novelty 7.0 of 10

    Using regression discontinuity, the paper shows that adding a subword to a tokenizer's vocabulary can raise the model's probability for that string by up to about 17 times in small models.

  2. Language Models Represent and Transform Concepts with Shared Geometry

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Contextual displacements of concepts in LLMs form semantically organized vector fields whose relational geometry is shared across models and predicts held-out displacements above chance.

  3. Natural Context Drift Undermines the Natural Language Understanding of Large Language Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    QA accuracy of open-weight LLMs drops as Wikipedia passages semantically drift from training-time content, while human accuracy stays flat.

  4. How Quantization Impacts Privacy Risk on LLMs for Code?

    cs.SE 2025-07 conditional novelty 6.0 of 10

    Quantizing code LLMs reduces membership inference effectiveness, with 4-bit compression giving larger privacy and performance drops than 8-bit.

Pith tools