REVIEW 8 cited by
The Files are in the Computer: On Copyright, Memorization, and Generative AI
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The New York Times's copyright lawsuit against OpenAI and Microsoft alleges OpenAI's GPT models have "memorized" NYT articles. Other lawsuits make similar claims. But parties, courts, and scholars disagree on what memorization is, whether it is taking place, and what its copyright implications are. These debates are clouded by ambiguities over the nature of "memorization." We attempt to bring clarity to the conversation. We draw on the technical literature to provide a firm foundation for legal discussions, providing a precise definition of memorization: a model has "memorized" a piece of training data when (1) it is possible to reconstruct from the model (2) a near-exact copy of (3) a substantial portion of (4) that piece of training data. We distinguish memorization from "extraction" (user intentionally causes a model to generate a near-exact copy), from "regurgitation" (model generates a near-exact copy, regardless of user intentions), and from "reconstruction" (the near-exact copy can be obtained from the model by any means). Several consequences follow. (1) Not all learning is memorization. (2) Memorization occurs when a model is trained; regurgitation is a symptom not its cause. (3) A model that has memorized training data is a "copy" of that training data in the sense used by copyright. (4) A model is not like a VCR or other general-purpose copying technology; it is better at generating some types of outputs (possibly regurgitated ones) than others. (5) Memorization is not a phenomenon caused by "adversarial" users bent on extraction; it is latent in the model itself. (6) The amount of training data that a model memorizes is a consequence of choices made in training. (7) Whether or not a model that has memorized actually regurgitates depends on overall system design. In a very real sense, memorized training data is in the model--to quote Zoolander, the files are in the computer.
Forward citations
Cited by 8 Pith papers
-
Probabilistic "Copies" in Generative AI Models
An LLM is an infringing copy of a work only when the work can be extracted from it with relatively little effort, so some models are copies of some works and no model is a copy of everything it trained on.
-
User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies
All six leading U.S. AI chatbot developers, as of May 2025, appear to train their models on users' chat data by default, often without clear opt-out options.
-
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
A new 8TB openly-licensed text corpus trains 7B LLMs that are competitive with Llama 1/2, showing that performant models need not depend on unlicensed web data.
-
Extracting memorized pieces of (copyrighted) books from open-weight language models
Book-level memorization varies sharply across LLMs: most books escape most models, yet Llama 3.1 70B can reproduce Harry Potter almost entirely from a six-token seed prompt.
-
How To Think About End-To-End Encryption and AI: Training, Processing, Disclosure, and Consent
Training shared AI models on end-to-end encrypted messages is incompatible with E2EE confidentiality; AI inference on encrypted content is compatible only with endpoint-local processing or strict per-user, no-third-pa...
-
Developer Perspectives on Licensing and Copyright Issues Arising from Generative AI for Software Development
A plurality of developers would place AI-generated code in the public domain, most see it as similar to reusing existing code, and few document AI usage or have copyright training.
-
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
The paper proposes that GenAI evaluation be standardized with a four-level social science measurement framework that separates concept definition from measurement and emphasizes validity testing.
-
A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
A framework that systematizes, operationalizes, and applies amounts, concepts, instances, and populations for valid GenAI measurement, extending Adcock and Collier's measurement theory.
Discussion (0). Continue with ORCID to comment.