REVIEW 3 major objections 3 minor 1 cited by
Valid extractable memorization requires a matched non-training baseline so generation probability exceeds predictability, not just occurrence.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-15 04:25 UTC pith:34IOCVMS
load-bearing objection A useful definitional cleanup for extractable memorization that forces matched baselines, but the load-bearing unmemorized-comparator assumption is fragile and we only have the abstract. the 3 major comments →
Extractable Memorization From First Principles
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A generation of a training sequence constitutes extractable memorization only when its probability exceeds a baseline established by matched non-training sequences that cannot themselves have been memorized; this baseline is obtained either by a conformal threshold calibrated to a target false-positive rate or by a census against a matched non-training document, and without it extraction claims lack validity.
What carries the argument
Matched comparison: either a conformal test that calibrates a probability threshold to a chosen false-positive rate against a non-training population, or a census that compares a single training document against a matched non-training document; both convert raw generation rates into calibrated evidence of memorization.
Load-bearing premise
Non-training sequences drawn as matched comparators must be free of memorization and supply an unbiased baseline for predictability; if they are themselves partially memorized or poorly matched, the calibrated thresholds and false-positive control fail.
What would settle it
Show that, for a fixed model and sequence length, non-training sequences carefully matched in domain and style are regenerated at rates statistically indistinguishable from training sequences after conformal calibration, so that no training sequence exceeds the resulting threshold.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript argues that valid extractable-memorization claims for LLMs require a matched comparison: a training sequence must be generated with probability high enough relative to comparable non-training sequences to exceed predictability rather than mere fluency. Non-training sequences are treated as an unmemorized baseline. This is formalized in two constructions described in the abstract: (1) a conformal test that samples training and non-training sequences from populations and calibrates a threshold to a chosen false-positive rate, and (2) a census that calibrates against a single matched non-training document (e.g., a book). The abstract reports that prior extraction setups omit such comparisons and therefore suffer validity problems—for example, OLMo 2 32B reproduces non-training 10-token Wikipedia suffixes roughly 24% as often as training ones—and that calibrated thresholds for Llama 3.1 70B on books can be as low as 1e-27. On this basis the paper refines “extractable memorization” to require a valid memorization claim and near-certain generation within a realistic budget.
Significance. If the formalizations and empirical claims hold under full scrutiny, the work would supply a clearer validity criterion for extractable-memorization studies and help adjudicate both overstated short-suffix extractions and claims that extraction is unreliable evidence of memorization. The conformal and census constructions, the concrete numerical illustrations (≈24% non-training generation share; thresholds ~1e-27), and the refined definition would be useful methodological contributions to the LLM memorization and privacy literature. Credit is due for framing the problem as an external-baseline comparison rather than an absolute generation-probability threshold, which is conceptually non-circular by construction. Assessment of significance is necessarily provisional: only the abstract is available, so derivations, experimental design, and robustness checks cannot be verified.
major comments (3)
- [Abstract (matched non-training baseline)] The central validity guarantee rests on the premise that matched non-training sequences “cannot have been memorized” and therefore supply an unbiased predictability floor (Abstract). For web-scale LLM corpora this premise is load-bearing and fragile: contamination, near-duplicates, or domain/style mismatch in the comparator set would invalidate both the conformal FPR calibration and the census threshold, and would re-interpret the reported ~24% Wikipedia non-training generation share. The abstract states no verification procedure (membership checks, near-duplicate filtering, or contamination audit) for the non-training baseline. A full manuscript must either provide such a procedure or quantify sensitivity of the calibrated thresholds and the 24% figure to residual contamination.
- [Abstract (conformal test / chosen FPR)] The conformal construction calibrates a threshold to a “chosen FPR” when sequences are sampled from populations (Abstract). The choice of FPR and of the matching/sampling distribution for non-training sequences are free parameters that directly determine which training sequences are declared memorized. Without a full experimental section specifying the sampling distribution, matching criteria, and sensitivity of reported claims (including the 1e-27 book thresholds) to these choices, the numerical illustrations cannot be assessed as robust. The manuscript should report sensitivity analyses over FPR and matching design, or justify a default that is not post-hoc.
- [Abstract (1e-27 thresholds; refined definition)] The abstract asserts that calibrated thresholds “as low as 1e-27” support memorization claims for sequences that “no feasible sampling budget would extract,” and on that basis refines extractable memorization to require near-certain generation within a realistic budget. That refinement is a definitional move whose empirical support depends on the same baseline and calibration steps. Because only the abstract is available, the derivation of the 1e-27 figure, the definition of “realistic budget,” and the mapping from calibrated probability to extractability under sampling cannot be checked. These steps are load-bearing for the refined definition and must be fully specified and stress-tested in the complete manuscript.
minor comments (3)
- [Abstract] The abstract is clear and well structured, but several technical terms (conformal test, census, matched document, FPR) are introduced without even a one-line formal sketch. A sentence or two of notation in the abstract or early introduction would help readers locate the constructions before the full methods section.
- [Abstract (numerical claims)] The 24% figure is reported as “roughly 24% as often”; the 1e-27 figure as “as low as 1e-27.” When the full text is available, these should be tied to exact table/figure entries, model checkpoints, and token-length definitions so that the claims are reproducible from the abstract’s numerical anchors.
- [Abstract (closing sentence)] The refined definition of extractable memorization is stated only at the end of the abstract. It would help to flag earlier that the paper’s contribution is both methodological (matched comparisons) and definitional, so readers know what is being revised versus what is being measured.
Circularity Check
No significant circularity: matched non-training baselines supply an external predictability floor; the refined definition and empirical rates are not forced by construction from the inputs.
full rationale
The abstract-only paper proposes a methodological criterion for valid extractable-memorization claims: generation probability of a training sequence must exceed a baseline obtained from matched non-training sequences (via a conformal FPR-calibrated threshold when sampling populations, or a census against a matched non-training document). Non-training sequences are treated as an external comparator that cannot have been memorized, so the comparison is not self-definitional. Reported quantities (e.g., Wikipedia non-training 10-token suffixes generated ~24% as often as training ones; Llama book thresholds as low as 1e-27) are empirical measurements under that design, not fitted parameters renamed as predictions. No uniqueness theorem, ansatz, or load-bearing self-citation appears in the abstract; the refinement of “extractable memorization” is a definitional proposal motivated by the validity argument rather than a derivation that reduces to its own inputs. The load-bearing premise that comparators are truly unmemorized and well-matched is an empirical assumption (correctness risk), not circularity. With only the abstract available, no equation-level reduction of a claimed first-principles result to its inputs can be exhibited; score 0 is therefore the warranted finding.
Axiom & Free-Parameter Ledger
free parameters (2)
- chosen FPR / conformal threshold
- matching / sampling distribution for non-training sequences
axioms (3)
- domain assumption Non-training sequences cannot have been memorized and therefore supply a valid predictability baseline.
- standard math Conformal prediction yields valid FPR control under the sampling assumptions used for training and non-training populations.
- domain assumption A single matched non-training document is a sufficient census baseline for claiming memorization of a target document (e.g., a book).
read the original abstract
Recent work on extractable memorization in LLMs suffers from two contrasting validity problems. Some studies overstate extraction, e.g., relying on sequences too short to distinguish memorization from predictability. Others imply that extraction is unreliable evidence of memorization, since models can also reproduce real-world text they weren't explicitly trained on. In different ways, both overlook what makes a valid extraction claim: the model must generate a training sequence with high enough probability to indicate memorization. To determine what's high enough, one has to perform a matched comparison: measuring the generation probabilities of both the training sequences of interest and comparable non-training sequences. Because non-training sequences cannot have been memorized, their probabilities provide a baseline for predictability; a training sequence exceeding this baseline provides evidence of memorization. We formalize matched comparisons in two ways: (1) a conformal test that calibrates a threshold to a chosen FPR when training and non-training sequences are sampled from populations, and (2) a census that calibrates against a matched non-training document when the object is a single document (e.g., a book). We show that matched comparisons enable rigorous, calibrated memorization claims, and reveal where prior setups have validity issues. For instance, on Wikipedia OLMo 2 32B reproduces non-training 10-token suffixes roughly 24% as often as training ones: that share of the training generation rate reflects false positives, not memorization. For Llama 3.1 70B on books, the thresholds we calibrate are as low as 1e-27, supporting memorization claims for sequences that no feasible sampling budget would extract. Based on these results, we refine "extractable memorization" to require a valid memorization claim and near-certain generation within a realistic budget.
Forward citations
Cited by 1 Pith paper
-
Probabilistic "Copies" in Generative AI Models
An LLM is an infringing copy of a work only when the work can be extracted from it with relatively little effort, so some models are copies of some works and no model is a copy of everything it trained on.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.