Pith. sign in

REVIEW 4 cited by

Proving Test Set Contamination in Black Box Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.17623 v2 pith:DNOH7Z4V submitted 2023-10-26 cs.CL cs.LG

classification cs.CLcs.LG
keywords contaminationmodelstestlanguagedatapretrainingaccessiblebenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models are trained on vast amounts of internet data, prompting concerns and speculation that they have memorized public benchmarks. Going from speculation to proof of contamination is challenging, as the pretraining data used by proprietary models are often not publicly accessible. We show that it is possible to provide provable guarantees of test set contamination in language models without access to pretraining data or model weights. Our approach leverages the fact that when there is no data contamination, all orderings of an exchangeable benchmark should be equally likely. In contrast, the tendency for language models to memorize example order means that a contaminated language model will find certain canonical orderings to be much more likely than others. Our test flags potential contamination whenever the likelihood of a canonically ordered benchmark dataset is significantly higher than the likelihood after shuffling the examples. We demonstrate that our procedure is sensitive enough to reliably prove test set contamination in challenging situations, including models as small as 1.4 billion parameters, on small test sets of only 1000 examples, and datasets that appear only a few times in the pretraining corpus. Using our test, we audit five popular publicly accessible language models for test set contamination and find little evidence for pervasive contamination.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 5 citations worldwide. Full citation record

  1. Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores

    cs.LG 2026-08 conditional novelty 6.0 of 10

    The standard pre/post cutoff check cannot separate memorization from recency, and a single external reference is needed to measure and adjust for temporal leakage.

  2. AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI

    cs.LG 2025-07 reject novelty 6.0 of 10

    AnalogFed combines federated learning with a generative analog-topology model, adding dummy-token input perturbation and partial homomorphic encryption to resist membership inference and model inversion attacks.

  3. Spectral Journey: How Transformers Predict the Shortest Path

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Two-layer transformers learn shortest paths on small graphs by building embeddings that correlate with spectral decomposition of the line graph, yielding an approximate spectral path-finding algorithm.

  4. Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning

    cs.CR 2025-06

Pith tools