Pith. sign in

REVIEW 3 cited by

Data Contamination Through the Lens of Time

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10628 v1 pith:B2W5U7HK submitted 2023-10-16 cs.CL

classification cs.CL
keywords datacontaminationllmsbenchmarksmodelstrainingevaluatingpublicly
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent claims about the impressive abilities of large language models (LLMs) are often supported by evaluating publicly available benchmarks. Since LLMs train on wide swaths of the internet, this practice raises concerns of data contamination, i.e., evaluating on examples that are explicitly or implicitly included in the training data. Data contamination remains notoriously challenging to measure and mitigate, even with partial attempts like controlled experimentation of training data, canary strings, or embedding similarities. In this work, we conduct the first thorough longitudinal analysis of data contamination in LLMs by using the natural experiment of training cutoffs in GPT models to look at benchmarks released over time. Specifically, we consider two code/mathematical problem-solving datasets, Codeforces and Project Euler, and find statistically significant trends among LLM pass rate vs. GitHub popularity and release date that provide strong evidence of contamination. By open-sourcing our dataset, raw results, and evaluation framework, our work paves the way for rigorous analyses of data contamination in modern models. We conclude with a discussion of best practices and future steps for publicly releasing benchmarks in the age of LLMs that train on webscale data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CurveShift: Is Agent Progress Scalar? Separating Level from Shape

    cs.CL 2026-07 conditional novelty 7.0 of 10

    A single rising-ability model explains most apparent AI gains on hard tasks; a smaller real hard-item gain remains in no-scaffold competitive programming.

  2. TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories

    cs.SE 2025-07 conditional novelty 6.0 of 10

    LLMs achieve roughly 0.80 TypeSim similarity to human type annotations but far more mypy consistency errors than a coherent system should have, on a new 50-repo benchmark.

  3. LLM Performance for Code Generation on Noisy Tasks

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LLMs solve heavily obfuscated benchmark tasks, and performance decay under obfuscation differs sharply between old and new datasets, which the authors interpret as a signature of training-data contamination.

Pith tools