Pith. sign in

REVIEW 3 cited by

Investigating Data Contamination for Pre-training Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.06059 v1 pith:WML3K5MZ submitted 2024-01-11 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords contaminationdatapre-trainingtextitcapabilitiesdownstreamevaluationlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models pre-trained on web-scale corpora demonstrate impressive capabilities on diverse downstream tasks. However, there is increasing concern whether such capabilities might arise from evaluation datasets being included in the pre-training corpus -- a phenomenon known as \textit{data contamination} -- in a manner that artificially increases performance. There has been little understanding of how this potential contamination might influence LMs' performance on downstream tasks. In this paper, we explore the impact of data contamination at the pre-training stage by pre-training a series of GPT-2 models \textit{from scratch}. We highlight the effect of both text contamination (\textit{i.e.}\ input text of the evaluation samples) and ground-truth contamination (\textit{i.e.}\ the prompts asked on the input and the desired outputs) from evaluation data. We also investigate the effects of repeating contamination for various downstream tasks. Additionally, we examine the prevailing n-gram-based definitions of contamination within current LLM reports, pinpointing their limitations and inadequacy. Our findings offer new insights into data contamination's effects on language model capabilities and underscore the need for independent, comprehensive contamination assessments in LLM studies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can Vision Language Models Understand Mimed Actions?

    cs.CL 2025-06 conditional novelty 7.0 of 10

    Vision-language models identify real actions with context far better than they identify mimed actions performed by 3D avatars, while humans are equally accurate on both.

  2. Investigating Training Data Detection in AI Coders

    cs.SE 2025-07 conditional novelty 6.0 of 10

    Most existing training-data detection methods perform poorly on code, while prefix-relative method ReCaLL consistently scores highest, though all degrade under code mutations.

  3. Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Continuing the pre-training of TabPFN on 71 curated real-world tables raises its average normalized ROC-AUC from 0.954 to 0.976 on 29 AutoML Benchmark datasets.

Pith tools