Pith. sign in

Paper Citation Record · LEDGER

Investigating Data Contamination for Pre-training Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2401.06059.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.06059 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:34.145164Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T21:17:24.167881Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 89a0abb8-07bf-4c42-a787-5cfdcd8e09cd · inbound

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting cites this paper.

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting Investigating Data Contamination for Pre-training Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:34.145164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:34.145164Z digest=sha256:5cd66890b4dc15d66c2c516d28ce581d35bc036e01b9d926fffa268f5817b6c3

Observation b460b024-0bbc-47a7-b8f4-f7159f6800f8 · inbound

Can Vision Language Models Understand Mimed Actions? cites this paper.

Can Vision Language Models Understand Mimed Actions? Investigating Data Contamination for Pre-training Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:41.260222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:23:41.260222Z digest=sha256:b332f9298588808a8f45373e8f3976f270f320e3c6022627c89f8b443a17f90f

Observation 367caea1-da5b-4b5a-98a1-cae7a28bf80c · inbound

Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data cites this paper.

Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data Investigating Data Contamination for Pre-training Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:23.208794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:23.208794Z digest=sha256:350f9f10e01f2a34ae8232e9c7ff9b544e83e7c3ad890c18b805fff6959e423c

Observation ede787cc-e818-4db0-be20-9d383381e5f6 · inbound

Investigating Training Data Detection in AI Coders cites this paper.

Investigating Training Data Detection in AI Coders Investigating Data Contamination for Pre-training Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:53.769972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:53.769972Z digest=sha256:ae87d914d1b96793996bcb3f8c65d789f48326b1d13b4868b983821f81334bcc

Observation bd36c6b1-c8ed-4339-b358-6869b3900492 · inbound

Dataset Watermarking for Closed LLMs with Provable Detection cites this paper.

Dataset Watermarking for Closed LLMs with Provable Detection Investigating Data Contamination for Pre-training Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:57.277115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:53:42.185498Z digest=sha256:23c240e0006778db07a426b89417bb94fb8aeac92d8bb244c4f9126976e48f49

Observation b54ef219-da74-42cc-9569-2fe133ae6f25 · inbound

LLM Benchmark Datasets Should Be Contamination-Resistant cites this paper.

LLM Benchmark Datasets Should Be Contamination-Resistant Investigating Data Contamination for Pre-training Language Models

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:23:07.050056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T07:19:50.354875Z digest=sha256:571eafcffdc12c9e7b5c1e49f2094fda601d07e6bcb8cca5c2ca910664593119

Observation 5377c182-e0ca-4f62-8394-9bcfe8a2581a · inbound

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness cites this paper.

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness Investigating Data Contamination for Pre-training Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:24.421044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T04:37:01.537128Z digest=sha256:38a63486806bc568a281d608e61be9c463f0f52dfd20f1ad342105d5535986fd

Observation 20189093-0b97-45f4-8f59-facd5b977006 · inbound

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models cites this paper.

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models Investigating Data Contamination for Pre-training Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:04:38.792275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T12:00:05.807004Z digest=sha256:3fb9dc6e53fd74cd8b54527d967e8f3f52be16c93fd39a5d270600aa66d14208

Observation f0e50035-da39-45b2-a795-5de66d88e973 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Investigating Data Contamination for Pre-training Language Models

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.570680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:bae0bd71efb566d5aa6fcce38831dcfcd66d2617f43d4495214e6c7195eebd83

Observation d690820f-a685-4ea9-92b7-ecfc5e511bf7 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Investigating Data Contamination for Pre-training Language Models

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.169298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:8623fc108e0b68e3d9e9134cc6d87f7404339aa2c0babccbfc523117d13da311