Pith. sign in

Paper Citation Record · LEDGER

Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 59 inbound Pith citation observations for arXiv:2311.04850.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.04850 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:09:46.693633Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 26dde097-e60d-47b9-ba7b-9ca5f4315064 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.134008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:f1611a02aafb7ce8d5f9533e58daa7295b0171919aff1285589aa8fedff51dbf

Observation 79ffe552-61e5-4673-bf9b-bc4fde5114c9 · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 208

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:58:17.334977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:9c9b22fab3ecb4f0af8fbf913962add2c42918c4dfde1759100b115c359656f7

Observation d04fd98e-06bd-41e4-af68-e6b519a0ec47 · inbound

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools cites this paper.

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:08:09.653347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T08:08:09.444352Z digest=sha256:e450da62525f231967f3d9e531c8f44c9826f6d18afce5b1148eac7be3b59e26

Observation 3e02f0b5-fa3c-42a0-9cb9-bbfa84a9327a · inbound

Text2Cypher: Bridging Natural Language and Graph Databases cites this paper.

Text2Cypher: Bridging Natural Language and Graph Databases Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:26:04.675767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:26:04.675767Z digest=sha256:93a62b0d2ad44c1cd03e94aac757ed24cc8ed82d45c0b2f3e0b8c2d894ffe842

Observation df4d665a-8012-494f-a772-a0c13eb4b3b4 · inbound

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge cites this paper.

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T12:58:02.289913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:58:02.289913Z digest=sha256:594e849f3432adc1c997395fa9f5d8527fc4595c3ea7e4f1ded17af07bcd17e5

Observation 08f10b90-0c52-45e5-b59e-f67d6f0a9ac9 · inbound

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark cites this paper.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.383219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.383219Z digest=sha256:63e868ac8dc9b83e9345d4487cb73f979d4fa17691553a1a4543b0978ca7141d

Observation 1c213eeb-cfae-46b5-9236-bb1ad4c61d7e · inbound

Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation cites this paper.

Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:09:39.173471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:09:39.173471Z digest=sha256:e450b561588f10bf23ef24a7a4e16ba2c0eb117e7371e257fddbba7742d21bef

Observation ac7606fa-bfa5-4291-b7d0-361a84b6553a · inbound

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation cites this paper.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.141387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.141387Z digest=sha256:031149e83dd81e9b584db697bb56d3031685bb7d53e3a6b6a8cce50ef2e1687c

Observation f0d39c96-8140-4530-b269-887097046d2b · inbound

How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence cites this paper.

How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T18:11:35.042417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:11:35.042417Z digest=sha256:92654da37e8018391e5e7d1b2816ff96cf026037dd3499404ad4b9d1f116a80f

Observation 721c1b03-93f3-4b00-a4dc-c92cdd4ae55f · inbound

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation cites this paper.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.384462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.384462Z digest=sha256:94d47967731902fdd1aadd568047f63ce789543f1799b671cf98520b6bc1e21a

Observation 246859ee-fbaf-4caa-ba72-24d790e7d46a · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.489455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.489455Z digest=sha256:872c7da6dd01201c8dd3ddf6d476b7f36db695d0c1b77b30261f7ee6b30046b3

Observation 281802d4-4a9d-4fd0-ae13-ad09cbb00999 · inbound

POPri: Private Federated Learning using Preference-Optimized Synthetic Data cites this paper.

POPri: Private Federated Learning using Preference-Optimized Synthetic Data Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:09:46.693633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:09:46.693633Z digest=sha256:e6b099f6170e5a907a150f8add6fe9dcb5db21e1adfb538a1724771cabe05d40

Observation 270df203-488c-4990-a75c-be6caf8aa4e4 · inbound

The Leaderboard Illusion cites this paper.

The Leaderboard Illusion Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-16T05:22:55.708255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:22:55.708255Z digest=sha256:f29786019732921fdbbd202d8c5766395483e9384ed2d3eb22cdb66272c4803b

Observation 2808fab0-37a0-48f2-a689-51a99df2f1a0 · inbound

Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches cites this paper.

Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:58.004610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:58.004610Z digest=sha256:d8e2995765493f021e6d017e276a970c0775ecd51d944f6bcfe8546c178384a2

Observation efb6c51e-5550-44c5-afc0-714182503268 · inbound

EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated Code cites this paper.

EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated Code Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:06.238476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:06.238476Z digest=sha256:593bbc412939fbc8ea5b61d491ef56eb487daedd9e170a336b59015d7c02643b

Observation 775566c9-a666-4c3c-a592-902782fcb963 · inbound

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting cites this paper.

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:35.529969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:35.529969Z digest=sha256:f90c4cf2debcc8cf70a174ac7374fa9e53cd298b6fe86f81dec02185b36a9308

Observation fb1ad3e0-5e03-4ab1-aeff-d7c4f5f9f56d · inbound

SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation cites this paper.

SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:18.525978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:18.525978Z digest=sha256:0a5f962cd72ce5460d47eb1404151eaaebe7af565e8a8df62f30a43ed9bcd155

Observation 0a0588e0-0c37-4bd2-95fb-f1c7e168d36e · inbound

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation cites this paper.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.369531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.369531Z digest=sha256:ca6dcb6e3922050267f7dce43d69e57ba6de1edea85c75cc174716ef6520697d

Observation 3826a019-05ea-4d9a-9365-5c3523deaad8 · inbound

CodeMorph: Mitigating Data Leakage in Large Language Model Assessment cites this paper.

CodeMorph: Mitigating Data Leakage in Large Language Model Assessment Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:37.265620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:37.265620Z digest=sha256:b85f6128a70471a7b12cea5a4d98139349096d876a86eade1a67108b1eda2155

Observation b690f495-6cb2-4638-a9e5-a77112b77557 · inbound

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique cites this paper.

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:20.143746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:11:20.143746Z digest=sha256:e6cf23faa3fddb38154736a6cf393822694e6f239d4d2c9e8c1f3e0ad4c759dc

Observation ea4ec8ab-bfca-45cd-acfa-51f14490acd4 · inbound

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework cites this paper.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.799239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.799239Z digest=sha256:996871b649da5f11239010279becfb50bcc1653d4db5020e4081673d5bbb334e

Observation 1bb22ce2-c68b-422d-948d-2f1f77c32e39 · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.341453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:05782f8240be6bac2cd3645b3409887886b4baccaccc27707d466229b01ed8d9

Observation 8fca8ffb-2cd4-4adf-b168-b5219a677eb7 · inbound

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report cites this paper.

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T05:57:29.511258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:57:29.511258Z digest=sha256:2f10876328652e00693420924054b02aa94f58d463bec18b4bf907cb89b5cb00

Observation 081e16bc-34c9-4643-bd6c-44f61643b2e9 · inbound

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation cites this paper.

Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:28.295748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:28.295748Z digest=sha256:2a2a9d31c283c6552becdb2d5043c1cf1b5523ea66c9a1d3240a506e313a0c62

Observation 54c99d7c-0932-4665-88df-4cafaffd7561 · inbound

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories cites this paper.

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:36:28.943124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-18T14:32:57.426996Z digest=sha256:ea1ec53ca44a046501f1e6f70a46febe50f3abe77333225558ee5fc95787b1d1

Observation 1cf0bda5-0b6e-4ddb-b0ff-0793c0142615 · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:51.012398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:51.012398Z digest=sha256:d378fc0902cb128be983436e824bca6b2b0088f46a0db878a470bc9115c0c839

Observation d25e5c67-f333-4fda-b471-faf6e2b0bdff · inbound

Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation cites this paper.

Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:40:51.520975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:57:34.386632Z digest=sha256:0e5b1e2f90a1466be6baf589e57b8e1b5aea793a78212d69377d04a39731e3e1

Observation 548a2693-0ac5-442b-aac0-89d3bec954ea · inbound

DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math? cites this paper.

DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math? Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:04.832167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:46:24.780270Z digest=sha256:d1ef5ea643a3cedf1a406a32ebf641eeda2821168684707af4cc793d6b0ed23e

Observation 47e2479c-fb1f-496d-827e-eaaf883b2921 · inbound

TPS-CalcBench: A Benchmark and Diagnostic Evaluation Framework for LLM Analytical Calculation Competence in Hypersonic Thermal Protection System Engineering cites this paper.

TPS-CalcBench: A Benchmark and Diagnostic Evaluation Framework for LLM Analytical Calculation Competence in Hypersonic Thermal Protection System Engineering Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.514127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T05:25:59.343884Z digest=sha256:74691614518b879e72c5b0b72710c964d76ee40283e24a62df0d04d1166d906d

Observation 22b7cd4d-8090-460d-9758-5f01a02de39b · inbound

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation cites this paper.

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:17:06.819446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T03:00:26.137401Z digest=sha256:233ddf041e6f8e1fb7048ac6fc37310bdc5f6f3b278c518971272220727c1e15

Observation 6fd5b37e-9661-4f13-97f0-0af6e7d0b045 · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:36.959600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:d85d21ace499938d505bb2f309a025ee07dccc8902e0e616d6c7358c7971444a

Observation 1b15f2a6-e35a-4bde-a1ba-9bf71adfead8 · inbound

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games cites this paper.

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:17.276350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T17:06:32.814188Z digest=sha256:ed6bf54e756f0aa936cc46f5ca595eda8a9524cee6c76de45dab1775ca62cfb7

Observation 9d53cd78-ddd4-4c96-9571-3b76c9fa1010 · inbound

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity cites this paper.

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T22:04:18.090953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T10:23:02.697982Z digest=sha256:b4c67f39986723c2cb0b76b10213f7ce3921d611baae6bbdddfc565ed0c740e1

Observation 1c5fe167-2a6f-4ffc-b264-cb3a266a33fe · inbound

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations cites this paper.

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:59.899239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T00:57:55.616036Z digest=sha256:9c25e444ac040a7ca103fc3ed75b2c96aed74269f3a77f3d157dc9a045701104

Observation bcaeda91-673a-474a-8963-2564111db7eb · inbound

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations cites this paper.

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:07.976186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T23:42:00.965554Z digest=sha256:35ae5f6a3accc65716dab1e299fe9de375d41d8eef0bbacb6b9f780288ca808e

Observation ec4241e4-e7c0-4b4c-8c06-4660896cd370 · inbound

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation cites this paper.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:23.900869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:1513f487293627cb3e3681c090ccfa1e4f63f2d2cc600bdc2ae80050c6fb4efe

Observation abfb013e-e714-4037-85ac-acc881089016 · inbound

Decaf: Improving Neural Decompilation with Automatic Feedback and Search cites this paper.

Decaf: Improving Neural Decompilation with Automatic Feedback and Search Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:07:08.093098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T02:03:19.836506Z digest=sha256:0f119ccfb44e39b79f888c2916d71e98e76cf11aa46b961d40579dffcac9e616

Observation 481aa828-3a0e-498f-9f0b-f56df90e7622 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.862291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:90e37d2a71d44861221f36f2918cc33f4dbc3c6a1c035ac9d050c52065ce33e3

Observation a0f175ac-6597-48d7-9bf4-51fc3977f5ec · inbound

LLM Benchmark Datasets Should Be Contamination-Resistant cites this paper.

LLM Benchmark Datasets Should Be Contamination-Resistant Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:23:07.058655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T07:19:50.354875Z digest=sha256:a1eb65d46efd2a2dbd064ea80c5af917f91f32b128086c30df1996a7a117cc03

Observation ae61037a-4a0c-4149-9764-5047fc27a22e · inbound

Provable Joint Decontamination for Benchmarking Multiple Large Language Models cites this paper.

Provable Joint Decontamination for Benchmarking Multiple Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 172

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T00:44:29.489044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-22T00:40:54.038367Z digest=sha256:e4ada678bb76318133e3d7e6c19078d7f89b0ed308202d92b19788806defb632

Observation e50e38d5-a330-4082-9fc1-b47897ecc449 · inbound

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation cites this paper.

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:06:15.310194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-22T08:05:42.212459Z digest=sha256:ea206535e57c1bd6f3ada88587b70d0f1c65f94d7c6ddd0ddb09ec82a679abe0

Observation 81a8cf3c-06b8-49b8-bbd7-3da728bad548 · inbound

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness cites this paper.

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:24.444660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-25T04:37:01.537128Z digest=sha256:10c58352111c17a1cfe62bbcfb7dc18981102b3893d4946b6858ac4feb0096eb

Observation e11c3d40-017b-429b-8690-ff2370e45d98 · inbound

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs cites this paper.

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:47.895008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:32:05.976335Z digest=sha256:5bd940951c1296c8d1952d44cf8edbb628d92e0819c1666c8a002e1e627458f5

Observation 04ea9e21-1c59-4575-a510-f812fd2b7aeb · inbound

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild cites this paper.

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.155954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T14:41:07.354007Z digest=sha256:56cc512d981864cb7e052f5f11fb604b8a4b5e1d8763f98ef32c1af6405bac36

Observation 6c2dc137-5f7d-4787-b49a-71cbffd49ef8 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:34:40.588543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:27:50.367497Z digest=sha256:b08166692849e0006b2604beccc644e2b51c0e787a4e4443fdd5b53fb769958c

Observation b5a659b1-05df-4a48-8f38-913a67b3cfe8 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:05:31.914578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-01T07:35:51.017797Z digest=sha256:03b738672ac9ebec106b14bd42a90e66270318b8d0243ea2ec71bafe1e3b7046

Observation 73c90e48-132c-4c28-9ee1-a4949e9206f0 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:27:26.731890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-02T23:22:08.567841Z digest=sha256:29dc5b7632af98537ee0dfab5289108b3ddb87c6bff1e66413a5d12cfb17d650

Observation dab2454a-666d-4842-b07a-aed17edf3fd0 · inbound

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models cites this paper.

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:04:38.780577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T12:00:05.807004Z digest=sha256:5397e7b2646f9dc9f3409096bc712dfcd28e221a259ab1ca2206f7cd7f37ab9a

Observation 429ec9d5-66e8-4def-a403-8144313ea449 · inbound

Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction cites this paper.

Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.880656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T09:48:49.786021Z digest=sha256:edd857d203a2c1aa7088856308a4a231f4f1e1501c0533d8dafdb37ed081ffb3

Observation dcbed4ea-62a8-48a1-b727-be0d3cb6004e · inbound

Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs cites this paper.

Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:37:56.593959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T09:57:14.328157Z digest=sha256:7c50afdb981179f6eb79fc2fc1e4350a97299c438e77bc2211f349d72ede546c

Observation df54faea-0bd2-4570-a54c-c347743948f7 · inbound

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models cites this paper.

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:24:18.367339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T06:20:02.046020Z digest=sha256:e2c3bea913d2c05ae13d187e0a8294da484bce8e2243643e81bf90931f92f64b

Observation e34d5daa-ff82-4dcf-8c9b-19274d08c5d6 · inbound

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI cites this paper.

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T05:03:42.877641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T05:03:42.877641Z digest=sha256:4c2e5d7b2506fea05d9c7077b1f74bcf6195c590ca4923c504b0060a9d6de428

Observation d05c235d-3146-48aa-b346-a37b6f91e5f6 · inbound

Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit Tests cites this paper.

Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit Tests Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T15:32:18.352655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:32:18.352655Z digest=sha256:00706e24bcfe902a1259a8436585bc9f86245d306f20040cdbb9018e2e735300

Observation 372bad00-77da-4547-981f-6351c5a94418 · inbound

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores cites this paper.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.418737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.418737Z digest=sha256:77124d485b02a0a71895aed0aa8f37af24b89a31165ff06488f21986c9c4b5db

Observation f312928c-1d49-415a-bc1e-92f5c4a57e6c · inbound

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks cites this paper.

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:05.614586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:05.614586Z digest=sha256:b65cdc20018ed9ff9664d190572a7bfe2b42a45fbda3760eeee4cc3561338bc6

Observation 1fd6bd95-accd-4fa4-b360-c9160579218f · inbound

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks cites this paper.

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T18:19:46.939354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:19:46.939354Z digest=sha256:f74c42d703f185159c018b23db65d734acee12d00a986f2157d34e5cd4c77e3c

Observation b58d0bba-0fb2-4354-b304-3f047d534dc5 · inbound

Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora? cites this paper.

Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora? Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:06.247862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:28:06.247862Z digest=sha256:3cdf59e23c756604d497cd0e09fdd43037d5ff5dc085cfcdd46ddd141f3bbbc8

Observation bc44b2f9-4fc9-4d31-8b5f-1ca9a405b30f · inbound

Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection cites this paper.

Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T00:08:49.091689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:08:49.091689Z digest=sha256:883dcd5ac030eed90604526f99c578d7bee9ba341c6228329e33db55941771fa

Observation d1f3c95e-7317-4c03-a922-21be04d45f4c · inbound

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure cites this paper.

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:12.482441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:12.482441Z digest=sha256:7d6874d4c35d404173d165ee0f64de37d2e33fcb067d793611a1898852bf3d82