Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2311.04850.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T11:19:51.012398Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 26dde097-e60d-47b9-ba7b-9ca5f4315064 · inbound
Benchmark Data Contamination of Large Language Models: A Survey Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 170
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 79ffe552-61e5-4673-bf9b-bc4fde5114c9 · inbound
DataComp-LM: In search of the next generation of training sets for language models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 208
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d04fd98e-06bd-41e4-af68-e6b519a0ec47 · inbound
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1bb22ce2-c68b-422d-948d-2f1f77c32e39 · inbound
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 54c99d7c-0932-4665-88df-4cafaffd7561 · inbound
LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1cf0bda5-0b6e-4ddb-b0ff-0793c0142615 · inbound
Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d25e5c67-f333-4fda-b471-faf6e2b0bdff · inbound
Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 548a2693-0ac5-442b-aac0-89d3bec954ea · inbound
DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math? Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 47e2479c-fb1f-496d-827e-eaaf883b2921 · inbound
TPS-CalcBench: A Benchmark and Diagnostic Evaluation Framework for LLM Analytical Calculation Competence in Hypersonic Thermal Protection System Engineering Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 22b7cd4d-8090-460d-9758-5f01a02de39b · inbound
When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6fd5b37e-9661-4f13-97f0-0af6e7d0b045 · inbound
Measuring AI Reasoning: A Guide for Researchers Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b15f2a6-e35a-4bde-a1ba-9bf71adfead8 · inbound
Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d53cd78-ddd4-4c96-9571-3b76c9fa1010 · inbound
Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1c5fe167-2a6f-4ffc-b264-cb3a266a33fe · inbound
GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bcaeda91-673a-474a-8963-2564111db7eb · inbound
GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ec4241e4-e7c0-4b4c-8c06-4660896cd370 · inbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation abfb013e-e714-4037-85ac-acc881089016 · inbound
Decaf: Improving Neural Decompilation with Automatic Feedback and Search Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 481aa828-3a0e-498f-9f0b-f56df90e7622 · inbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a0f175ac-6597-48d7-9bf4-51fc3977f5ec · inbound
LLM Benchmark Datasets Should Be Contamination-Resistant Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ae61037a-4a0c-4149-9764-5047fc27a22e · inbound
Provable Joint Decontamination for Benchmarking Multiple Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 172
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e50e38d5-a330-4082-9fc1-b47897ecc449 · inbound
The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81a8cf3c-06b8-49b8-bbd7-3da728bad548 · inbound
How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e11c3d40-017b-429b-8690-ff2370e45d98 · inbound
TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04ea9e21-1c59-4575-a510-f812fd2b7aeb · inbound
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6c2dc137-5f7d-4787-b49a-71cbffd49ef8 · inbound
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b5a659b1-05df-4a48-8f38-913a67b3cfe8 · inbound
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 73c90e48-132c-4c28-9ee1-a4949e9206f0 · inbound
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dab2454a-666d-4842-b07a-aed17edf3fd0 · inbound
TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 429ec9d5-66e8-4def-a403-8144313ea449 · inbound
Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dcbed4ea-62a8-48a1-b727-be0d3cb6004e · inbound
Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df54faea-0bd2-4570-a54c-c347743948f7 · inbound
SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e34d5daa-ff82-4dcf-8c9b-19274d08c5d6 · inbound
Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.