Pith. sign in

Paper Citation Record · LEDGER

Investigating Data Contamination in Modern Benchmarks for Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2311.09783.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.09783 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:42:38.666546Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ec73b96e-3e48-43e5-844d-8cc1f4d85384 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:40.906325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:215f948292342d4cfc8ae078af10e4fc11bb2326a22663b0fc4f234091df1318

Observation 233cf14c-0080-43bf-9aba-8f16ba7b6c33 · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:48:26.394739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:9fb37bbcb6ba508339b9fd17749c3661062666b02f00879a37d1ea78f808ffff

Observation 0e558609-c205-45a7-9d6b-e7a4d63507bd · inbound

PRIMETIME : Limits of LLMs in Temporal Primitives cites this paper.

PRIMETIME : Limits of LLMs in Temporal Primitives Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:36:59.121794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T18:36:48.376877Z digest=sha256:e93cf44db79ad95800ddc8726fbb80273872f04d0d8f2197453079fd2a8e7c08

Observation ebacc265-7c29-422e-9e7c-1ed6523700ef · inbound

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality cites this paper.

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:38.666546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:38.666546Z digest=sha256:bf4895f20c2cfad9855de494a58f983c33ca41319957eb29d55ffaf018617c5c

Observation eb8820db-3ad5-485b-a6b4-cb0519f1b63c · inbound

SciDA: Scientific Dynamic Assessor of LLMs cites this paper.

SciDA: Scientific Dynamic Assessor of LLMs Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:17.481039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:17.481039Z digest=sha256:47ca4b129b99ac074e40d7074f66669841512d52ac4fb6299af718689b04de99

Observation c1a274ae-e16c-4f1c-b373-528e1058307a · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.291053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:320aa1fba8c39eca951ea019bdf2673cd7de3b72bcb126a39f9577d870f8c153

Observation 6a902d75-6f93-4099-86b0-e9a87819967a · inbound

ZoFia: Zero-Shot Fake News Detection with Entity-Guided Retrieval and Multi-LLM Interaction cites this paper.

ZoFia: Zero-Shot Fake News Detection with Entity-Guided Retrieval and Multi-LLM Interaction Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:50:38.285574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T01:47:35.232468Z digest=sha256:9a6b1b28a3acd6d2b8d1530405e0f836cc65ea09eb3c0ad7d2da77309ef42713

Observation 91ac6d3a-7590-46fd-9add-094893313853 · inbound

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks cites this paper.

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:05.144360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:35:24.397273Z digest=sha256:b4aab06339a2215e7de5bc5c41e858418eb994f40507bbc24c0643029c08edf1

Observation 3d6b53c8-ead6-415c-a590-edb84a71c7a6 · inbound

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation cites this paper.

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:14.401063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T01:41:42.003483Z digest=sha256:ba8e0b985fa1679eba0f4184da6a22952443a1e53a08bfe8ec5632202d65e82e

Observation 4a340f5a-cdf5-44da-8454-28278d68da5a · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:24:56.558582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:1bc8d00298f70ee27895630d85f22d5592c69b233496def4bd85caf648b5f15f

Observation 4b27821f-d48b-413e-bfa5-d9f344fe198b · inbound

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection cites this paper.

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:36:30.302577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T09:48:09.688745Z digest=sha256:acd55ad7d54127e5a7bc2edc70484b55412fcf002e023d986f4c42728b4e77d0

Observation 7dfa333e-46d4-421f-90ba-f63ecc16cedb · inbound

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection cites this paper.

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:39:16.485829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-04T00:30:13.665405Z digest=sha256:1266307d12bbdb8d03f6021e19bc29e67ff7633e51a06745862a5acc40c23cf6

Observation 86763adc-0a71-42c0-969b-9498eb69b3d8 · inbound

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier cites this paper.

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:57:48.182085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:36:09.211639Z digest=sha256:9b4f25570a55d686157dff2f95792203d7cde98adb11afe056f1e6b7310c9f36

Observation b65907d7-b88a-45de-9d53-32fb3f9e9b5d · inbound

SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks cites this paper.

SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T18:37:16.481985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-02T18:19:43.146102Z digest=sha256:6ffdde0ba2bd9e5eb7d2de8febfdafc2da39713f5729c02a11bdb449c2d2caee

Observation 832cdfff-61ee-400e-9c61-dcf0b491a644 · inbound

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains cites this paper.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.374252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.374252Z digest=sha256:30987bd5db4e166f1a58d35c71a27ca28f960fcd77c5901304c9e590d14185d2