Pith. sign in

Paper Citation Record · LEDGER

Investigating Data Contamination in Modern Benchmarks for Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2311.09783.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.09783 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:42:38.666546Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ec73b96e-3e48-43e5-844d-8cc1f4d85384 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:40.906325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:d92d3299e365c6f0934de09270fa80d5ac0b0d0b250e508cd45f67d9eabc5ce3

Observation 233cf14c-0080-43bf-9aba-8f16ba7b6c33 · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:48:26.394739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:1a6510afbbf057239b77177ec3cc26e0458b1019ad46452d134f6c7bbb7c8824

Observation 0e558609-c205-45a7-9d6b-e7a4d63507bd · inbound

PRIMETIME : Limits of LLMs in Temporal Primitives cites this paper.

PRIMETIME : Limits of LLMs in Temporal Primitives Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:36:59.121794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T18:36:48.376877Z digest=sha256:89147d520f35a2c9b2d8721e04954ad0b916ec506f6e69e94cbf40092fc4833e

Observation ebacc265-7c29-422e-9e7c-1ed6523700ef · inbound

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality cites this paper.

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:38.666546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:38.666546Z digest=sha256:5a0c424f50a3941e9e20507e48095305e519668993bb632e36d1a18985a470fa

Observation eb8820db-3ad5-485b-a6b4-cb0519f1b63c · inbound

SciDA: Scientific Dynamic Assessor of LLMs cites this paper.

SciDA: Scientific Dynamic Assessor of LLMs Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:17.481039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:17.481039Z digest=sha256:e7e8080cfcebbc8c83422da08ca51b997ed34f62078bb9e41a52a8a3c8cde4e3

Observation c1a274ae-e16c-4f1c-b373-528e1058307a · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.291053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:0578de9c89902d8907a6529b728d0dc94db18f807b722cd6ed4fff5aee74e117

Observation 6a902d75-6f93-4099-86b0-e9a87819967a · inbound

ZoFia: Zero-Shot Fake News Detection with Entity-Guided Retrieval and Multi-LLM Interaction cites this paper.

ZoFia: Zero-Shot Fake News Detection with Entity-Guided Retrieval and Multi-LLM Interaction Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:50:38.285574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T01:47:35.232468Z digest=sha256:9d2e55f8b82890e6f3d7c860aab361d6aef43ed3cc94acaebfafc4dbac645a18

Observation 91ac6d3a-7590-46fd-9add-094893313853 · inbound

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks cites this paper.

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:05.144360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:35:24.397273Z digest=sha256:37361c28832e626096957f0a8b10c6f32fc721ef4b1466e82de2179f7f900195

Observation 3d6b53c8-ead6-415c-a590-edb84a71c7a6 · inbound

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation cites this paper.

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:14.401063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T01:41:42.003483Z digest=sha256:2477cfed12ddbe1e394ed9f081cab1d01cc0331be6f8830519a96397e13109ac

Observation 4a340f5a-cdf5-44da-8454-28278d68da5a · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:24:56.558582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:6d85c68fd15cfb1270265fe276efd9560b1d49cd58755d6de4b5889b1e4d82de

Observation 4b27821f-d48b-413e-bfa5-d9f344fe198b · inbound

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection cites this paper.

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:36:30.302577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T09:48:09.688745Z digest=sha256:e0b40e018dee3542465b7d79d980c8606f198640d5c817726ab5c3eb0450510b

Observation 7dfa333e-46d4-421f-90ba-f63ecc16cedb · inbound

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection cites this paper.

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:39:16.485829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-04T00:30:13.665405Z digest=sha256:ef4568a303bfe58d9926a42e656149750da78d7fbfd1c87b28db049d71c19384

Observation 86763adc-0a71-42c0-969b-9498eb69b3d8 · inbound

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier cites this paper.

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:57:48.182085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:36:09.211639Z digest=sha256:b19f3b167c699428d15228a77f242b4ee754e1bfda28d76c54fb1d2b9751c9e9

Observation b65907d7-b88a-45de-9d53-32fb3f9e9b5d · inbound

SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks cites this paper.

SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T18:37:16.481985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-02T18:19:43.146102Z digest=sha256:054d579dc1286825944329e00674732e09aa543d949731a85caf10954e289bfb

Observation 832cdfff-61ee-400e-9c61-dcf0b491a644 · inbound

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains cites this paper.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.374252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.374252Z digest=sha256:86c5c3c40909766c7c7e4f919892b20524748f85a58c103817a9dfc69cf8b29b