Pith. sign in

Paper Citation Record · LEDGER

Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2402.03927.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.03927 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:42:08.081999Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

18
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 11011c9c-31f9-4e97-8420-c2beeba8b61b · inbound

The Impossible Test: A 2024 Unsolvable Dataset and A Chance for an AGI Quiz cites this paper.

The Impossible Test: A 2024 Unsolvable Dataset and A Chance for an AGI Quiz Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T17:00:12.341990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:00:12.341990Z digest=sha256:e6184ec036f8bfc0bed0af7434468a2787fb9753fc5fec40d3cd90d9877bb2ad

Observation c8023431-371b-4128-96f9-6d4dd42ca5d2 · inbound

Engineering AI Judge Systems cites this paper.

Engineering AI Judge Systems Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.194285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.194285Z digest=sha256:f766f4f1ce64a11202d04b4437a7fe71d53664ff7bbb5afe22fe45e2dce3b350

Observation bfaadb59-3516-481a-a599-dc4e5d82bdab · inbound

Large Language Models show both individual and collective creativity comparable to humans cites this paper.

Large Language Models show both individual and collective creativity comparable to humans Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:46:58.224749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:46:58.224749Z digest=sha256:94cc1481e0c97d390ed8d8b5b89e4451794a3b5d1b1eed5020a7c5df7d40192a

Observation 0dd69ec0-e528-408e-954f-d2ee398554f5 · inbound

The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance? cites this paper.

The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance? Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:55.381433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:55.381433Z digest=sha256:65ea27a41f6b3df4100aa57c75885d84f364d447e9f876e7ccab8beede7640f9

Observation bf6a3346-154b-4b3c-bf07-28ab8783b1b8 · inbound

Political-LLM: Large Language Models in Political Science cites this paper.

Political-LLM: Large Language Models in Political Science Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-11T19:52:04.657165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:52:04.657165Z digest=sha256:d6cc30597e13d285b8df923a0822e3db9c30a5417302501e2f117c9c57d54bec

Observation fe8ad168-8f8b-4b7f-a8bc-f3fbbb9668b1 · inbound

Dynamic Skill Adaptation for Large Language Models cites this paper.

Dynamic Skill Adaptation for Large Language Models Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:46:46.022253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:46:46.022253Z digest=sha256:94ac0b807cccb8dc7e606e2ff26ba3df95d72538e2a73d5ab0fcfeeb3dc9a897

Observation 7cafa042-47d3-4793-8996-cf0500754b4d · inbound

Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media Platforms cites this paper.

Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media Platforms Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T16:25:32.874279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:25:32.874279Z digest=sha256:165293c57311579bed4aabc42f246c98428cf31cf475c7ac6264d9fb9c2fccea

Observation aa10e7b4-9a06-4dc0-b1da-7b7e75627246 · inbound

LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks cites this paper.

LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T16:26:32.043593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:26:32.043593Z digest=sha256:54cfb3346c16bf1dc411e2e2b5bd0547bc43994ef4a2fcde2832f76e37f07d01

Observation 921a5f53-8fbf-46ae-bcb6-104c57184294 · inbound

AIMS.au: A Dataset for the Analysis of Modern Slavery Countermeasures in Corporate Statements cites this paper.

AIMS.au: A Dataset for the Analysis of Modern Slavery Countermeasures in Corporate Statements Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T14:04:30.135125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:04:30.135125Z digest=sha256:ba3692d66dffb6228e023329d9d6bd5628df28c85c031b83110ce61e3e11675a

Observation c2aeb9ec-465f-4ccf-9404-aee96c36798b · inbound

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation cites this paper.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.081999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.081999Z digest=sha256:35117020bc7498915d57319f7fe5d7ec241fe727ffa0a38f307cfb2671c6aacc

Observation ddc01214-47f7-460b-9fbd-5013d2f2b2e6 · inbound

Emergent LLM behaviors are observationally equivalent to data leakage cites this paper.

Emergent LLM behaviors are observationally equivalent to data leakage Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:14.101878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:14.101878Z digest=sha256:ab6b14abed6fcbd548de65c9fb6d409a5bc357caff734273732d1bf2ef185439

Observation 75c04252-af6e-42d3-a26c-52a51f0b002d · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:20:34.441271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T01:18:44.523602Z digest=sha256:c9b8b419cca8676d6a7a86492cce93e1a0b6a720a66f35f8f6af65918db66425

Observation 062d1044-d54a-40c0-a1ca-11441b92821b · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T00:11:50.195992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:11:50.195992Z digest=sha256:f07dffa8283ae3738fcb453f3a837d5a6d5e9bcad7f0925fed106e19b63f2e47

Observation b55e2365-36b5-4ea6-90cc-8c6fa82d3493 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.114913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.114913Z digest=sha256:1703d65d487bb23d38b771f4118e29d7c088e35ec67a463742499b65c893fdf2

Observation bfaa887b-ade1-4646-a0a7-6f870dfecc1a · inbound

Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework cites this paper.

Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:55:44.564388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:53:12.473010Z digest=sha256:eb3d9654cd88b9bd86119f08a745a6c7461ddb24ae88c1a4abda417f43746d0a

Observation 25c2bb43-9552-4f13-828e-94a4ddb9dc64 · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:14.113824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:c74adc94cd7615e4b3f8764fa3cf95e5904908f8faddacf0a611e6a05a85c1bc

Observation 4419564c-2605-4cd4-bf4d-3f18199ce9b6 · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:24:56.552879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:2a3f0d383e0ea76fd30a79f39974a68d8d1f2a43a11138f276960c78b1a2c275

Observation e555ed24-e24f-4bbe-9cde-1e3c834064bc · inbound

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection cites this paper.

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:36:30.289844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T09:48:09.688745Z digest=sha256:2779a8039b78d5d323da009bd1c498ccc47b85e004dd1622c2428870d8407f18

Observation bbf8c6c6-a404-471a-bc98-cb11d8e6cc14 · inbound

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection cites this paper.

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:39:16.467562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-04T00:30:13.665405Z digest=sha256:f3ebd13f22498c07b52d7a161d01e5276d730a794d1dc7089881595ede0cdca0