Pith. sign in

Paper Citation Record · LEDGER

Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2411.03923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.03923 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:31:04.401902Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:16:58.130004Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bb55b6ed-45c0-445b-a95d-ebfd06e968c7 · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:48:26.481297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:bde608fa121ebf876427a2094d01035d82b84ca8c3b50c25de0cf390f0bff8ed

Observation 9d1eb138-e830-4a39-bfb2-be530c42a6d8 · inbound

StockSim: A Dual-Mode Order-Level Simulator for Evaluating Multi-Agent LLMs in Financial Markets cites this paper.

StockSim: A Dual-Mode Order-Level Simulator for Evaluating Multi-Agent LLMs in Financial Markets Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:26.695956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:26.695956Z digest=sha256:900d4231e7ae43f45450e06970ef1673403c5f75e13e961963a2f90386506a9e

Observation e2a72f36-9d7c-4e11-a269-6549549feb13 · inbound

Position: AI Evaluations Should be Grounded on a Theory of Capability cites this paper.

Position: AI Evaluations Should be Grounded on a Theory of Capability Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:00:41.475033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T21:57:55.834632Z digest=sha256:de6ef52aaa71524607a0d0a1e148b31e7fdddee4c642e1813285b00a28106871

Observation ca0e5472-5a92-4f1e-80c4-6a08e2760cb1 · inbound

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks cites this paper.

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:29.688964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T04:27:11.735657Z digest=sha256:7764c200a9683e33059d081f7b31ae499236e743a8e293b2082dc451a0c66f7b

Observation 60f42c4b-1ba6-476c-a6b6-5419dd0d799f · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:05:36.930166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:6bc305ffb8c6976e7dba97d62defee97919a17e5becf0f1f90c2cef10f5d5045

Observation 09cbb8ab-6e55-4a22-91c5-76b151197650 · inbound

Dataset Watermarking for Closed LLMs with Provable Detection cites this paper.

Dataset Watermarking for Closed LLMs with Provable Detection Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:57.253215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T00:53:42.185498Z digest=sha256:14d05803b5c9ec9ff43771e634b3d357d069ad3bc7d03958ac0cc453714db8e5

Observation c3d3fcce-b28a-408a-bbfd-68acd20b10e8 · inbound

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild cites this paper.

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.153124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T14:41:07.354007Z digest=sha256:113009c36bf572a97230146894c80593da612d28818bb2f6a2fe661d3c8671d1

Observation a8712fb1-cb99-4579-95f3-03c93b80a763 · inbound

Amplifying Membership Signal Through Chained Regeneration cites this paper.

Amplifying Membership Signal Through Chained Regeneration Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:39.288910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T06:22:08.140403Z digest=sha256:45e53c95ad07af9996797c5cfb294acce635fa1098448d00e62120e351f806b0

Observation fd18870b-ebe7-4d86-91f6-677dc74517e0 · inbound

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages cites this paper.

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.131717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-02T13:14:15.684960Z digest=sha256:beeb73f3170fb1f00ab67dfce982f04c0730a9b0ef125bbe97f2d6fe3341043f

Observation e9a46123-ac97-4ffe-a532-9b4ac9c3f628 · inbound

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy cites this paper.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.669375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.669375Z digest=sha256:4b26c20a200e5c385eba816d68f263b9bd21c256374a29206524a57f91388c5e

Observation c0328c06-8a03-4415-88fd-e754b03f040c · inbound

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores cites this paper.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.401902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.401902Z digest=sha256:559968bc17be1b0d6f67b77f18ad59fe3263b1795d502a6a312864fdcd23cb25