Pith. sign in

Paper Citation Record · LEDGER

Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2305.10160.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.10160 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:01:31.426943Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 37f0bed8-eb3f-4f69-b81e-727a2200282e · inbound

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive cites this paper.

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 124

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T23:04:44.484284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-17T23:04:44.287660Z digest=sha256:1708a0ebc6f4a382a64955a28cdd96241b1c26e823757ba01d497344cc4c39c4

Observation 0f038c9f-2969-4ee3-bd78-e5e79d30df6e · inbound

BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices cites this paper.

BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T17:01:31.426943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:01:31.426943Z digest=sha256:9089ab41f3b968d58d5dd2ea56ad7781201067efc77f0ec05f53e39c86cda598

Observation c2402e30-e15d-44c7-95eb-752fd06122a3 · inbound

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge cites this paper.

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:58:02.061334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:58:02.061334Z digest=sha256:3343e7bb193a0955f5215e1abdf1f69c3def3897d9b6e6a95e11b4c7b656f24f

Observation def9805a-17ce-494c-bcb4-8975ff197ef5 · inbound

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis cites this paper.

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:26.282453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:26.282453Z digest=sha256:d3ff30f37d8077128f7bc2550c22089cb81d88c5810022a2acb4d47857cfac37

Observation 8948bcd2-cb98-4b5b-a743-8de1248f04a9 · inbound

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks cites this paper.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:07.979885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:8c0848be790e7526ea01182552878e15ac255b7c33d6294072a3753d6de93ffe

Observation 1b7beccb-dbfb-4fc2-913d-02deb0567a92 · inbound

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications cites this paper.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.230145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.230145Z digest=sha256:b91ee0934325932c103dab53a19b2780a03a9a77e49c7272eeb8d6af96e3d112

Observation 6dfe038a-0acd-4db3-ba54-d5683f83287b · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.721746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.721746Z digest=sha256:0d04003468b3af06b803f6ccbd88ab04b3b3d2f61680fe497954ad327d55a8f2

Observation cff115f9-91d8-4e04-92c3-93e4e40e7104 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.943868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:584386c4e1b891f275fb057d26a8705c88dcc7fb336003c8d64c8902e64fa613

Observation 5ff2b34a-7ad9-403e-a716-6ce5d3514417 · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:24:56.608658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:2d4f1d57a5498514ac33b7e559d70c8020c122ca403037f4b966dfa55d5da40a

Observation f279528b-74e9-4044-a744-ad2286e9e365 · inbound

Measuring Progress Toward AGI: A Cognitive Framework cites this paper.

Measuring Progress Toward AGI: A Cognitive Framework Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:03:23.843852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-29T12:00:26.249340Z digest=sha256:5440f6844d9bd56227a734b83110e235c56aac0a0a4f5be459eaf38cda059574

Observation 13615bf3-77df-4b8d-b153-3737578c7886 · inbound

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination cites this paper.

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T06:02:27.886353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T06:02:27.886353Z digest=sha256:082d9a4481881d1fcd1069897d20e0cc18e334ec800bfe62e720a6b58913052e