Pith. sign in

Paper Citation Record · LEDGER

Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2309.10677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.10677 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:38:42.292215Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c820bea0-77f0-4cc3-a4b9-a0b06817bc97 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.028002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:28dec3f26841e81b9aaec91d6db7d929bb011410193891ad1bed92c3073b42cb

Observation 4f9e806b-50c5-4845-a892-601a9d8ad91c · inbound

Are Large Language Models Memorizing Bug Benchmarks? cites this paper.

Are Large Language Models Memorizing Bug Benchmarks? Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.288484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.288484Z digest=sha256:870c5715aa9e936a3233fa081df93bf15c4a37c4f360afe41e7ccfc073e2805b

Observation 860f927a-9d1b-4c35-b4fc-9d5cea66c2cc · inbound

Are Large Language Models Memorizing Bug Benchmarks? cites this paper.

Are Large Language Models Memorizing Bug Benchmarks? Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.292215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.292215Z digest=sha256:0c939773b81f4775c630c76f864c9fe7d96d711530d8096e9af7753672cfa907

Observation 033cd38c-224c-43f2-9ddd-98674589c3f8 · inbound

How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence cites this paper.

How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-09T18:11:34.992365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:11:34.992365Z digest=sha256:17440757e159cb55a97f1b2a1f6670756169115d87fb5eb1bf23dcb5d4883b6b

Observation 3a55477a-f0cc-4719-8e26-4ed2979c2896 · inbound

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation cites this paper.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.977795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.977795Z digest=sha256:d07d5fcadbc093d5d2709c3bb6abc5e555c92261aa01062fed4f8154046a659e

Observation 86067b5c-a2f2-46c8-b25f-e75f0c369c44 · inbound

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis cites this paper.

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:26.295693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:26.295693Z digest=sha256:1c610ff0c5c080c831beb18d7fbb42fb82a231e70cbffc85037e2317a1b98d41

Observation 3a23f700-7fd3-4219-90c1-be9535d467ff · inbound

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity cites this paper.

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T22:04:18.052583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-08T10:23:02.697982Z digest=sha256:d55e15f900176e5293732740d28cc83eb8cd7516fcfbf55c79cab3e0c7a1891b

Observation 5834c8d5-4fd1-4e13-ad15-cccda5b6f6e0 · inbound

An Interpretable and Scalable Framework for Evaluating Large Language Models cites this paper.

An Interpretable and Scalable Framework for Evaluating Large Language Models Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:41:01.570612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:08:25.577363Z digest=sha256:f22b759045bed18905e774dfa7dd33ec0e727998ff824347d8b80694e18846b5

Observation c16934d6-2bc7-45dd-aa53-6b6d19c1c62e · inbound

Provable Joint Decontamination for Benchmarking Multiple Large Language Models cites this paper.

Provable Joint Decontamination for Benchmarking Multiple Large Language Models Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:44:29.427759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-22T00:40:54.038367Z digest=sha256:e70ef22bf2359978a8c99e7aa9f47fdcb09a7ccb99fa7e6b4bb5f717924d9b8e

Observation f2182f77-fe4c-47f6-b4cb-36b19fd18d4f · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:57.288015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:2038d1caa890b73806677467e2e202201194de2bf34a722864680197c0f014aa

Observation a9a3e6b1-6af2-40f7-8702-d757e1c18d07 · inbound

Predicting Task Difficulty Without Rollouts cites this paper.

Predicting Task Difficulty Without Rollouts Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.442787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.442787Z digest=sha256:d1d87c66aeaa9cde9094112dfd2c4dd44d77fa588385c922a7ca854fb2b31b03

Observation d04425ab-053e-4ec3-8b66-33e6ea411669 · inbound

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination cites this paper.

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T06:02:27.958702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T06:02:27.958702Z digest=sha256:201bb9274fb5cac9aec45641e5aaa268eafe2e668cb0df3e26a1a296c3f5d457