Pith. sign in

Paper Citation Record · LEDGER

Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2309.10677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.10677 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:25:20.442787Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c820bea0-77f0-4cc3-a4b9-a0b06817bc97 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.028002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:96766875e91d5f179d45b9ef6d2217d230d824a9b9828d2568db9800d1f06e38

Observation 3a55477a-f0cc-4719-8e26-4ed2979c2896 · inbound

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation cites this paper.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.977795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.977795Z digest=sha256:fd6c851d082d81675bbcfd22acfce475d3dcb1c7d991070f861ccc12ef995e28

Observation 86067b5c-a2f2-46c8-b25f-e75f0c369c44 · inbound

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis cites this paper.

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:26.295693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:26.295693Z digest=sha256:bf03535d8cb57db731333707aed2d0c904a711adbf3aab896e221aee2719220b

Observation 3a23f700-7fd3-4219-90c1-be9535d467ff · inbound

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity cites this paper.

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T22:04:18.052583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T10:23:02.697982Z digest=sha256:afb955c5e58c522295f6fa4c3113cad783e2f3c2bfb4153883c009cf7f99c02a

Observation 5834c8d5-4fd1-4e13-ad15-cccda5b6f6e0 · inbound

An Interpretable and Scalable Framework for Evaluating Large Language Models cites this paper.

An Interpretable and Scalable Framework for Evaluating Large Language Models Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:41:01.570612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:08:25.577363Z digest=sha256:eb456a1328ee0921d91f8106d866cd176f83a039f8787daef3e6a83a62fab946

Observation c16934d6-2bc7-45dd-aa53-6b6d19c1c62e · inbound

Provable Joint Decontamination for Benchmarking Multiple Large Language Models cites this paper.

Provable Joint Decontamination for Benchmarking Multiple Large Language Models Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:44:29.427759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T00:40:54.038367Z digest=sha256:37058f0298c222a4335adccf126017aaa126f8a5e7cc2479f7483f976bc398af

Observation f2182f77-fe4c-47f6-b4cb-36b19fd18d4f · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:57.288015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:38ffe0c8d65643de2711e71866c50704a288a101ac3a33c39f06b86d9fb51d18

Observation a9a3e6b1-6af2-40f7-8702-d757e1c18d07 · inbound

Predicting Task Difficulty Without Rollouts cites this paper.

Predicting Task Difficulty Without Rollouts Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.442787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.442787Z digest=sha256:93e1914972f736f5590bab7296aee50f8d7a1220eee38c019681d84ad3ba8b23