Pith. sign in

Paper Citation Record · LEDGER

NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2310.18018.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.18018 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:51:15.453440Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

8
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d38930be-bc42-42ce-8362-ea6b0e0573df · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.453440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.453440Z digest=sha256:ffc1d529fb3cf16313e315fdfb93c49ac787c6937c9602df1038fc78623dc60a

Observation 9171a74f-991e-4ff4-b18d-8ed9ca0db66c · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.003448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.003448Z digest=sha256:12026719930e786a50b63619732de9cc64ed8748e38e4f9ba0a30b1502c5cb50

Observation 6ff839f5-66d4-4b38-947c-22946783829d · inbound

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis cites this paper.

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:26.345390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:26.345390Z digest=sha256:ed7dfee2c33d4a81836ef268dbe02faa3dbc224c849cfd02f257b40a3f3d2faa

Observation 67128382-2d9a-44be-83dd-fd934e1ae9bd · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T03:22:01.356975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:7304c96cb194e0e3bc421cdd64c0a4b5de7fc5ba3f77364c75082d686ef80ba5

Observation 0f330cad-9e76-467d-95ef-727506a63474 · inbound

On the Fitness Landscape in the $NK$ Model cites this paper.

On the Fitness Landscape in the $NK$ Model NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T19:28:54.967333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:28:54.967333Z digest=sha256:b0434bf96687ad91fff35361bde6a5513d4a1a1150280626a30d682dc82b0c33

Observation 56c0d6cd-85ae-444a-96c4-7198018eabf1 · inbound

Artificial Phantasia: Emergent Mental Imagery in Large Language Models cites this paper.

Artificial Phantasia: Emergent Mental Imagery in Large Language Models NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:44:22.869054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T21:41:39.111769Z digest=sha256:2c75b1b06f06cfbea960321ef07b43859dd0fdb9bbd08304c18829f671e5753e

Observation e5b6228e-fa83-4191-b596-db823287d718 · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:49.706763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:49.706763Z digest=sha256:8a1dac7d18893d344997fb6ebb8d1ba560f8ae8e68da45cf2c2348ceed9bd9e9

Observation 293cbe29-27d6-4e41-91e6-87ae083e33c7 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:20:34.337660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T01:18:44.523602Z digest=sha256:0f53b4c120e23b9f7d6e2ba3fc745129ed0aafd99882174fd8bcfb4cec68fec9

Observation 5fe6ffde-e7d3-4b18-9162-da15dd263a19 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T00:11:51.863397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:11:51.863397Z digest=sha256:b0c75234ddb83475613c9365cc589c839f6585a7187bac752b7f902df27a01c2

Observation d6048266-81af-4ed5-9c79-d42da57ac594 · inbound

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications cites this paper.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.337035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.337035Z digest=sha256:bae1c5780c9c1cb340fb4a5705be25b1b5ee69569fff0a578af89167a54442aa

Observation 14e051e7-4ac8-4f34-8c88-9452415addd0 · inbound

Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest cites this paper.

Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:41:01.820540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T03:20:12.777878Z digest=sha256:e23a75862b683276ec6b3ac03b9523bcc2d7e74a4c1f9c939cf8d19e7bd788d3

Observation 0c62ac00-20a6-47ee-9c7c-338a0ba8f1a1 · inbound

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks cites this paper.

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:05.187194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:35:24.397273Z digest=sha256:915fe68fb5fc70d17a964224b25ca4bd6fae80ad5b9925507e30c3d261a35fe3

Observation 69286f6b-21dd-4371-9c4e-7f613ba6fab7 · inbound

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild cites this paper.

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.135213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T14:41:07.354007Z digest=sha256:bc2508e11791d4b4a65f05a174929d8e9cbf2ea35c165526573920bda7379e88

Observation 66eb410d-85f7-4e23-b259-15edd22a1667 · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:24:56.613744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:2c103a62de27afbe40d2c2acfad3691653494369f280fc5c08db4324b67c1447

Observation 8bbe3091-2d39-4119-8607-e5f378ccfd67 · inbound

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations? cites this paper.

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations? NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:33:45.019663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T17:33:03.397468Z digest=sha256:85f0fd261d96ac6493ec774ff397f10243e044b8a2e8b083e2e1410b8b4f670c

Observation 8ba959b7-1a87-4e52-8441-eaf435884bdf · inbound

The Case for Model Science: Verify, Explore, Steer, Refine cites this paper.

The Case for Model Science: Verify, Explore, Steer, Refine NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:16:13.563798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:24:32.311565Z digest=sha256:50b5dd6dda029334484a78338d975130a4d22ecf3794ca0481b71077767b1b33

Observation 66ad5f14-c0ad-47f2-80ba-698b5e72f080 · inbound

Dissecting model behavior through agent trajectories cites this paper.

Dissecting model behavior through agent trajectories NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.917233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T01:27:39.812496Z digest=sha256:379a4ec2c80ad0291f4592b16a7505d666bd4e7c37d845d079810b82c3b21576

Observation 17835370-b4a7-4b61-bee8-ff4db2e32631 · inbound

Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds cites this paper.

Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:27:18.679688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-02T19:18:43.558804Z digest=sha256:09336840740cf572eba8f017eb6b5521c597f17105411aac3ab10a4dbe4009ee

Observation 4482a391-6e89-4f76-9714-9d72df27b199 · inbound

Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge cites this paper.

Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:38:18.445574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T13:36:31.189451Z digest=sha256:80ef7136e5e34ff313ef426d211cf5b6f2598b226b329d058b6280e5a1f74923

Observation f07d090d-375a-4152-ad26-b306a029ab47 · inbound

Information Discernment in Large Language Models cites this paper.

Information Discernment in Large Language Models NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T13:26:26.237664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:26:26.237664Z digest=sha256:3e08ca9f04c987faea3bbf91b737d3373eea38966b42e276cad5b95462278eac