Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Benchmark Leakage in Large Language Models

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2404.18824.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.18824 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:57:15.748808Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 35845f36-317b-421a-94c2-f2807d182c79 · inbound

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models cites this paper.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Benchmarking Benchmark Leakage in Large Language Models

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:09:15.062242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:030805de1a38b0e710cc5d7eac563285980a7d383c8941055b4e90c923e529f9

Observation e6a2f3b7-9e67-434d-96b4-8ebd827c55a1 · inbound

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models cites this paper.

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models Benchmarking Benchmark Leakage in Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:09.658023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:07:09.658023Z digest=sha256:e84bcac51f983aa30f5f667eff3f3f2365d2928eb789461b97dcdc956a85dbb8

Observation 91903aa1-6de5-4571-be83-ce728ec2225b · inbound

Are Large Language Models Memorizing Bug Benchmarks? cites this paper.

Are Large Language Models Memorizing Bug Benchmarks? Benchmarking Benchmark Leakage in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:42.284289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:42.284289Z digest=sha256:b95e2b98083021c3e507c1c825fe57077f9c80b19cfb429a9a93da1e8076b66c

Observation de6c23c1-c11c-4110-9e8a-2f8e4e62b384 · inbound

CNNSum: Exploring Long-Context Summarization with Large Language Models in Chinese Novels cites this paper.

CNNSum: Exploring Long-Context Summarization with Large Language Models in Chinese Novels Benchmarking Benchmark Leakage in Large Language Models

Reference 5

Resolution
malformed identifier
no resolver link, observed 2026-08-11T23:09:32.524149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:09:32.524149Z digest=sha256:41912cdd9dc5b82bbcaca61f9d1629c2aea6ef4e15afb567247455e85b85d363

Observation f589d2df-966e-4b15-a73d-9f80d7ae1311 · inbound

QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs cites this paper.

QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs Benchmarking Benchmark Leakage in Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T14:40:29.599620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:40:29.599620Z digest=sha256:1595d2c605f06f2900bb2513dfdecf474f1afcaf7dc7753f0d3d441b575104ba

Observation b0987000-d762-46b5-9b84-d3d3b0efc085 · inbound

Large Language Models for Mathematical Analysis cites this paper.

Large Language Models for Mathematical Analysis Benchmarking Benchmark Leakage in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:27:39.903073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:27:39.903073Z digest=sha256:de1d8a0e4ff27eb592781d6f087ae3fefec8f3349b6a9a840d864e4b925827cc

Observation c062c923-6be8-43bf-878c-a578ffa12f7a · inbound

Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards cites this paper.

Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards Benchmarking Benchmark Leakage in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:44:33.033248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:44:33.033248Z digest=sha256:adef9c8a4045e826e2170528bc5b2fb5e907c92456cf73df355319c1250c5814

Observation 3730f6b9-d06d-41b0-b116-0aca4d3b5cc3 · inbound

JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models cites this paper.

JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models Benchmarking Benchmark Leakage in Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T15:04:55.640338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:04:55.640338Z digest=sha256:206aacecd5f0169919e424762d20ec466e20de9a2fe308659f323c8e1498d64d

Observation fc7cd146-896d-4efd-88a2-9f264f2ea236 · inbound

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models cites this paper.

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models Benchmarking Benchmark Leakage in Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T19:27:46.877517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:27:46.877517Z digest=sha256:be0009dfbce79469ed4bfaead71928805067071598fd5eb2029a31691021ab6b

Observation f36a13b3-4485-4c8e-9e26-7ec039e89033 · inbound

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation cites this paper.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Benchmarking Benchmark Leakage in Large Language Models

Reference 135

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.379586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.379586Z digest=sha256:6097b9bbbc1f0a5a2647150379fcbcedb2de435bc8e6fb6c431757de7965350c

Observation e677b2b1-af21-474d-a58f-b061c5d09c2f · inbound

Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion cites this paper.

Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion Benchmarking Benchmark Leakage in Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:57:15.748808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:57:15.748808Z digest=sha256:a87d2c20113898f20a1890de2533e3c5390b731bb5b86a454608e2d17dda2bb8

Observation 608bd7f2-ffaf-47f5-8cb9-d0a78d2b027b · inbound

Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs cites this paper.

Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs Benchmarking Benchmark Leakage in Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.597486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:24.597486Z digest=sha256:25ad4538aa64a51fb0efbbef2b01af0f305fb2bc6b4ba1945d09e4539656ebf0

Observation cbd67ca9-a106-43a1-a3b5-d6a351c50b4f · inbound

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation cites this paper.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Benchmarking Benchmark Leakage in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.982799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.982799Z digest=sha256:e4e97dd3e657821ecf3aaf7c3ad5d67e72e785183d6480968a13273ea131910b

Observation e5f296d6-5c29-4bf5-a846-f1b33b22266a · inbound

Evaluating the Sensitivity of LLMs to Prior Context cites this paper.

Evaluating the Sensitivity of LLMs to Prior Context Benchmarking Benchmark Leakage in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.260189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.260189Z digest=sha256:0d738ed5e20c13b29aeaaec7e4ee8a3bf4488fde460f8d16e8509c8cb765355b

Observation 61f9a667-6db2-4864-ae3c-66699921e2b4 · inbound

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis cites this paper.

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis Benchmarking Benchmark Leakage in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:26.386550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:26.386550Z digest=sha256:2e92697b48004753cf04412cb1f14c43901e918c97dab14b68aebdc1dc3ee136

Observation 79604d1c-895d-4767-aba9-8317fdb7eda2 · inbound

Chain of Methodologies: Scaling Test Time Computation without Training cites this paper.

Chain of Methodologies: Scaling Test Time Computation without Training Benchmarking Benchmark Leakage in Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:37.198751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:50:37.198751Z digest=sha256:695f741223d9442fd1c159f48d1d807a38cb51467eff0c23e1080e201d38ecfe

Observation 9a361e93-bc4a-4fe3-8ff8-29e3ba0a6029 · inbound

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics cites this paper.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Benchmarking Benchmark Leakage in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.477391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.477391Z digest=sha256:0785c6679617400fd00a9b9a57c5fbd981bb3b53cdca478607f1c74eee137233

Observation 5349fefa-a554-421c-82ef-0da9cf739b54 · inbound

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities cites this paper.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Benchmarking Benchmark Leakage in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.929888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.929888Z digest=sha256:92f110da25e80733f0cd42493067f0ee9e093fde707d00d9dca1edf1fbef6307

Observation d9b1cdac-411b-4f68-9f04-6e3dc58cf728 · inbound

Can Vision Language Models Understand Mimed Actions? cites this paper.

Can Vision Language Models Understand Mimed Actions? Benchmarking Benchmark Leakage in Large Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:50.139856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:23:50.139856Z digest=sha256:f863f2c978fb232055522d973bb70c4007d42c58b23a204b038c42eaec5eb930

Observation ad11074b-4823-46d4-91e6-e3be25d48af1 · inbound

Deprecating Benchmarks: Criteria and Framework cites this paper.

Deprecating Benchmarks: Criteria and Framework Benchmarking Benchmark Leakage in Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:07:40.526852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:07:40.526852Z digest=sha256:ccb2f3a0dfeb52069db574e68c1a68d6c87e6e2f8f1af310a34edcb3a10b9e2f

Observation 1481c7fa-2543-4237-a743-10a1e134b5b1 · inbound

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models cites this paper.

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models Benchmarking Benchmark Leakage in Large Language Models

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:15.997858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:16:15.997858Z digest=sha256:e4c388bdc0bcd78a5578cee0edc4b05db26eab78b9de34fabc2d0f29a4bced5a

Observation e5a2ae8d-e388-4217-b3eb-eedbb3fa2312 · inbound

GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines cites this paper.

GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines Benchmarking Benchmark Leakage in Large Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T00:41:56.045788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T00:38:43.897231Z digest=sha256:5f6ba7b679e04b228035bfea3fa6ff342748cbe1aeb4ca87fd2bef5a96085e30

Observation 4cdd65dc-7dfe-4e5b-8d6e-e89609178a30 · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering Benchmarking Benchmark Leakage in Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:50.832030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:50.832030Z digest=sha256:4fe4d4c621b0d1d6b4625a784f30b0af71be6348f90a119a8d790745a85ee115

Observation a2ede590-2b75-4e82-a160-b3e8b2954827 · inbound

MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science cites this paper.

MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science Benchmarking Benchmark Leakage in Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T10:00:42.785560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:00:42.785560Z digest=sha256:1a889f184f72872a8cba41fab49cced647984b0eadcb5871d2925f04e8955c2c

Observation 9d533bd6-19d6-4c41-9c84-3633a1f2fddf · inbound

LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation cites this paper.

LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation Benchmarking Benchmark Leakage in Large Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:20:59.061010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T06:16:01.603353Z digest=sha256:da66ef1354e363afa55320094b0b240e3f9f24379d34b9311c4a9d54809b2b62

Observation a32d5988-584d-4f29-9cb3-540f7386a92a · inbound

FiMMIA: scaling semantic perturbation-based membership inference across modalities cites this paper.

FiMMIA: scaling semantic perturbation-based membership inference across modalities Benchmarking Benchmark Leakage in Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T19:00:46.849166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:00:46.849166Z digest=sha256:78b6c42bb1b9fe98bcf148bae6b4eaab7ebac4491cf694a2538416e9f5f9df9e

Observation 9ae0c76b-85e0-4c01-be3a-10fb040c358b · inbound

Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners cites this paper.

Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners Benchmarking Benchmark Leakage in Large Language Models

Reference 10

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T17:28:10.516390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T17:24:05.849450Z digest=sha256:abec0efd20062f78e51567bbd65a601fc4cac6d674b2229996e012859b3be151

Observation 5c566624-e304-4be1-af2a-30111054f033 · inbound

SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora cites this paper.

SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora Benchmarking Benchmark Leakage in Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T01:02:46.482669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:02:46.482669Z digest=sha256:50936c549e796b62a93b95d6c13e272bcfd4424a3f93f9cd3c2afb323e3a0b45

Observation d9ccfd0b-1ad1-45a5-8e21-17c30ff436fd · inbound

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics cites this paper.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Benchmarking Benchmark Leakage in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:57.424441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:57.424441Z digest=sha256:e914e3efaade02a169d16ddd19e7e4b728825213df767df1c8cd5a89dd18dba7

Observation c35f61c7-01d6-473a-9633-244972149695 · inbound

MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events cites this paper.

MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events Benchmarking Benchmark Leakage in Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:10.039975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T11:19:15.920545Z digest=sha256:42c866f1cbc39857776130a4739ce57406940316264eb15988b27ed469b7ab0a

Observation bc4a7044-8616-4f93-b7da-af4e042ed8db · inbound

SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks cites this paper.

SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks Benchmarking Benchmark Leakage in Large Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:05:23.386113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T04:40:42.747298Z digest=sha256:04a0858f15ffa39e5fa9b23a3f9b8eaa6af7b3ccab6a873fcc042337d14c9bcd

Observation 2a72359c-2c5f-47d6-a70a-8e728236879b · inbound

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors cites this paper.

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors Benchmarking Benchmark Leakage in Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:18.734289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T22:07:58.614654Z digest=sha256:7bd730025fa59aae8b73cc3107864d9f215f4b97acecbddd6b67d951adb6d1d9

Observation b5e0477a-8a71-4b88-a2c8-c5c88ac8a16a · inbound

AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models cites this paper.

AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models Benchmarking Benchmark Leakage in Large Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:51:05.376351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T20:59:23.365434Z digest=sha256:13d0a4b2170f7fbb2aa90736f1b06f18eb920e84249de1ca487673bd839ca208

Observation 79bdc75f-a803-40bf-8fcb-ae2fd06e01c5 · inbound

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity cites this paper.

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Benchmarking Benchmark Leakage in Large Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T22:04:18.120684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T10:23:02.697982Z digest=sha256:d88dccc3d627a8452eeef5b7e9cb1a68b0900577ea20dc838f9659c78b9b9ab1

Observation 88f9965b-edb3-4822-a956-bfb4c39b339f · inbound

PITMuS: A Tool for Automated Bug Dataset Generation via Source-Level Mutant Reconstruction cites this paper.

PITMuS: A Tool for Automated Bug Dataset Generation via Source-Level Mutant Reconstruction Benchmarking Benchmark Leakage in Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:24:38.150937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T05:24:18.375671Z digest=sha256:dbeaf3fce61ef3d374b80e1711b743768e56e08d72133961f2daca66ac264f6e

Observation c7b760cf-d7b0-4618-b25d-e4cca102112c · inbound

PITMuS: A Tool for Automated Bug Dataset Generation via Source-Level Mutant Reconstruction cites this paper.

PITMuS: A Tool for Automated Bug Dataset Generation via Source-Level Mutant Reconstruction Benchmarking Benchmark Leakage in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T05:04:55.845629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:04:55.845629Z digest=sha256:3e5864eb3c24e18ef4003eb702b55cafc8bec76259226ba5fcd2b96d85dafa4a

Observation 110f064d-adfe-4ee6-af45-a46623d8e8eb · inbound

VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation cites this paper.

VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation Benchmarking Benchmark Leakage in Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:21:16.994334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T08:17:03.221396Z digest=sha256:2854c6af6ca33dd423f322d1751e2a439ee75ef512c73f94b48d857b5ad2b643

Observation 5e7df553-c4ce-4300-bb33-ea55a9c64edf · inbound

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs cites this paper.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Benchmarking Benchmark Leakage in Large Language Models

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T22:15:04.884435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T22:13:58.020838Z digest=sha256:f8d71aa51fa69c1e889fad9a7ff1257ec94095a68d883850356cf64d6d11b960

Observation 3732a0e4-37ef-43f7-948a-ee90bf729cdc · inbound

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs cites this paper.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Benchmarking Benchmark Leakage in Large Language Models

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:29:21.760196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:420923b5b8fed94090c38b7f945ece3520d73b032e6c9db3dbadc3f3b1a5dd0d

Observation 7725151a-1a91-4028-a8ae-bdefd39a0a07 · inbound

Can AI Agents Synthesize Scientific Conclusions? cites this paper.

Can AI Agents Synthesize Scientific Conclusions? Benchmarking Benchmark Leakage in Large Language Models

Reference 132

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:47:41.425924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T13:05:07.718882Z digest=sha256:3df5bf37ae10910892a3be4f082f9e656cffebc35b39b9d02e35623b0a73227d

Observation e5930678-901f-4aff-a538-6f425fcd1788 · inbound

Bridging Functional Correctness and Runtime Efficiency Gaps in LLM-Based Code Translation cites this paper.

Bridging Functional Correctness and Runtime Efficiency Gaps in LLM-Based Code Translation Benchmarking Benchmark Leakage in Large Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:08:57.953176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T00:56:23.711897Z digest=sha256:e94304addc461551c04d194f2aadec4d9794112d8204d1c2dcf4c660ea910804

Observation 8970d1b1-991c-44bf-aad8-faefee56bfd7 · inbound

Uncertainty-based Debiasing and Unlearning for Decontamination cites this paper.

Uncertainty-based Debiasing and Unlearning for Decontamination Benchmarking Benchmark Leakage in Large Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:49:52.411522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T05:57:44.576411Z digest=sha256:fc0aff94b4ec17660940ff04fb0c4879aafc10831571e5cc7a34c7b38cf906f5

Observation 027965fd-2235-4d53-a6cf-2e475773e4e1 · inbound

OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents cites this paper.

OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents Benchmarking Benchmark Leakage in Large Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:39:57.055022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T01:29:30.732471Z digest=sha256:9cdbdf3ecd5fec0a58e31a9aaf36d1f1505d35365e06b27bcf69b1b48695521c

Observation e127fd19-8449-43fe-89da-9f0f04b55cb4 · inbound

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models cites this paper.

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models Benchmarking Benchmark Leakage in Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:24:18.372738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T06:20:02.046020Z digest=sha256:ec6137d8b269734d666070a2b24d71e465c82b0fe83c0fcf41df99dfd6c21503

Observation c4746815-ff22-47ec-924e-2d076714c243 · inbound

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG cites this paper.

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG Benchmarking Benchmark Leakage in Large Language Models

Reference 97

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T01:15:13.184152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-01T01:13:13.329619Z digest=sha256:f43498d0fe9df520a27ed8f1c81ca9f9e2d009a9e86182a358e6dfd49614575e

Observation 71a2da4b-1397-486a-9975-79a067ba5313 · inbound

Meta-Benchmarks for Financial-Services LLM Evaluation cites this paper.

Meta-Benchmarks for Financial-Services LLM Evaluation Benchmarking Benchmark Leakage in Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.673308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-03T14:08:20.432931Z digest=sha256:f99a4e8e62d1dc2938606af695ba96d95322a2fa7da874c53729e3ceb69e180d

Observation f0c8bc16-9060-40f1-86cf-5ceed93a0191 · inbound

Navigating the Open-Source Model Ecosystem: An Empirical Study of Creator Practices in Artistic Image Generation cites this paper.

Navigating the Open-Source Model Ecosystem: An Empirical Study of Creator Practices in Artistic Image Generation Benchmarking Benchmark Leakage in Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T10:58:56.777252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:58:56.777252Z digest=sha256:aac6e65f295367962de7f8179285ba76000f2f8e8dfffbf916440fb224d6b013

Observation f42e648e-7f65-4ff9-8770-9e6a013df65d · inbound

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents cites this paper.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Benchmarking Benchmark Leakage in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:37.126253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:37.126253Z digest=sha256:02da510340ea38d2cb247339097f25a5742963a5cfa3b809624e4733f622f216

Observation 403211e6-ef5c-494e-a568-3b88a48bfefb · inbound

SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data cites this paper.

SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data Benchmarking Benchmark Leakage in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T11:16:11.626141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:16:11.626141Z digest=sha256:74d6d104fc59ca3408ee3b74d773934d0c376acf9294253fbd3ffb3764cb75aa

Observation 1e3a030c-b9bf-4faf-a812-f2a6152c4c70 · inbound

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers cites this paper.

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Benchmarking Benchmark Leakage in Large Language Models

Reference 236

Resolution
unresolved
no resolver link, observed 2026-08-15T14:34:14.741919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:34:14.741919Z digest=sha256:aeed46d556696a88bee77a680a7ad6d8453c292b6662956eb1ad316fbbd7d509