Pith. sign in

Paper Citation Record · LEDGER

Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2311.04850.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.04850 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:19:51.012398Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 26dde097-e60d-47b9-ba7b-9ca5f4315064 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.134008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:dee4b4bb301e5b746ee9222eb87ac171a2060cfeea8af19b8d7f530bed43bce6

Observation 79ffe552-61e5-4673-bf9b-bc4fde5114c9 · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 208

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:58:17.334977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:a758c589fafcd06998f61b463b2ade72b8005d75afb47333462c9429b0409e5e

Observation d04fd98e-06bd-41e4-af68-e6b519a0ec47 · inbound

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools cites this paper.

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:08:09.653347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T08:08:09.444352Z digest=sha256:510778642ccb379e633a0870e2262bdcd44ab0968f57dc4adf5b48a7a9585513

Observation 1bb22ce2-c68b-422d-948d-2f1f77c32e39 · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.341453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:f208a005a3d4048dd717c0ec2853cc63d20ccedf028c35ff28761508d4f0e93e

Observation 54c99d7c-0932-4665-88df-4cafaffd7561 · inbound

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories cites this paper.

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:36:28.943124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T14:32:57.426996Z digest=sha256:5eaaf2458030f82207f5f6be2ae76d98e972040378d123b45706d4062aecbc20

Observation 1cf0bda5-0b6e-4ddb-b0ff-0793c0142615 · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:51.012398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:51.012398Z digest=sha256:f1081328abef3a2401976b44cd475de8ccc5c26d56301aee500d08b4f1b9932f

Observation d25e5c67-f333-4fda-b471-faf6e2b0bdff · inbound

Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation cites this paper.

Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:40:51.520975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:57:34.386632Z digest=sha256:18f68c5cfba9981a1898072c308578bbf1d94ba33c9597c3f5a6a0d521e8451d

Observation 548a2693-0ac5-442b-aac0-89d3bec954ea · inbound

DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math? cites this paper.

DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math? Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:04.832167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:46:24.780270Z digest=sha256:805b913da6e8ff10700cacb8af2a914e675b35b716a809cddf20cf1328d731d0

Observation 47e2479c-fb1f-496d-827e-eaaf883b2921 · inbound

TPS-CalcBench: A Benchmark and Diagnostic Evaluation Framework for LLM Analytical Calculation Competence in Hypersonic Thermal Protection System Engineering cites this paper.

TPS-CalcBench: A Benchmark and Diagnostic Evaluation Framework for LLM Analytical Calculation Competence in Hypersonic Thermal Protection System Engineering Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.514127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T05:25:59.343884Z digest=sha256:dabfc65c14501244203a6ee18bb82620174b413bb3fd6e89cde375d72d34f9ac

Observation 22b7cd4d-8090-460d-9758-5f01a02de39b · inbound

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation cites this paper.

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:17:06.819446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T03:00:26.137401Z digest=sha256:18cd0d1ea11c00d2e52b99d3e4d8d65cc5f500788a892e22bec016704403755e

Observation 6fd5b37e-9661-4f13-97f0-0af6e7d0b045 · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:36.959600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:86d9391813e4bf8c727ca3b9b2e9b75f7e8508735e10d9bc2ca95e56d60ce3a4

Observation 1b15f2a6-e35a-4bde-a1ba-9bf71adfead8 · inbound

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games cites this paper.

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:17.276350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T17:06:32.814188Z digest=sha256:a3f8b1ad974c067538b946eb582ed26b4a3a89489a27d78e253fc682f8821775

Observation 9d53cd78-ddd4-4c96-9571-3b76c9fa1010 · inbound

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity cites this paper.

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T22:04:18.090953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T10:23:02.697982Z digest=sha256:a8c4a46358384b5810065ed489f3db407f0f357e1ce32f69fdafe627c3969865

Observation 1c5fe167-2a6f-4ffc-b264-cb3a266a33fe · inbound

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations cites this paper.

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:59.899239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T00:57:55.616036Z digest=sha256:622c1952ccefb179e8f94aba1ce3eae7ca3d527a06c1dee980800e0225a68d23

Observation bcaeda91-673a-474a-8963-2564111db7eb · inbound

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations cites this paper.

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:07.976186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:42:00.965554Z digest=sha256:5eb425f2ed52eb27a53e5ae8e8ee82fefc33cc0bc245c5e14580e3907f507e43

Observation ec4241e4-e7c0-4b4c-8c06-4660896cd370 · inbound

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation cites this paper.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:23.900869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:a3cf3befcc0088bdb38a72360a0837c3ae31e51a88c0ff0bb593808b0d140a11

Observation abfb013e-e714-4037-85ac-acc881089016 · inbound

Decaf: Improving Neural Decompilation with Automatic Feedback and Search cites this paper.

Decaf: Improving Neural Decompilation with Automatic Feedback and Search Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:07:08.093098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:03:19.836506Z digest=sha256:1d2b494e2784717579bacf667fce3f3c4bcea9bc94b4f99970fe89974a0b309f

Observation 481aa828-3a0e-498f-9f0b-f56df90e7622 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.862291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:b9b840409efd2f0fd5497b2e42af7133d83748f60077e910849851e331f8b9d4

Observation a0f175ac-6597-48d7-9bf4-51fc3977f5ec · inbound

LLM Benchmark Datasets Should Be Contamination-Resistant cites this paper.

LLM Benchmark Datasets Should Be Contamination-Resistant Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:23:07.058655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T07:19:50.354875Z digest=sha256:f4f4bb4d2e01333d425430b2674677e99c507757809e9518c5dd71025f9ed5e7

Observation ae61037a-4a0c-4149-9764-5047fc27a22e · inbound

Provable Joint Decontamination for Benchmarking Multiple Large Language Models cites this paper.

Provable Joint Decontamination for Benchmarking Multiple Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 172

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T00:44:29.489044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-22T00:40:54.038367Z digest=sha256:d852e9faa603a14cc9371a4dfadd5fcfd29bc85e9c1af45b97a6ef4e3f3401e0

Observation e50e38d5-a330-4082-9fc1-b47897ecc449 · inbound

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation cites this paper.

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:06:15.310194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-22T08:05:42.212459Z digest=sha256:0135a8a57b4128601df85383b01354d645f30adfa72a6f9f64e00f926a9a5bdd

Observation 81a8cf3c-06b8-49b8-bbd7-3da728bad548 · inbound

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness cites this paper.

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:24.444660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T04:37:01.537128Z digest=sha256:2fa63aa7cf608d41506385f74051b96e8d069bd105dfa551dc11baefb99e3d0f

Observation e11c3d40-017b-429b-8690-ff2370e45d98 · inbound

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs cites this paper.

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:47.895008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T15:32:05.976335Z digest=sha256:946baf5798622ed586b82fec5a9504b5109f1cb39409645123968229904339f0

Observation 04ea9e21-1c59-4575-a510-f812fd2b7aeb · inbound

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild cites this paper.

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.155954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T14:41:07.354007Z digest=sha256:fcfa2db970d28f8edc6e31c51141c5e100e94dce3307a755eb8fcb7d410dbb4c

Observation 6c2dc137-5f7d-4787-b49a-71cbffd49ef8 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:34:40.588543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T13:27:50.367497Z digest=sha256:8b6961d3ace60ac60dfa382d16eafec3bc26ec6600ec4e3c1bba9f58e00a6e3b

Observation b5a659b1-05df-4a48-8f38-913a67b3cfe8 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:05:31.914578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T07:35:51.017797Z digest=sha256:272068cc0e5e37eca159bea03783d120849679f6ed50d5873e550b0889f787fe

Observation 73c90e48-132c-4c28-9ee1-a4949e9206f0 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:27:26.731890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T23:22:08.567841Z digest=sha256:2488b67e7a6870261d3bfdeb5fc60a79b97300d30160ca601db7e27d8a752054

Observation dab2454a-666d-4842-b07a-aed17edf3fd0 · inbound

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models cites this paper.

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:04:38.780577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T12:00:05.807004Z digest=sha256:146d21bdc274c1122e6bb9b7d56c59d355edee89855b1259a32f11f5175221a7

Observation 429ec9d5-66e8-4def-a403-8144313ea449 · inbound

Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction cites this paper.

Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.880656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T09:48:49.786021Z digest=sha256:39b46454424bee37eaa612f5f550579e9ef98118bb84d4318c503a27e854e0d7

Observation dcbed4ea-62a8-48a1-b727-be0d3cb6004e · inbound

Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs cites this paper.

Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:37:56.593959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T09:57:14.328157Z digest=sha256:40079312286dcf4667dbb10c78818f84aafd8f97072393d36bf3b5d2c03c1df0

Observation df54faea-0bd2-4570-a54c-c347743948f7 · inbound

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models cites this paper.

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:24:18.367339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T06:20:02.046020Z digest=sha256:134d6672d3448c9f48b6d7c406b52a76f690d17310a2302046bf14400dff0006

Observation e34d5daa-ff82-4dcf-8c9b-19274d08c5d6 · inbound

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI cites this paper.

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T05:03:42.877641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T05:03:42.877641Z digest=sha256:e07df4983b2f28677f5dea2ce6d46102304e02f4ffda012a041169ba59159105