Pith. sign in

Paper Citation Record · LEDGER

Don't Make Your LLM an Evaluation Benchmark Cheater

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2311.01964.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.01964 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:50.793745Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

18
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3aaa96ce-fe7c-4b16-930d-13eb3909db90 · inbound

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders cites this paper.

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:05:56.990257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-24T03:03:51.053556Z digest=sha256:b60b7ac162b84a826875a2978db1f9c8cbe43339e5aa71e6e70e5f99eecbb00d

Observation 09a21ebc-9711-4355-941f-f7a01f2d0c76 · inbound

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code cites this paper.

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 223

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:34:42.839749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T17:34:42.565806Z digest=sha256:85cb52efb6acb947be4aaf6781aa5e93597842fcb02ed80fbfb8cca8739ac2bd

Observation 01ef215d-f693-4a35-a813-56451877a25c · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 187

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.160577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:5c260b3b5744a0ebc4084236f3771f90cc86cca8302e2355904a4517e2658da7

Observation 6e87ec1d-90ca-4af7-9894-a408d1a97434 · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:05:03.901359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:d9a42197c6ddb4fe8df76f71a93db927f2d47ef52ee9b72380573e7cbbac7895

Observation 11e37873-e86e-4360-8c9d-92532dc29dd4 · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 197

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:51:13.131940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:d7a7ca4362177d17214ddc771d985a7bf42c6016e1100f467beb8c8dc88dab75

Observation 3983b9a3-1339-405a-8c93-7ae58305dce3 · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:53.137574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:53.137574Z digest=sha256:3a608f2f56d493d0b937e4e4ff007f3b8abb4e3134773a034d7b7d3701278397

Observation c201a5ad-d818-43c7-b7a5-0128eba0b6ae · inbound

Can Vision Language Models Understand Mimed Actions? cites this paper.

Can Vision Language Models Understand Mimed Actions? Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:50.793745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:23:50.793745Z digest=sha256:668ee67d0c5e45a516376ba042bc0de3a386178562bd2d10ca91bc4ad22052b6

Observation 11652142-6a23-4b3c-b5d8-bee15e27fe51 · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.438581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.438581Z digest=sha256:149a72abecedd5cc9f3378540738f1c64d4a7ac088b476aeccb3056504ab95a7

Observation 5a5eeb3d-ad4e-4eb4-807a-a8526dd0d0aa · inbound

Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances cites this paper.

Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:49.856525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:47:49.856525Z digest=sha256:fcdb3f63a4103fe67f98f54a3f35453506f8dcbd231735df5a46a1a7e7e44c3b

Observation 5ecc2e53-7f59-435d-afc4-b5bb1194e4e6 · inbound

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs cites this paper.

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:43:41.764235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:43:41.764235Z digest=sha256:c98705445f85daa84c0922feb23393325cf87abc4e395c7525d6d04ce801651c

Observation a680c7fb-277d-4b9c-a183-19a2afd2d214 · inbound

Deprecating Benchmarks: Criteria and Framework cites this paper.

Deprecating Benchmarks: Criteria and Framework Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:07:40.530184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:07:40.530184Z digest=sha256:b1d7d9c708e0f5336382cf0d19e443679d3c8401130d44a227b609088afeb467

Observation 97aa3bc0-878f-47d6-8b62-1e86e5b135dd · inbound

ISACL: Internal State Analyzer for Copyrighted Training Data Leakage cites this paper.

ISACL: Internal State Analyzer for Copyrighted Training Data Leakage Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T16:50:28.964070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:50:28.964070Z digest=sha256:6421ffe8344a0d28ae0e9ca460b86d71d577ed9c61308470e07b113e43d42ca5

Observation 814f58bc-4f4b-435d-92ee-04ac5bae3038 · inbound

InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity cites this paper.

InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:44:28.055272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:44:28.055272Z digest=sha256:d31036ab684cbe7427df71519e32a954f57413b4f65712e8bdd3f481d9a8b1e2

Observation ca9ba4eb-c67c-4ba0-94b8-5b5319d0bcc5 · inbound

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories cites this paper.

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.964642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T14:32:57.426996Z digest=sha256:6f0ebf6cfac35e54d5390a8fcc72a61134039cc25bc66abd31bd48bd03726654

Observation 62e80328-de9e-46ef-8346-c51a5dec4830 · inbound

Make a Video Call with LLM: A Measurement Campaign over Six Mainstream Apps cites this paper.

Make a Video Call with LLM: A Measurement Campaign over Six Mainstream Apps Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T13:29:04.649764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:29:04.649764Z digest=sha256:79f3597bc25005c7bc3aa90094dbda130a439a014d11f66916dda6f33cb52649

Observation 13edee3e-fa11-45a1-a31e-00ca70fb6dd7 · inbound

SciCoQA: Quality Assurance for Scientific Paper--Code Alignment cites this paper.

SciCoQA: Quality Assurance for Scientific Paper--Code Alignment Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 11

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T13:20:57.867847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T13:20:07.741347Z digest=sha256:3be36b6c4919a86172a820e092d91849d455e7c17f0f3b809f899c29fccab4f4

Observation 0be77a86-62d8-4578-a0e6-0e23617115be · inbound

Benchmark Leakage Trap: Can We Trust LLM-based Recommendation? cites this paper.

Benchmark Leakage Trap: Can We Trust LLM-based Recommendation? Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T23:32:09.809875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:32:09.809875Z digest=sha256:0922145c8d09f7580914e6a276e04c107f43f752da1a5f553998aef93e94dd1e

Observation fb03d784-d5ec-47bb-b80b-073e1a2b49b2 · inbound

Agentic Business Process Management: A Research Manifesto cites this paper.

Agentic Business Process Management: A Research Manifesto Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:25:18.811678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T08:24:56.763438Z digest=sha256:54a33e96921de1e5f88a90c8894290ec62c7f551e7ae2ce9df46a26110ae6e0d

Observation 22ed82ce-384d-4802-8fd1-d4c1ae345d65 · inbound

LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection cites this paper.

LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:55:49.393337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T20:30:53.094207Z digest=sha256:322cc04de7e04c33a47f086712009b599c49c3f0e79064c0921a16991bec44e3

Observation a69ccd46-f3ac-4193-967f-d2b183c8a583 · inbound

Riemann-Bench: A Benchmark for Moonshot Mathematics cites this paper.

Riemann-Bench: A Benchmark for Moonshot Mathematics Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:00.811359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:02:42.607682Z digest=sha256:3119b37fe16d22d17e6093cdaa0d15aa7733a23d330133038b6a00ad42465e85

Observation 227c3cac-82a2-40f0-be47-0d92e3231ded · inbound

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles cites this paper.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:20:58.932401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:9983adbc2928a0ec7844d177c8e6a10df99533fcf87f291270b9e82b7cb940d8

Observation 5d80a5c7-a48b-4c8f-831f-22fbfb20d00a · inbound

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks cites this paper.

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:05.207324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T00:35:24.397273Z digest=sha256:8377efefecd5b70a3ee6cc77a820e9c2f52fde2590862c70ba638618dbe42544

Observation d1555bcf-9371-468c-b0c0-4efe0b8c70e2 · inbound

Training a General Purpose Automated Red Teaming Model cites this paper.

Training a General Purpose Automated Red Teaming Model Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 5

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T19:41:08.319868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T11:20:09.224745Z digest=sha256:8234ddbcd9f200f7a4b2208845b07f96ab4f013358ff9ad19ef58634cf58b373

Observation dbdfad1f-74b3-4fa4-9dc4-5dd2f0084784 · inbound

STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator cites this paper.

STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:01:10.945701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:39:30.528601Z digest=sha256:2918e7d041c91e50c7b4849302688d3a6fd2c940b8406bad801433f3c3c84167

Observation 2e23b564-b627-4049-8ba4-b2fd10035453 · inbound

Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora cites this paper.

Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:16.280192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:13:30.648190Z digest=sha256:5d89c05ab0cf6e06a4c989a34415d9d52e029ed7eac437a3a9b40c7d8a6bf500

Observation f8720662-6515-4a69-bcce-5408c93bda67 · inbound

Generating Leakage-Free Benchmarks for Robust RAG Evaluation cites this paper.

Generating Leakage-Free Benchmarks for Robust RAG Evaluation Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:06:18.948679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:04:39.253429Z digest=sha256:3f9c898cca6b4935dde3108477689c8580bdfc6eee2c096757ba3ae390e54705

Observation d5bf708a-2ecb-4cc9-8afa-52f0f61be15f · inbound

Provable Joint Decontamination for Benchmarking Multiple Large Language Models cites this paper.

Provable Joint Decontamination for Benchmarking Multiple Large Language Models Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 181

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:44:29.503089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-22T00:40:54.038367Z digest=sha256:775456c5277a512d605584e13aa357f629d8f923725db602388510544e3552dc

Observation 6ac7c5d0-9059-4a44-9a04-cdcff63d2890 · inbound

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness cites this paper.

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:24.435108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T04:37:01.537128Z digest=sha256:e19e385d949fd663d4009d02b3174aac639addc180593531066bf2d6a2134383

Observation 7feb40be-8f95-4344-97f7-80e1ed66da49 · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:24:56.540890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:8fd2d0aeeb6bfca546738871e704f71d0c0e2a76efc25354ef01dcd3640e07e9

Observation 3f65e975-0d9d-4cb3-a563-06e43cae517f · inbound

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation cites this paper.

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:52:49.465940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T23:49:57.580051Z digest=sha256:f91cb62a51f68915c98734f8168a7c1c2db0fa8216d20408641697acaf4860d2

Observation 119018cb-5232-4f83-8943-cac3dbc2a450 · inbound

Flaws in the LLM Automation Narrative cites this paper.

Flaws in the LLM Automation Narrative Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.290264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T10:52:36.252919Z digest=sha256:5cf03c8e0103f52bd45898784a621061b1ababd6cab45bb10f957addf95646f8

Observation c4eb3bdd-b34a-46fb-aca3-ecf7dea39295 · inbound

Defeat Devices in AI Systems cites this paper.

Defeat Devices in AI Systems Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:44:27.875364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-30T08:34:58.879346Z digest=sha256:1b9720cbbf0b5c79585ab464d306bc3e077ef6564522111bd8efe5ec8b75ecf3

Observation be1c851b-6620-4735-9841-8d4ecfb0c51c · inbound

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs cites this paper.

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 208

Resolution
verified exact
local_arxiv, observed 2026-07-10T13:57:06.719623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-10T13:49:17.343893Z digest=sha256:520a2a6e9b52e380f730a4c43825904ba0a4cb9b4ab68647aa48f69970b58e28

Observation 4bd18683-eceb-42fc-acc7-3c6ea3c4f451 · inbound

Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent cites this paper.

Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:02.934623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:07:02.934623Z digest=sha256:efd037b6dcf96e02a3d26b6dc92f0f559c7274a1e81d6835f795c1dcdfcfc038