Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2311.01964.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:50.793745Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
18
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 3aaa96ce-fe7c-4b16-930d-13eb3909db90 · inbound
Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 09a21ebc-9711-4355-941f-f7a01f2d0c76 · inbound
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 223
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01ef215d-f693-4a35-a813-56451877a25c · inbound
Benchmark Data Contamination of Large Language Models: A Survey Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 187
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6e87ec1d-90ca-4af7-9894-a408d1a97434 · inbound
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 155
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 11e37873-e86e-4360-8c9d-92532dc29dd4 · inbound
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 197
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3983b9a3-1339-405a-8c93-7ae58305dce3 · inbound
A Conceptual Framework for AI Capability Evaluations Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c201a5ad-d818-43c7-b7a5-0128eba0b6ae · inbound
Can Vision Language Models Understand Mimed Actions? Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11652142-6a23-4b3c-b5d8-bee15e27fe51 · inbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a5eeb3d-ad4e-4eb4-807a-a8526dd0d0aa · inbound
Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ecc2e53-7f59-435d-afc4-b5bb1194e4e6 · inbound
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a680c7fb-277d-4b9c-a183-19a2afd2d214 · inbound
Deprecating Benchmarks: Criteria and Framework Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97aa3bc0-878f-47d6-8b62-1e86e5b135dd · inbound
ISACL: Internal State Analyzer for Copyrighted Training Data Leakage Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 814f58bc-4f4b-435d-92ee-04ac5bae3038 · inbound
InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca9ba4eb-c67c-4ba0-94b8-5b5319d0bcc5 · inbound
LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 62e80328-de9e-46ef-8346-c51a5dec4830 · inbound
Make a Video Call with LLM: A Measurement Campaign over Six Mainstream Apps Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13edee3e-fa11-45a1-a31e-00ca70fb6dd7 · inbound
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0be77a86-62d8-4578-a0e6-0e23617115be · inbound
Benchmark Leakage Trap: Can We Trust LLM-based Recommendation? Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb03d784-d5ec-47bb-b80b-073e1a2b49b2 · inbound
Agentic Business Process Management: A Research Manifesto Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 22ed82ce-384d-4802-8fd1-d4c1ae345d65 · inbound
LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a69ccd46-f3ac-4193-967f-d2b183c8a583 · inbound
Riemann-Bench: A Benchmark for Moonshot Mathematics Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 227c3cac-82a2-40f0-be47-0d92e3231ded · inbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5d80a5c7-a48b-4c8f-831f-22fbfb20d00a · inbound
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d1555bcf-9371-468c-b0c0-4efe0b8c70e2 · inbound
Training a General Purpose Automated Red Teaming Model Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dbdfad1f-74b3-4fa4-9dc4-5dd2f0084784 · inbound
STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2e23b564-b627-4049-8ba4-b2fd10035453 · inbound
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f8720662-6515-4a69-bcce-5408c93bda67 · inbound
Generating Leakage-Free Benchmarks for Robust RAG Evaluation Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d5bf708a-2ecb-4cc9-8afa-52f0f61be15f · inbound
Provable Joint Decontamination for Benchmarking Multiple Large Language Models Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 181
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6ac7c5d0-9059-4a44-9a04-cdcff63d2890 · inbound
How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7feb40be-8f95-4344-97f7-80e1ed66da49 · inbound
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3f65e975-0d9d-4cb3-a563-06e43cae517f · inbound
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 119018cb-5232-4f83-8943-cac3dbc2a450 · inbound
Flaws in the LLM Automation Narrative Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c4eb3bdd-b34a-46fb-aca3-ecf7dea39295 · inbound
Defeat Devices in AI Systems Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation be1c851b-6620-4735-9841-8d4ecfb0c51c · inbound
From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 208
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4bd18683-eceb-42fc-acc7-3c6ea3c4f451 · inbound
Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.