Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2411.03923.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T04:31:04.401902Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T13:16:58.130004Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation bb55b6ed-45c0-445b-a95d-ebfd06e968c7 · inbound
LiveBench: A Challenging, Contamination-Limited LLM Benchmark Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d1eb138-e830-4a39-bfb2-be530c42a6d8 · inbound
StockSim: A Dual-Mode Order-Level Simulator for Evaluating Multi-Agent LLMs in Financial Markets Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2a72f36-9d7c-4e11-a269-6549549feb13 · inbound
Position: AI Evaluations Should be Grounded on a Theory of Capability Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ca0e5472-5a92-4f1e-80c4-6a08e2760cb1 · inbound
QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 60f42c4b-1ba6-476c-a6b6-5419dd0d799f · inbound
Measuring AI Reasoning: A Guide for Researchers Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 09cbb8ab-6e55-4a22-91c5-76b151197650 · inbound
Dataset Watermarking for Closed LLMs with Provable Detection Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c3d3fcce-b28a-408a-bbfd-68acd20b10e8 · inbound
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a8712fb1-cb99-4579-95f3-03c93b80a763 · inbound
Amplifying Membership Signal Through Chained Regeneration Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fd18870b-ebe7-4d86-91f6-677dc74517e0 · inbound
MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e9a46123-ac97-4ffe-a532-9b4ac9c3f628 · inbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0328c06-8a03-4415-88fd-e754b03f040c · inbound
Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.