Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2310.18018.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:51:15.453440Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
8
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation d38930be-bc42-42ce-8362-ea6b0e0573df · inbound
Unbiased Evaluation of Large Language Models from a Causal Perspective NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9171a74f-991e-4ff4-b18d-8ed9ca0db66c · inbound
RewardAnything: Generalizable Principle-Following Reward Models NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ff839f5-66d4-4b38-947c-22946783829d · inbound
Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67128382-2d9a-44be-83dd-fd934e1ae9bd · inbound
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0f330cad-9e76-467d-95ef-727506a63474 · inbound
On the Fitness Landscape in the $NK$ Model NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56c0d6cd-85ae-444a-96c4-7198018eabf1 · inbound
Artificial Phantasia: Emergent Mental Imagery in Large Language Models NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e5b6228e-fa83-4191-b596-db823287d718 · inbound
Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 293cbe29-27d6-4e41-91e6-87ae083e33c7 · inbound
DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5fe6ffde-e7d3-4b18-9162-da15dd263a19 · inbound
DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6048266-81af-4ed5-9c79-d42da57ac594 · inbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14e051e7-4ac8-4f34-8c88-9452415addd0 · inbound
Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c62ac00-20a6-47ee-9c7c-338a0ba8f1a1 · inbound
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 69286f6b-21dd-4371-9c4e-7f613ba6fab7 · inbound
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 66eb410d-85f7-4e23-b259-15edd22a1667 · inbound
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8bbe3091-2d39-4119-8607-e5f378ccfd67 · inbound
LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations? NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8ba959b7-1a87-4e52-8441-eaf435884bdf · inbound
The Case for Model Science: Verify, Explore, Steer, Refine NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 66ad5f14-c0ad-47f2-80ba-698b5e72f080 · inbound
Dissecting model behavior through agent trajectories NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 17835370-b4a7-4b61-bee8-ff4db2e32631 · inbound
Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4482a391-6e89-4f76-9714-9d72df27b199 · inbound
Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f07d090d-375a-4152-ad26-b306a029ab47 · inbound
Information Discernment in Large Language Models NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.