Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2104.14337.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:21:08.222059Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation b61e2e05-8d89-44cd-a85b-3293219884a3 · inbound
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games Dynabench: Rethinking Benchmarking in NLP
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f15163c3-9e9c-41e5-a705-858294c34b73 · inbound
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations Dynabench: Rethinking Benchmarking in NLP
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49ea3421-f8b6-4f38-b7d7-95ecb42e611c · inbound
CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++? Dynabench: Rethinking Benchmarking in NLP
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17846667-5c48-497e-9546-8845f1211c2f · inbound
AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge Dynabench: Rethinking Benchmarking in NLP
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10374146-0427-4a94-9227-3d428057b41f · inbound
What makes a good metric? Evaluating automatic metrics for text-to-image consistency Dynabench: Rethinking Benchmarking in NLP
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb36d922-8591-4171-8ea2-9d9ce961197d · inbound
WeAudit: Scaffolding User Auditors and AI Practitioners in Auditing Generative AI Dynabench: Rethinking Benchmarking in NLP
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 422a061a-ceeb-4687-91e2-7c693a4ba007 · inbound
AgoraSpeech: A multi-annotated comprehensive dataset of political discourse through the lens of humans and AI Dynabench: Rethinking Benchmarking in NLP
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bfebf8d-203d-47db-bea3-2d1bb1b32492 · inbound
Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities Dynabench: Rethinking Benchmarking in NLP
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39a34301-2c01-4c36-8d7d-52060d5d47d8 · inbound
Humanity's Last Exam Dynabench: Rethinking Benchmarking in NLP
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4b504f72-8c3c-41cc-9278-205d85bd7d02 · inbound
When Incentives Backfire, Data Stops Being Human Dynabench: Rethinking Benchmarking in NLP
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b929376-432b-4282-9211-7a2ad3be4b8f · inbound
Thinking beyond the anthropomorphic paradigm benefits LLM research Dynabench: Rethinking Benchmarking in NLP
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc4387f9-e218-4af0-a674-4f10df6c376e · inbound
LLM Performance for Code Generation on Noisy Tasks Dynabench: Rethinking Benchmarking in NLP
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c072be9-1d61-4f5e-97e5-5ac8b700aab8 · inbound
Potemkin Understanding in Large Language Models Dynabench: Rethinking Benchmarking in NLP
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 011ad500-ec58-4f4c-8ec8-45ea96369d38 · inbound
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models Dynabench: Rethinking Benchmarking in NLP
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2282e2fc-22ea-45e8-86ae-724b13ad9049 · inbound
Agentic Web: Weaving the Next Web with AI Agents Dynabench: Rethinking Benchmarking in NLP
Reference 115
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cf568cf-ac71-422f-8d4f-e25c6af4ebed · inbound
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Dynabench: Rethinking Benchmarking in NLP
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b6ad4775-1b20-42e1-9058-1c72ab173879 · inbound
Private, Verifiable, and Auditable AI Systems Dynabench: Rethinking Benchmarking in NLP
Reference 153
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91451675-098c-4626-8d54-ab661566828f · inbound
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M Dynabench: Rethinking Benchmarking in NLP
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c5adf30-311c-4f75-b382-796f0467f70f · inbound
Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation Dynabench: Rethinking Benchmarking in NLP
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 12d933b8-0880-4d12-86f6-c66322eed828 · inbound
EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning Dynabench: Rethinking Benchmarking in NLP
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fa4cbc4-7498-4013-9463-71b8fa7e0024 · inbound
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI Dynabench: Rethinking Benchmarking in NLP
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 142def31-9698-490a-a6c7-a17862a62603 · inbound
RoboPlayground: Democratizing Robotic Evaluation through Structured Physical Domains Dynabench: Rethinking Benchmarking in NLP
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 04341ae4-8797-4e6b-9c8b-15d01ebd9d6c · inbound
Too long; didn't solve Dynabench: Rethinking Benchmarking in NLP
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1382ce5a-86b0-4d76-9a13-c56ed3d8f611 · inbound
Too long; didn't solve Dynabench: Rethinking Benchmarking in NLP
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00e9179d-a40f-4902-a1e7-dfa5a8904107 · inbound
CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction Dynabench: Rethinking Benchmarking in NLP
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8cbb151b-15f8-4fd4-999b-69396e71f8ff · inbound
QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks Dynabench: Rethinking Benchmarking in NLP
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b8ad867c-3b0a-4c9f-b96a-8b3ae646d7ba · inbound
TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation Dynabench: Rethinking Benchmarking in NLP
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8c63ef9a-691e-4399-b0ed-37d88cdde605 · inbound
Analysis and Explainability of LLMs Via Evolutionary Methods Dynabench: Rethinking Benchmarking in NLP
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 10663faa-a7ed-4d7c-b0e8-5359b4e29507 · inbound
Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games Dynabench: Rethinking Benchmarking in NLP
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3f039523-3804-412f-a2bc-39b04018becd · inbound
Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks Dynabench: Rethinking Benchmarking in NLP
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f0e8d7ea-342f-4d28-bfcd-1ae36ec07607 · inbound
Interactive Evaluation Requires a Design Science Dynabench: Rethinking Benchmarking in NLP
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e6026c7b-c1ff-4fa7-b33c-fd7ec7ec5ce6 · inbound
Open-World Evaluations for Measuring Frontier AI Capabilities Dynabench: Rethinking Benchmarking in NLP
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9137d3af-0e49-425f-865c-49b6bf62c79c · inbound
CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Dynabench: Rethinking Benchmarking in NLP
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3610d9eb-e494-4f55-87ed-ae935d5d4b59 · inbound
Meta-Benchmarks for Financial-Services LLM Evaluation Dynabench: Rethinking Benchmarking in NLP
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d18b0e1a-4d83-4284-a968-2162598b044d · inbound
Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation Dynabench: Rethinking Benchmarking in NLP
Reference 306
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b805c310-4960-4f48-ab76-5432e72d93d6 · inbound
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks Dynabench: Rethinking Benchmarking in NLP
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.