Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2503.05336.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:35.042357Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation f44943f3-cb76-4bfa-aeeb-bf86bf08a589 · inbound
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Toward an Evaluation Science for Generative AI Systems
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00bf41d9-f51f-42f5-b628-c437707f753c · inbound
Adultification Bias in LLMs and Text-to-Image Models Toward an Evaluation Science for Generative AI Systems
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67010805-7da1-477f-b792-654e356efed1 · inbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Toward an Evaluation Science for Generative AI Systems
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2f468dc-1180-48b3-b383-c073f710ffd0 · inbound
Correlated Errors in Large Language Models Toward an Evaluation Science for Generative AI Systems
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea0c3d7-609e-468a-b979-7bd66ba0adf6 · inbound
A Conceptual Framework for AI Capability Evaluations Toward an Evaluation Science for Generative AI Systems
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec3598ac-b114-4790-9e78-9e1066a36c88 · inbound
Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead Toward an Evaluation Science for Generative AI Systems
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 828d7cf6-478b-4f3c-8fac-08bb24bb253c · inbound
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Toward an Evaluation Science for Generative AI Systems
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30b72dfc-149c-4062-9953-cf756c0ac8fa · inbound
Participatory AI: A Scandinavian Approach to Human-Centered AI Toward an Evaluation Science for Generative AI Systems
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1c9d7d6-75c0-484d-90eb-1f5255204b8b · inbound
Participatory AI: A Scandinavian Approach to Human-Centered AI Toward an Evaluation Science for Generative AI Systems
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6a775fc-c6c1-4af5-afe9-ea84f218b4db · inbound
An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications Toward an Evaluation Science for Generative AI Systems
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40b6930b-71b5-46b5-a009-952795f8cf87 · inbound
The Persistence of Cultural Memory: Investigating Multimodal Iconicity in Diffusion Models Toward an Evaluation Science for Generative AI Systems
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eff0a74f-f146-4145-8d11-e8b26b292e0d · inbound
Making AI Evaluation Deployment Relevant Through Context Specification Toward an Evaluation Science for Generative AI Systems
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e3f3e52-cd17-4f64-82e5-d5df6f151ba4 · inbound
Towards an Evaluation Methodology for AI in Second Language Education: Lessons Learned from Developing L2-Bench Toward an Evaluation Science for Generative AI Systems
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec26a659-b467-44f0-963d-666ff420317c · inbound
Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research Toward an Evaluation Science for Generative AI Systems
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2da2f11-fed6-4085-aeee-ee23ff3e76bf · inbound
Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing Toward an Evaluation Science for Generative AI Systems
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5deee66-5baf-4f67-bdf2-f7cf1ddf2c66 · inbound
BenCSSmark: Making the Social Sciences Count in LLM Research Toward an Evaluation Science for Generative AI Systems
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63ac6204-afa9-429b-b0d0-b83fed439a25 · inbound
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Toward an Evaluation Science for Generative AI Systems
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22ee5cdf-d680-43df-a261-32de1b528991 · inbound
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Toward an Evaluation Science for Generative AI Systems
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d80f729-8b14-47c4-8441-e48c95297b47 · inbound
Unsteady Metrics and Benchmarking Cultures of AI Model Builders Toward an Evaluation Science for Generative AI Systems
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5fdd04f2-1eb7-4d1e-8b2a-1f4fae503f23 · inbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Toward an Evaluation Science for Generative AI Systems
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f756615-b9db-4248-ba42-eb1faca750f8 · inbound
HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Toward an Evaluation Science for Generative AI Systems
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af1ad875-3399-41f1-a4ac-2653c7c2e5cd · inbound
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Toward an Evaluation Science for Generative AI Systems
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d834dcae-40d0-4cbc-95ab-648cf68eb808 · inbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Toward an Evaluation Science for Generative AI Systems
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.