Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T16:11:36.483820Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 100 of 155 outbound references and 0 inbound Pith citation observations for arXiv:2606.09809.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T16:11:36.483820Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 155 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7b245c6e-5c7c-42ae-8c18-1581876bfa32 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Developing and maintaining an open- source repository of AI evaluations: Challenges and insights
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 768ab1bc-4dd9-4f56-856c-e3874910ec0a · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 09a70119-7f73-4524-a805-c41a03f73f33 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Audit and Assurance of AI Algorithms: A framework to ensure ethical algorithmic practices in Artificial Intelligence
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 300eba02-62bf-4e26-a423-27cb2c662207 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Lessons from the trenches on evaluating machine-learning systems in materials science.Computational Materi- als Science, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07de6629-35e0-4160-aa89-04d53f06453f · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting AllMetrics: A Unified Python Library for Standardized Metric Evaluation and Robust Data Validation in Machine Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fddfe3a7-b55c-4ae0-9750-b4d84e401043 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Arnstein
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d739621-4023-464d-a898-50de56e202e6 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Comparison of AI models across intelligence, performance, and price,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5ca1dac-2be0-4b75-85b2-5cd6b57304d3 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17a62291-90e3-4113-b81b-97ecc6dfe3ba · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting AI Risk Atlas: Taxonomy and Tooling for Navigating AI Risks and Resources
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ea6ae664-fb70-4f10-85d7-0fffb38b3ebc · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e6c38c0-6447-45b6-9866-3d149e859c37 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 220186bc-b6cf-4092-87cc-f29cb8de6b89 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Every eval ever: Toward a common language for AI eval reporting
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c79f37c2-f434-48e5-8986-bb51d6b42e8d · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting arXiv preprint arXiv:2511.04703; presented at Neurips 2025 , year=
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 76c7d314-f787-4bf5-903c-8c0438e44c77 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Absolute Evaluation Measures for Machine Learning: A Survey
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f60233f5-bd3c-45ba-bf68-fe3356b03fd0 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Open llm leaderboard (2023- 2024)
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8159c697-2f6f-473f-a423-8e39246aa914 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Lessons from the Trenches on Reproducible Evaluation of Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation facbbe95-f13d-4bde-b884-df235d47cfd0 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting A metrological framework for uncertainty evaluation in machine learning classification models.Metrologia, 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52ce2210-9dc6-4fbe-af4d-2125cc502f24 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting The impact of standardisation and standards on innovation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41d0cab6-d5b4-422c-936d-348cc8619a9e · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Assessing ai: Surveying the spectrum of approaches to understanding and auditing ai systems, 2025
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b8a0ff5-d55f-4092-b74e-5c7c4aa83d55 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Evaluation for change
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b784dba0-3e0f-4119-9cf7-790e0f2aeffa · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Bordes, C
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3756fd09-6a51-404c-a723-06998069018d · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 809552e1-9e2c-49d6-a9ce-34459b9f84be · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Cohn, and Jose Hernandez- Orallo
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9562ffa5-7302-4efd-937f-cd72e94c3bb2 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting best fit
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8fcc0fc-ff7a-4583-9ce0-3a662040d04d · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Black-box access is insufficient for rigorous AI audits, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fea5f83f-3ffc-4564-a951-9d7608954e35 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting ISBN 978-1-4503-7110-0
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d5236e58-a0c7-4cc9-a54f-dd3d2d21c171 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Managing misuse risk for dual-use foundation models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7ff8246-6aca-4a75-a1ad-594b2a8564d1 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Test & Evaluation Best Practices for Machine Learning-Enabled Systems
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 738ae33f-8811-4c89-8585-260ba7a98ee0 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Evaluating machine expertise: How graduate students develop frameworks for assessing GenAI content, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 613320f0-72c3-489b-ae48-3616f67f4de5 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Gonzalez, and Ion Stoica
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7dd991b-4739-4eaf-a6d5-4b1e569a3d21 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Collins, Karel G
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4728629b-1ad2-4939-99cc-83d70a6f65f9 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Who Audits the Auditors? Recommendations from a field scan of the algorithmic auditing ecosystem
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 52dda918-d769-4483-b073-651ec09383ea · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Evalcards: A framework for standardized evaluation reporting, 2025
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5287ce39-eb69-4e26-8f6b-e5a182d71ff0 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Nahab, and Xiao Hu
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a57e7f02-52f4-4b32-bcc2-485640696c9e · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Reconsideration on evaluation of machine learning models in continuous monitoring using wearables
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation de13c229-28b5-460e-8bc0-67b9e5c9cf20 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Introducing Epoch AI’s AI benchmarking hub, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b4b035c-af93-4529-9ec6-f64841e35e2b · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Can we trust AI benchmarks? An interdisci- plinary review of current issues in AI evaluation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53ada467-ae6f-44a5-a908-1db68dcba33c · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting The general-purpose AI code of practice: Safety & security chapter, July 2025
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f400d197-4754-4cab-bf45-ba0286b68d98 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting The general-purpose AI code of practice: Trans- parency chapter, July 2025
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae0b8826-358a-44c9-ba96-fc0b3c821c9c · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting EvalEval: Every eval ever shared task, 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b10e715-edaf-4f83-8432-d3c49d9e84c8 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Good practices for evaluation of machine learning systems
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 16a21a07-681d-48d3-8186-bdb157931f7d · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Frontier capability assessment
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e977c2d2-ef7f-4c93-8630-5dc0dd0dd403 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Datasheets for datasets.Commun
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69090bd6-e9ac-4361-bbef-53f3c68e9e69 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text.Journal of Artificial Intelligence Research, 2023
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187eaf4c-0767-4ac9-9c9c-5e6dd4e50397 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5247466e-5473-45ad-8301-eddc3cce436b · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Stress-Testing Capability Elicitation With Password-Locked Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6d2dc62f-528e-4866-81d7-d551b8797e4c · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Olmes: A standard for language model evaluations
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0742dfa1-161e-4604-b4f6-09fc8b534b28 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Gupta, Jessica Hullman, and Hari Subramonyam
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4dba60e-6eb6-47b0-b0a5-a631bcea00bc · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting System Cards for AI-Based Decision-Making for Public Policy
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6cabc94-c87a-4c08-90a0-e8c0dc6df617 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Empirical Privacy Evaluations of Generative and Predictive Machine Learning Models -- A review and challenges for practice
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 88aeeb68-936f-43be-a5db-b62d9ca0884f · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Bernstein, and Mykel John Kochenderfer
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2d40dca-467b-4097-86e9-7d10adf88655 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 846636de-5d90-4d31-a485-24fe2a3e4881 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Auto-benchmarkcard: Automated synthesis of benchmark documentation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7d06c8c-206c-475b-a101-3bc49ba1521a · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c838f560-df5b-48bf-b216-70c9b87f9898 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Evaluation gaps in machine learning practice
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78aaf0f0-e42a-45c9-a479-6b23b4a90f84 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Rethinking Machine Learning Model Evaluation in Pathology
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 10fa77d5-1123-4b07-b294-40ba69195ff5 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Deprecating benchmarks: Criteria and framework
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ec1306f-fbf5-46f7-8e94-18bfe0825f5d · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Cantrell, Keiran Peng, Thanh Huy Pham, Christopher A
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c8c2b13-2bfe-4c2e-9801-167b8ec3a8a8 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Benchmark profiling: Mechanistic diagnosis of LLM benchmarks, 2025
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3525567a-a7c4-416d-b5e0-5c026955eac2 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Had- field, Lukas Heim, Marianela Rodriguez, Jonas B
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abb5f37e-6155-4945-96e0-c23b990aeb37 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Richard Landis and Gary G
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a58b9389-fa75-416b-b92b-fe1336010982 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Towards explainable evaluation metrics for machine translation.Journal of Machine Learning Research, 2024
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c99d7223-2015-4d88-a724-58467ba04a48 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Frangi, Antonio R
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84f8a9e-b3a1-44d0-a18d-336e07395aad · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Holistic Evaluation of Language Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fa987c16-36fc-4970-98eb-7890c5f31ba9 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Manning, Christopher Ré, Diana Acosta-Navas, Drew A
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab6b7e59-b904-4f46-b24d-3c8807e41fa3 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Are we learning yet? a meta review of evaluation failures across machine learning
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a96ef9-77b9-4134-88d0-e5db5a517309 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting A safe harbor for AI evaluation and red teaming
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b25eb4e2-3da2-4935-97cf-bf1dda8b9930 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting LLM cyber evaluations don’t capture real-world risk,
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd764713-2be5-4852-b2a9-32a041fabb52 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting LLM Cyber Evaluations Don't Capture Real-World Risk
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7404c52-2c9e-4187-8065-84534df7ec12 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Data contamination: From memorization to exploitation
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6790727a-79ba-4a09-953a-7d2c552da715 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Building less-flawed metrics: Understanding and creating better measurement and incentive systems.Patterns, 2023
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aae7fab-7a26-4247-9fd2-fda3ad46b397 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Oard, Luca Soldaini, Ian Soboroff, Orion Weller, Efsun Kayi, Kate Sanders, Marc Mason, and Noah Hibbler
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f0053b5-d011-4038-8aaf-224e8c215a18 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e96cffc0-b2b3-4478-b5be-ecdd3c63ba0c · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Adding error bars to evals: A statistical approach to language model evaluations,
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74551f00-3c90-46f7-ba22-fef4cc9354a3 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aa00e479-e585-4cb3-a540-baa559133031 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Model cards for model reporting
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d68ccd6-9171-44d5-b05f-9d322032e90c · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting State of What Art? A Call for Multi-Prompt LLM Evaluation
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ba2db502-f4e6-4dc8-9bb7-9b329e9e870b · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Extrinsic evaluation of machine translation metrics
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d3c5ae-5318-4bc0-bbe8-66ad1d569c5b · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Unresolved cited work
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4546810-6750-4ab7-ab34-ec27b1103b63 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting A Survey on Large Language Model Benchmarks
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f5e1abd3-f75f-4508-bd04-b077584b989f · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Evaluation of DeepSeek AI models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eda49157-95e1-4994-ac0a-88e683ced263 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Unresolved cited work
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 447529a4-fd12-4e37-8abf-2c6f998706cb · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Byun, Kevin Wei, and Toby Webster
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf33c5a0-7ccf-490a-9ae6-53808fd41cfb · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Unresolved cited work
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8ddff62-d616-403d-9d3b-f918f010c1ca · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Toward best practices for AI evaluation and governance: A proposal for a european union general-purpose AI model evaluation standards task force
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81a91e57-c3b8-4afd-ba90-afd8f2d8b6eb · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Data cards: Purposeful and transparent dataset documentation for responsible ai
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34d37655-022d-4497-b11e-0505caafc797 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting The AI Model Risk Catalog: What Developers and Researchers Miss About Real-World AI Harms
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 34f6a5ad-d167-4e2d-b9d4-a22bf6d56317 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Mellor, Jonathan Uesato, Po-Sen Huang, Johannes Welbl, Laura Weidinger, Sumanth Dathathri, Amelia Glaese, Geoffrey Irving, Iason Gabriel, William Isaac, and Lisa Anne Hendricks
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34045775-d44e-43ae-bedc-d26990cc8319 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2327c442-bf51-4758-a333-70bcceac9caf · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Kochenderfer
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76033481-4fb4-41c0-90c1-3edb3c4d1ca4 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Measur- ing what matters: Connecting ai ethics evaluations to system attributes, hazards, and harms
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e89ba13-7217-408b-81c6-8bc2f39aeee6 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Measurement to meaning: A validity-centered framework for AI evaluation, 2025
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78190728-5411-439c-a543-619cec0a4454 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Unresolved cited work
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 241f9a46-824b-4056-b059-6e49503cc97a · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Unresolved cited work
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e2dd374-9a2f-4173-96dd-0a61d0dfe4e9 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Reality check: A new evaluation ecosystem is necessary to understand ai’s real world effects, 2025
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6616e616-cfbf-45a5-802e-8df446a2038d · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Improving methodologies for agentic evaluations across domains: Leakage of sensitive information, fraud and cybersecurity threats, 2026
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aad5fe4-fc00-42c9-9231-205c1d2855b6 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Model evaluation for extreme risks
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 55f2fff8-51cc-4ae6-958c-cbabd9dabc08 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02c5708e-1869-4364-9c27-a627fa4201ce · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Daly, Michael Hind, David Piorkowski, Xiangliang Zhang, Nuno Moniz, and Nitesh V Chawla
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83d0a000-3cbf-4c26-9c86-6e7225b07dc3 · outbound
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Verifiable evaluations of machine learning models using zkSNARKs
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.