Pith. sign in

Paper Citation Record · LEDGER

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

As of 5 August 2026, this Paper Citation Record lists 100 of 155 outbound references and 0 inbound Pith citation observations for arXiv:2606.09809.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.09809 v1

Coverage vector

measured 100 of 155 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T16:11:36.483820Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 155 outbound references displayed

  • verified exact27
  • verified fuzzy0
  • unresolved66
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7b245c6e-5c7c-42ae-8c18-1581876bfa32 · outbound

This paper cites Developing and maintaining an open- source repository of AI evaluations: Challenges and insights.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Developing and maintaining an open- source repository of AI evaluations: Challenges and insights

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:2412c6a32a366791825b1d4f03ae09212d4ae6e3e80bb8d8090ae45a7fa846ec

Observation 768ab1bc-4dd9-4f56-856c-e3874910ec0a · outbound

This paper cites When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:57:32.210991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:aa999b1c50d7b1b5ff99f415e24db53852fbb363a0d92db11d9e068069fd4256

Observation 09a70119-7f73-4524-a805-c41a03f73f33 · outbound

This paper cites Audit and Assurance of AI Algorithms: A framework to ensure ethical algorithmic practices in Artificial Intelligence.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Audit and Assurance of AI Algorithms: A framework to ensure ethical algorithmic practices in Artificial Intelligence

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.195823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:5fbda812138a7eb63d50e74921323105576f8bfd16fe524c8f0365bc4fb0369e

Observation 300eba02-62bf-4e26-a423-27cb2c662207 · outbound

This paper cites Lessons from the trenches on evaluating machine-learning systems in materials science.Computational Materi- als Science, 2025.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Lessons from the trenches on evaluating machine-learning systems in materials science.Computational Materi- als Science, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:77a0531861c6e799665d01bb89c9ee4998e2eb1422acbdf52d7b621b3ec894c8

Observation 07de6629-35e0-4160-aa89-04d53f06453f · outbound

This paper cites AllMetrics: A Unified Python Library for Standardized Metric Evaluation and Robust Data Validation in Machine Learning.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting AllMetrics: A Unified Python Library for Standardized Metric Evaluation and Robust Data Validation in Machine Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.204731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:ef532b9faf24c7b328335546c6e6b136a88beb113a1d8fbfb2264d51151b73fa

Observation fddfe3a7-b55c-4ae0-9750-b4d84e401043 · outbound

This paper cites Arnstein.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Arnstein

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:fe5fbc884946b26513596ca8acba7f004d10bf187e81e9e0e77590898ba7b3c6

Observation 7d739621-4023-464d-a898-50de56e202e6 · outbound

This paper cites Comparison of AI models across intelligence, performance, and price,.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Comparison of AI models across intelligence, performance, and price,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:141d48c1e23fbabb0c048ba226c041526a551077428a474d4de6da74bd3e0e2d

Observation b5ca1dac-2be0-4b75-85b2-5cd6b57304d3 · outbound

This paper cites an unresolved cited work.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:3912f17e111930240373b63f8d059925fbcbc4f13ff726aa32a24a07d11b5550

Observation 17a62291-90e3-4113-b81b-97ecc6dfe3ba · outbound

This paper cites AI Risk Atlas: Taxonomy and Tooling for Navigating AI Risks and Resources.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting AI Risk Atlas: Taxonomy and Tooling for Navigating AI Risks and Resources

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.198709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:257a5cb04921ef7056f5333b7d57af8e5bd47ef48b5c524e2d257dbefdf7c5d2

Observation ea6ae664-fb70-4f10-85d7-0fffb38b3ebc · outbound

This paper cites Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:f0e644af8576b46551a79f557a899912e2a26213c3b16adea53d43ee0761d6e0

Observation 2e6c38c0-6447-45b6-9866-3d149e859c37 · outbound

This paper cites When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.207757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:f2a846ba0129db19fe3162d0043e67aab64cfe31017c02d5ddaf1b7a56d23516

Observation 220186bc-b6cf-4092-87cc-f29cb8de6b89 · outbound

This paper cites Every eval ever: Toward a common language for AI eval reporting.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Every eval ever: Toward a common language for AI eval reporting

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:b23c5155acf7be94b165d61cab445cb07361581c0d7ca07ffa797e85b3794080

Observation c79f37c2-f434-48e5-8986-bb51d6b42e8d · outbound

This paper cites arXiv preprint arXiv:2511.04703; presented at Neurips 2025 , year=.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting arXiv preprint arXiv:2511.04703; presented at Neurips 2025 , year=

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.638182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:ee30f1ee170074e37d62577074eb49af95e60a78cb13d79f4ed7e367da801b96

Observation 76c7d314-f787-4bf5-903c-8c0438e44c77 · outbound

This paper cites Absolute Evaluation Measures for Machine Learning: A Survey.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Absolute Evaluation Measures for Machine Learning: A Survey

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.294769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:d1c5e9954748f61b7d000c96cc9739837a25cad99e3d4fe1bbd8269ac0f5a383

Observation f60233f5-bd3c-45ba-bf68-fe3356b03fd0 · outbound

This paper cites Open llm leaderboard (2023- 2024).

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Open llm leaderboard (2023- 2024)

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:86544f4f2a212418ade7f2af73a3a9654b26355c0c976b8f4e0cd18aa337daea

Observation 8159c697-2f6f-473f-a423-8e39246aa914 · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T02:07:33.240679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:e713833f9ed735f1fb25804cdf651859d3dd778f458740d6eb57e76961f01cb8

Observation facbbe95-f13d-4bde-b884-df235d47cfd0 · outbound

This paper cites A metrological framework for uncertainty evaluation in machine learning classification models.Metrologia, 2025.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting A metrological framework for uncertainty evaluation in machine learning classification models.Metrologia, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:524615c0716ec2edd5e44c05b81eadec99148da507f35415aba5c3595d49f18d

Observation 52ce2210-9dc6-4fbe-af4d-2125cc502f24 · outbound

This paper cites The impact of standardisation and standards on innovation.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting The impact of standardisation and standards on innovation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:54bfb58bcc58a70648ba491c7ea520afa6c524f7e9705c358382fc614393612f

Observation 41d0cab6-d5b4-422c-936d-348cc8619a9e · outbound

This paper cites Assessing ai: Surveying the spectrum of approaches to understanding and auditing ai systems, 2025.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Assessing ai: Surveying the spectrum of approaches to understanding and auditing ai systems, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:c380a5bb58b3c0d1eb07717b3cc7bddb0c9afb18df86b9e08b54a3726c661592

Observation 3b8a0ff5-d55f-4092-b74e-5c7c4aa83d55 · outbound

This paper cites Evaluation for change.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Evaluation for change

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:54be811e7f74c0fb5b24e542182bb6d94586e73f92c75bf5e99f73985ad66260

Observation b784dba0-3e0f-4119-9cf7-790e0f2aeffa · outbound

This paper cites Bordes, C.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Bordes, C

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T02:07:33.244421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:a74d5ce91d59d2935fa6adaddc15c2f3efe67b09f3afd1c5ee7ca3c0a03cb90b

Observation 3756fd09-6a51-404c-a723-06998069018d · outbound

This paper cites an unresolved cited work.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:9c574cee5b150b3837f520be25d554efd7d30f67880f2744a2bdd3b50235ecbc

Observation 809552e1-9e2c-49d6-a9ce-34459b9f84be · outbound

This paper cites Cohn, and Jose Hernandez- Orallo.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Cohn, and Jose Hernandez- Orallo

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:112868fb577d3e2deda7ae613fa0a11b998725c56858c3a519ddbd64b59fff90

Observation 9562ffa5-7302-4efd-937f-cd72e94c3bb2 · outbound

This paper cites best fit.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting best fit

Reference 24

Resolution
malformed identifier
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:27f4e27460a79e6b03bb69ff9c118aa62438221be219d62e4e9f10b0348c99c3

Observation a8fcc0fc-ff7a-4583-9ce0-3a662040d04d · outbound

This paper cites Black-box access is insufficient for rigorous AI audits, 2024.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Black-box access is insufficient for rigorous AI audits, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:628772fa13b4751a454a6eac052bd79d15ab1be049c989ab6804935d9a3c1444

Observation fea5f83f-3ffc-4564-a951-9d7608954e35 · outbound

This paper cites ISBN 978-1-4503-7110-0.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting ISBN 978-1-4503-7110-0

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T16:21:02.278927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:4318dc2517b9fb00bd41a3bfb43f635b5a7264eca6b6d44861b68b4a1e85b6d8

Observation d5236e58-a0c7-4cc9-a54f-dd3d2d21c171 · outbound

This paper cites Managing misuse risk for dual-use foundation models.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Managing misuse risk for dual-use foundation models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:556e97b9df937c1ffeca01d12ee7151bbc9e0b9a8fcdd1e62ac79e4d6c4f8eb4

Observation c7ff8246-6aca-4a75-a1ad-594b2a8564d1 · outbound

This paper cites Test & Evaluation Best Practices for Machine Learning-Enabled Systems.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Test & Evaluation Best Practices for Machine Learning-Enabled Systems

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.273535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:7c793077f2329ddba8d87b1c930114f1eb4a8fc9455ee8e46637490439011482

Observation 738ae33f-8811-4c89-8585-260ba7a98ee0 · outbound

This paper cites Evaluating machine expertise: How graduate students develop frameworks for assessing GenAI content, 2025.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Evaluating machine expertise: How graduate students develop frameworks for assessing GenAI content, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:ae6a54826e4d498dd08ff4d8ad7ac2984d0bb1220f4c248c426a00c817ffb2ce

Observation 613320f0-72c3-489b-ae48-3616f67f4de5 · outbound

This paper cites Gonzalez, and Ion Stoica.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Gonzalez, and Ion Stoica

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:a724d1fa62a090d11d010c86ebd1afba1ef4bb4fbf8c250fae5cea87324002a2

Observation c7dd991b-4739-4eaf-a6d5-4b1e569a3d21 · outbound

This paper cites Collins, Karel G.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Collins, Karel G

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:e03e5d47cff18ef84f408f12652329d69229de4587b24437ddc2803129ca418c

Observation 4728629b-1ad2-4939-99cc-83d70a6f65f9 · outbound

This paper cites Who Audits the Auditors? Recommendations from a field scan of the algorithmic auditing ecosystem.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Who Audits the Auditors? Recommendations from a field scan of the algorithmic auditing ecosystem

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.270816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:ab4a9511ed49b3dd4d2bbc710f0e701a408d5a489afcd0d8e0ce3ad564d7351b

Observation 52dda918-d769-4483-b073-651ec09383ea · outbound

This paper cites Evalcards: A framework for standardized evaluation reporting, 2025.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Evalcards: A framework for standardized evaluation reporting, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:aed6a6dabfe6eabca21291085e04f329983eca603e004ab1688c42fb232b7428

Observation 5287ce39-eb69-4e26-8f6b-e5a182d71ff0 · outbound

This paper cites Nahab, and Xiao Hu.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Nahab, and Xiao Hu

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:3762eb78eaf036bcc795980989d2b2ee9f07ad2bd8e7b81f7d40e88de93beb9b

Observation a57e7f02-52f4-4b32-bcc2-485640696c9e · outbound

This paper cites Reconsideration on evaluation of machine learning models in continuous monitoring using wearables.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Reconsideration on evaluation of machine learning models in continuous monitoring using wearables

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.635442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:8d3014e5bd9daa1fa366d0cc582251337cfdbbe7e60cbcf503e54a09ecc7cc31

Observation de13c229-28b5-460e-8bc0-67b9e5c9cf20 · outbound

This paper cites Introducing Epoch AI’s AI benchmarking hub, 2024.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Introducing Epoch AI’s AI benchmarking hub, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:86e1854ddd5aa7dee5ac084ccb1d74959e2fec697f0aa2292694afa53db87afd

Observation 2b4b035c-af93-4529-9ec6-f64841e35e2b · outbound

This paper cites Can we trust AI benchmarks? An interdisci- plinary review of current issues in AI evaluation.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Can we trust AI benchmarks? An interdisci- plinary review of current issues in AI evaluation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:6741a31be2baa9456858c86620cbc912d97aed2c5cabc9a8d838a89a0415a5c4

Observation 53ada467-ae6f-44a5-a908-1db68dcba33c · outbound

This paper cites The general-purpose AI code of practice: Safety & security chapter, July 2025.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting The general-purpose AI code of practice: Safety & security chapter, July 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:a688c60b8abc9f36f39bae325485816c479f71e9efdbdde046d0e68b59b24441

Observation f400d197-4754-4cab-bf45-ba0286b68d98 · outbound

This paper cites The general-purpose AI code of practice: Trans- parency chapter, July 2025.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting The general-purpose AI code of practice: Trans- parency chapter, July 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:d0c5c08beaaed62b11290de24d532bb09d8c1523a69b019d2c9d4e6ee1389762

Observation ae0b8826-358a-44c9-ba96-fc0b3c821c9c · outbound

This paper cites EvalEval: Every eval ever shared task, 2024.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting EvalEval: Every eval ever shared task, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:5b7b7cd4102f03c74e7c030f7c29b139ca7fe670d8aafd620c96c2f7621e3c2e

Observation 9b10e715-edaf-4f83-8432-d3c49d9e84c8 · outbound

This paper cites Good practices for evaluation of machine learning systems.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Good practices for evaluation of machine learning systems

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.629761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:4627405bf0c7e31677f177d19243df71aec1faa7e1d4b63a0deea2d72ccda2f3

Observation 16a21a07-681d-48d3-8186-bdb157931f7d · outbound

This paper cites Frontier capability assessment.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Frontier capability assessment

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:29af3a03e930b421c30e45b38e210709c04043a332f9129687e006bbadfa8b62

Observation e977c2d2-ef7f-4c93-8630-5dc0dd0dd403 · outbound

This paper cites Datasheets for datasets.Commun.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Datasheets for datasets.Commun

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:5316fea72c67cb2b8c4e10c36669952e2fa46f0bf1b804f7f62fb6c0fe69cb1e

Observation 69090bd6-e9ac-4361-bbef-53f3c68e9e69 · outbound

This paper cites Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text.Journal of Artificial Intelligence Research, 2023.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text.Journal of Artificial Intelligence Research, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:2165f02a411af09ce24db02e24be441bf029a96515f3a941e1874f93582cc60f

Observation 187eaf4c-0767-4ac9-9c9c-5e6dd4e50397 · outbound

This paper cites AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.632648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:fa4d8dcf6510eb259c48b062833fee2fe5943d7e28eddc869658dc381319f3a6

Observation 5247466e-5473-45ad-8301-eddc3cce436b · outbound

This paper cites Stress-Testing Capability Elicitation With Password-Locked Models.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Stress-Testing Capability Elicitation With Password-Locked Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T02:07:33.261481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:b2507fb44567dccdc9e240dc6aa62daa9f90d4f5f4a038d4536131cbd3ecbf92

Observation 6d2dc62f-528e-4866-81d7-d551b8797e4c · outbound

This paper cites Olmes: A standard for language model evaluations.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Olmes: A standard for language model evaluations

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:332cf294e7031371fcfc9d329be94c2d9d90ade62bc7c40ca158cc86da9ab225

Observation 0742dfa1-161e-4604-b4f6-09fc8b534b28 · outbound

This paper cites Gupta, Jessica Hullman, and Hari Subramonyam.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Gupta, Jessica Hullman, and Hari Subramonyam

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:4ffedbb83e9588b8c898a10036d3b024f766b0c32f07461e55ea996937526e68

Observation a4dba60e-6eb6-47b0-b0a5-a631bcea00bc · outbound

This paper cites System Cards for AI-Based Decision-Making for Public Policy.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting System Cards for AI-Based Decision-Making for Public Policy

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.624032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:7fc95edaab8e0396c2ab068605034cc9eea395508e36825f3412b97d0c18682c

Observation f6cabc94-c87a-4c08-90a0-e8c0dc6df617 · outbound

This paper cites Empirical Privacy Evaluations of Generative and Predictive Machine Learning Models -- A review and challenges for practice.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Empirical Privacy Evaluations of Generative and Predictive Machine Learning Models -- A review and challenges for practice

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.280148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:47437793968f923a75196e31ff412192677c00899bf26ae04ce08b2a439cdb74

Observation 88aeeb68-936f-43be-a5db-b62d9ca0884f · outbound

This paper cites Bernstein, and Mykel John Kochenderfer.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Bernstein, and Mykel John Kochenderfer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:d8e465e33f57f09f04ccd27a6a156cbf2df102ecd702f78443e9fd2ed87f744e

Observation e2d40dca-467b-4097-86e9-7d10adf88655 · outbound

This paper cites A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:ceab897da366525fe4ec743e117e9c963162e217298725b734a75ebdbe4bb024

Observation 846636de-5d90-4d31-a485-24fe2a3e4881 · outbound

This paper cites Auto-benchmarkcard: Automated synthesis of benchmark documentation.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Auto-benchmarkcard: Automated synthesis of benchmark documentation

Reference 53

Resolution
verified exact
doi, observed 2026-06-27T16:21:02.272109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:39e733530ab41ddc5a07a80530f0d2cad15f98baa1b3c84689040bc6df69c312

Observation e7d06c8c-206c-475b-a101-3bc49ba1521a · outbound

This paper cites Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.627001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:5ab12ae9996f40ce6bc89bdfe8472f143efa4e1ad7c50da118b308930ef95328

Observation c838f560-df5b-48bf-b216-70c9b87f9898 · outbound

This paper cites Evaluation gaps in machine learning practice.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Evaluation gaps in machine learning practice

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:69b6db8e6f064e4258846310bf021b902698329605f1fade6812cb7ba87a5800

Observation 78aaf0f0-e42a-45c9-a479-6b23b4a90f84 · outbound

This paper cites Rethinking Machine Learning Model Evaluation in Pathology.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Rethinking Machine Learning Model Evaluation in Pathology

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.620424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:5e2007c9e2b8d9c2d1b71f855eb57fbaea4b36cdc1aad5a813db4d49be0303d4

Observation 10fa77d5-1123-4b07-b294-40ba69195ff5 · outbound

This paper cites Deprecating benchmarks: Criteria and framework.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Deprecating benchmarks: Criteria and framework

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:70063c4196d1b708cc839df30f61ff77707302d47d250e67ec8c19abd84c0ac8

Observation 5ec1306f-fbf5-46f7-8e94-18bfe0825f5d · outbound

This paper cites Cantrell, Keiran Peng, Thanh Huy Pham, Christopher A.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Cantrell, Keiran Peng, Thanh Huy Pham, Christopher A

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:afa66dbfd70ca2d3e0d663ee9ef35585419e23545dc8d90e0f7e288db3d7804e

Observation 8c8c2b13-2bfe-4c2e-9801-167b8ec3a8a8 · outbound

This paper cites Benchmark profiling: Mechanistic diagnosis of LLM benchmarks, 2025.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Benchmark profiling: Mechanistic diagnosis of LLM benchmarks, 2025

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.296898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:35270a62a195d7f379275ea96d58d2b2c8f0f4cfd2cf2c6419166301a747d8ee

Observation 3525567a-a7c4-416d-b5e0-5c026955eac2 · outbound

This paper cites Had- field, Lukas Heim, Marianela Rodriguez, Jonas B.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Had- field, Lukas Heim, Marianela Rodriguez, Jonas B

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:77b544f1592079d383060f6277cbb5469354a8ef7f5e66e9054585d924b5a415

Observation abb5f37e-6155-4945-96e0-c23b990aeb37 · outbound

This paper cites Richard Landis and Gary G.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Richard Landis and Gary G

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:07058190172c0338be6687176503c4cf4b97aa8c775cdf4e85ec33cce7aa836a

Observation a58b9389-fa75-416b-b92b-fe1336010982 · outbound

This paper cites Towards explainable evaluation metrics for machine translation.Journal of Machine Learning Research, 2024.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Towards explainable evaluation metrics for machine translation.Journal of Machine Learning Research, 2024

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:9b7fe33936df1358fc2ee0e43e99a67b4e96b0e6bb31fd0fd5d8524f92d04358

Observation c99d7223-2015-4d88-a724-58467ba04a48 · outbound

This paper cites Frangi, Antonio R.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Frangi, Antonio R

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:b423c4bee0311c2f0bc66438a87168880d07e858f46846b05234f22c224b50e0

Observation b84f8a9e-b3a1-44d0-a18d-336e07395aad · outbound

This paper cites Holistic Evaluation of Language Models.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Holistic Evaluation of Language Models

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T02:07:33.254030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:3d29eac4250227257d931d43a0451a0713509400fe8b09b693e2f5849cd608da

Observation fa987c16-36fc-4970-98eb-7890c5f31ba9 · outbound

This paper cites Manning, Christopher Ré, Diana Acosta-Navas, Drew A.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Manning, Christopher Ré, Diana Acosta-Navas, Drew A

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:6c0bc6a6f760e2cd5eb269dc1cd1aea904b6b852a63f5e19a0c6dc279e6b3c81

Observation ab6b7e59-b904-4f46-b24d-3c8807e41fa3 · outbound

This paper cites Are we learning yet? a meta review of evaluation failures across machine learning.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Are we learning yet? a meta review of evaluation failures across machine learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:b50194345f59c06ad508e8cc20828ac51604ad4d6ebfc880fe20907b1ba2cd05

Observation f7a96ef9-77b9-4134-88d0-e5db5a517309 · outbound

This paper cites A safe harbor for AI evaluation and red teaming.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting A safe harbor for AI evaluation and red teaming

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:bd9a3b7006db2986ea7e7ad6b1f79f6de384439690935143b862e97da5953402

Observation b25eb4e2-3da2-4935-97cf-bf1dda8b9930 · outbound

This paper cites LLM cyber evaluations don’t capture real-world risk,.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting LLM cyber evaluations don’t capture real-world risk,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:4556fea2185869e98ec9e3207c8e6c5ea21aa610e4ca439ff4a13e08d577df3f

Observation dd764713-2be5-4852-b2a9-32a041fabb52 · outbound

This paper cites LLM Cyber Evaluations Don't Capture Real-World Risk.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting LLM Cyber Evaluations Don't Capture Real-World Risk

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.283723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:08c411295e3530cb894da6761c972b588194d3a6dc819cd0c97b89ecb8865bfe

Observation e7404c52-2c9e-4187-8065-84534df7ec12 · outbound

This paper cites Data contamination: From memorization to exploitation.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Data contamination: From memorization to exploitation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:8f5e643448946ed0f6abee14626ba017b35983568a9b362788ed819b7c0291aa

Observation 6790727a-79ba-4a09-953a-7d2c552da715 · outbound

This paper cites Building less-flawed metrics: Understanding and creating better measurement and incentive systems.Patterns, 2023.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Building less-flawed metrics: Understanding and creating better measurement and incentive systems.Patterns, 2023

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:94488d00f40a2f3151275473e4a226dfdc0796065f3e72a63c11cd78935c4f05

Observation 3aae7fab-7a26-4247-9fd2-fda3ad46b397 · outbound

This paper cites Oard, Luca Soldaini, Ian Soboroff, Orion Weller, Efsun Kayi, Kate Sanders, Marc Mason, and Noah Hibbler.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Oard, Luca Soldaini, Ian Soboroff, Orion Weller, Efsun Kayi, Kate Sanders, Marc Mason, and Noah Hibbler

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:3e4709657db3083dca74e6264cd022ace948ba4e2b2a32c179867ef932688eab

Observation 5f0053b5-d011-4038-8aaf-224e8c215a18 · outbound

This paper cites STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.277542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:ad7155c92a00bc50ca2aff2a6645f2e3014044d7f9e51aadaf278fa5fc09341a

Observation e96cffc0-b2b3-4478-b5be-ecdd3c63ba0c · outbound

This paper cites Adding error bars to evals: A statistical approach to language model evaluations,.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Adding error bars to evals: A statistical approach to language model evaluations,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:5532745b2f0c959959efbff92022e6f275d6fc45438dd085e0a7504d437160c4

Observation 74551f00-3c90-46f7-ba22-fef4cc9354a3 · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.250825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:9baa2675c73ca5cfe705a99dfe1683cede6b00c49523dfa2333364294aca0800

Observation aa00e479-e585-4cb3-a540-baa559133031 · outbound

This paper cites Model cards for model reporting.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Model cards for model reporting

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:c9050aa139997350c52fbfa4e75652e56d7334472c055a3c6046b6b04a52030f

Observation 7d68ccd6-9171-44d5-b05f-9d322032e90c · outbound

This paper cites State of What Art? A Call for Multi-Prompt LLM Evaluation.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T02:07:33.257770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:5abbce715321ff9006bf4de8d7dde8e416d3e64f812498ff1826dc787394c540

Observation ba2db502-f4e6-4dc8-9bb7-9b329e9e870b · outbound

This paper cites Extrinsic evaluation of machine translation metrics.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Extrinsic evaluation of machine translation metrics

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:c9d8db5b16713f8b51fe47ac19b1a552d33d6ade1dd42bdcdb09db71ea3cdba7

Observation f4d3c5ae-5318-4bc0-bbe8-66ad1d569c5b · outbound

This paper cites an unresolved cited work.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:b0cbcc639221190abc1ed3177eab208ab8429b825dbe4e5512845acfd235392b

Observation b4546810-6750-4ab7-ab34-ec27b1103b63 · outbound

This paper cites A Survey on Large Language Model Benchmarks.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting A Survey on Large Language Model Benchmarks

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.617219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:81c10c17473be22872fadd16666b002b4557b7d35aedfc05394131bd1d093be3

Observation f5e1abd3-f75f-4508-bd04-b077584b989f · outbound

This paper cites Evaluation of DeepSeek AI models.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Evaluation of DeepSeek AI models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:5809f987df73cb21579793760e1d3158ee0328ad27ebfc3c071dde0959b5bb2f

Observation eda49157-95e1-4994-ac0a-88e683ced263 · outbound

This paper cites an unresolved cited work.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:fee31c61aa0e525b9b8ee784c939e14b2dc0fd9fbb02debc23138c7ae02eaf38

Observation 447529a4-fd12-4e37-8abf-2c6f998706cb · outbound

This paper cites Byun, Kevin Wei, and Toby Webster.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Byun, Kevin Wei, and Toby Webster

Reference 83

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:8c09160ac5e60f2ee23bef04bcf2f88814c1cb090606357a0c01e9b5b1e6e57e

Observation bf33c5a0-7ccf-490a-9ae6-53808fd41cfb · outbound

This paper cites an unresolved cited work.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:00aecbde5d27c1030a71c72e90298535a80ab30d8845c988ae7182f21384eb8f

Observation f8ddff62-d616-403d-9d3b-f918f010c1ca · outbound

This paper cites Toward best practices for AI evaluation and governance: A proposal for a european union general-purpose AI model evaluation standards task force.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Toward best practices for AI evaluation and governance: A proposal for a european union general-purpose AI model evaluation standards task force

Reference 85

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:d215cea1b01009b69b8eb195ec12f4c65404757524ece6a799f848a5db87972b

Observation 81a91e57-c3b8-4afd-ba90-afd8f2d8b6eb · outbound

This paper cites Data cards: Purposeful and transparent dataset documentation for responsible ai.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Data cards: Purposeful and transparent dataset documentation for responsible ai

Reference 86

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:31204ced8b7442d092420ec34a161f07f703a6fc8943f653f3fddd5e60a1ace0

Observation 34d37655-022d-4497-b11e-0505caafc797 · outbound

This paper cites The AI Model Risk Catalog: What Developers and Researchers Miss About Real-World AI Harms.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting The AI Model Risk Catalog: What Developers and Researchers Miss About Real-World AI Harms

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.246972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:1e7c2bc8874694351db76da07b8cad51d299570eff58e9d25a974eed5aa54e88

Observation 34f6a5ad-d167-4e2d-b9d4-a22bf6d56317 · outbound

This paper cites Mellor, Jonathan Uesato, Po-Sen Huang, Johannes Welbl, Laura Weidinger, Sumanth Dathathri, Amelia Glaese, Geoffrey Irving, Iason Gabriel, William Isaac, and Lisa Anne Hendricks.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Mellor, Jonathan Uesato, Po-Sen Huang, Johannes Welbl, Laura Weidinger, Sumanth Dathathri, Amelia Glaese, Geoffrey Irving, Iason Gabriel, William Isaac, and Lisa Anne Hendricks

Reference 88

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:f07abe65fc760a4f6523b6d4fa8e3100b45d10b2c6aa65a5bfbb3d1d54c4eaa0

Observation 34045775-d44e-43ae-bedc-d26990cc8319 · outbound

This paper cites Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.268541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:9a630fd01bd61a7c6cd77a338dd8c08c59061566bef0e2836f67ba7e0970e462

Observation 2327c442-bf51-4758-a333-70bcceac9caf · outbound

This paper cites Kochenderfer.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Kochenderfer

Reference 90

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:01e0f98d55998b47eafb560194e390f52cd1c5095b4b668fdebc7c18374c833d

Observation 76033481-4fb4-41c0-90c1-3edb3c4d1ca4 · outbound

This paper cites Measur- ing what matters: Connecting ai ethics evaluations to system attributes, hazards, and harms.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Measur- ing what matters: Connecting ai ethics evaluations to system attributes, hazards, and harms

Reference 91

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:1926a0fd691a78df974b9aac00a6fa3c8619a4344b0e75abc2bf4e7d48381181

Observation 2e89ba13-7217-408b-81c6-8bc2f39aeee6 · outbound

This paper cites Measurement to meaning: A validity-centered framework for AI evaluation, 2025.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Measurement to meaning: A validity-centered framework for AI evaluation, 2025

Reference 92

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:c46d9cef365487caec57874553b1bd5bfd7880ea96de9338046ea1392499cd95

Observation 78190728-5411-439c-a543-619cec0a4454 · outbound

This paper cites an unresolved cited work.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Unresolved cited work

Reference 93

Resolution
malformed identifier
doi_truncated, observed 2026-06-27T16:21:02.273864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:81c4f1984f9416be0ae9eb8358470f4dbc11a93f7d54b79cc58d6e7867e84581

Observation 241f9a46-824b-4056-b059-6e49503cc97a · outbound

This paper cites an unresolved cited work.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Unresolved cited work

Reference 94

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:7b2d6798a27833bb3c8831013bbd6264b8535261245c742caa73d7a35a87e6ad

Observation 1e2dd374-9a2f-4173-96dd-0a61d0dfe4e9 · outbound

This paper cites Reality check: A new evaluation ecosystem is necessary to understand ai’s real world effects, 2025.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Reality check: A new evaluation ecosystem is necessary to understand ai’s real world effects, 2025

Reference 95

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:76b07402ae3b9afbe4893bee0ad79843312ee01f907370c879b37806695fc793

Observation 6616e616-cfbf-45a5-802e-8df446a2038d · outbound

This paper cites Improving methodologies for agentic evaluations across domains: Leakage of sensitive information, fraud and cybersecurity threats, 2026.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Improving methodologies for agentic evaluations across domains: Leakage of sensitive information, fraud and cybersecurity threats, 2026

Reference 96

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:17718c98cba35959788ccde15a597034e060e54e30ec85da543f770e12fb11f6

Observation 2aad5fe4-fc00-42c9-9231-205c1d2855b6 · outbound

This paper cites Model evaluation for extreme risks.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Model evaluation for extreme risks

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.287703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:1ae4cda0a038fe4e8b4df273ac4d592150eb5dbb4b3f0e3dc745fb86a3d5c0ee

Observation 55f2fff8-51cc-4ae6-958c-cbabd9dabc08 · outbound

This paper cites Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker

Reference 98

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:a8a149db0bf89e23943f1b2a286ee3db6257ea4dce8294a715bfa73596804cb9

Observation 02c5708e-1869-4364-9c27-a627fa4201ce · outbound

This paper cites Daly, Michael Hind, David Piorkowski, Xiangliang Zhang, Nuno Moniz, and Nitesh V Chawla.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Daly, Michael Hind, David Piorkowski, Xiangliang Zhang, Nuno Moniz, and Nitesh V Chawla

Reference 99

Resolution
unresolved
no resolver link, observed 2026-06-27T16:11:36.483820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:5e6a8c850126c47c76358cd685035a08536dec2a6580ffe3a14b160575723f85

Observation 83d0a000-3cbf-4c26-9c86-6e7225b07dc3 · outbound

This paper cites Verifiable evaluations of machine learning models using zkSNARKs.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Verifiable evaluations of machine learning models using zkSNARKs

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.614700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:5578a55c4304e0038dc0d18c0c68a4d76457b37eac5a8fd870076d49e52c74d5

Pith citing papers

No inbound Pith citation observations are available.