Pith. sign in

REVIEW 2 major objections 2 minor 155 references

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

T0 review · 2 major / 2 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read EvalCards composes benchmark metadata, evaluation run data, and model metadata into unified records that supply four interpretive signals.

desk verdict EvalCards gives a workable schema and reader modes for AI eval reports plus real deployment numbers, but supplies no test that the signals actually improve interpretation. read the letter →

arxiv 2606.09809 v1 pith:CB25XQ62 submitted 2026-06-08 cs.AI

classification cs.AI
keywords AIevaluationreportingcardsreproducibilitybenchmarkmetadatamodeldocumentationcompletenessscorecomparabilityprovenance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs, so readers cannot reliably compare results, identify omissions, or trace aggregate claims to evidence. EvalCards addresses these gaps by deriving a reporting schema from a structured review of 52 papers and 10 stakeholder interviews, then implementing an operational layer that combines three data sources into one record. The record carries four signals—reproducibility, documentation completeness, provenance and risk, and score comparability—delivered through reader modes for research and non-research audiences. The system has been deployed to process 5,816 models, 635 benchmarks, and 101,843 results, exposing systematic shortfalls in current practice.

What carries the argument

The EvalCards unified record, which integrates benchmark metadata, evaluation run data, and model metadata and renders four signals through audience-calibrated reader modes.

What would settle it

Finding that a substantial share of new evaluation results still cannot be compared across sources or traced to evidence even after the EvalCards schema is applied because required details lie outside the defined fields.

Watch

Extended reading notes

Core claim

EvalCards is an operational reporting layer that composes benchmark metadata, evaluation run data, and model metadata into a unified record. It supplies four interpretive signals—reproducibility, documentation completeness, provenance and risk, and score comparability—rendered through reader modes calibrated to research and non-research audiences, and it has been deployed across 5,816 models, 635 benchmarks, and 101,843 results to surface gaps in existing reporting.

Load-bearing premise

A schema derived from reviewing 52 papers and interviewing 10 stakeholders will be comprehensive enough and widely adoptable to close the identified gaps in evaluation reporting.

Editorial extensions

If this is right

  • Readers can compare results across sources with greater reliability.
  • Omissions in individual reports become visible to users.
  • Aggregate claims can be traced back to their supporting evidence.
  • Systematic gaps in reporting practice across the field become measurable.
  • Different stakeholder groups receive interpretations matched to their questions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same schema could be extended to track how reporting completeness changes over successive model releases.
  • Automated extraction pipelines built on EvalCards might flag incomplete reports before they reach leaderboards.
  • The provenance signal could support audits that link published scores to specific training or evaluation conditions.
  • Adoption might encourage benchmark maintainers to align their output formats with the four signals from the start.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper presents EvalCards as an operational reporting layer that composes benchmark metadata, evaluation run data, and model metadata into a unified record. It derives a reporting schema from a structured review of 52 papers and 10 stakeholder interviews, implements four interpretive signals (reproducibility, documentation completeness, provenance and risk, and score comparability) rendered through reader modes for different audiences, and deploys a monitoring tool that applies the schema across 5,816 models, 635 benchmarks, and 101,843 results to surface gaps in current reporting practice.

Significance. If the schema and signals prove adoptable, EvalCards could reduce inconsistencies in AI evaluation reporting by providing a composable, stakeholder-calibrated interpretive layer. The scale of the deployment (thousands of models and results) is a concrete strength, demonstrating extraction infrastructure and empirically documenting reporting shortfalls across the field.

major comments (2)
  1. [Abstract and deployment section] Abstract and deployment section: the claim that the four signals 'surface systematic gaps' rests on descriptive statistics from the 101,843 results; the manuscript supplies no validation data, error analysis, inter-rater study, or comparison against expert judgments to show that the signals correctly identify the intended interpretive omissions.
  2. [Schema derivation] Schema derivation (literature review + interviews): while the process is described, the manuscript does not provide a traceable mapping from the 52 papers/10 interviews to the exact four signals or to the reader-mode distinctions, leaving open whether the schema is comprehensive or whether alternative signals were considered and rejected.
minor comments (2)
  1. Notation for the four signals is introduced without an explicit table summarizing their definitions, inputs, and reader-mode renderings; adding such a table would improve clarity.
  2. The manuscript cites the 52 papers and 10 interviews but does not include a supplementary table or appendix listing the reviewed sources or interview protocol.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback. We address each major comment below with clarifications and indicate revisions that will strengthen the presentation of the schema and deployment without overstating the signals' validation.

read point-by-point responses
  1. Referee: [Abstract and deployment section] Abstract and deployment section: the claim that the four signals 'surface systematic gaps' rests on descriptive statistics from the 101,843 results; the manuscript supplies no validation data, error analysis, inter-rater study, or comparison against expert judgments to show that the signals correctly identify the intended interpretive omissions.

    Authors: The deployment section uses descriptive statistics to illustrate how the signals, once operationalized, reveal patterns across the corpus; the primary aim is to demonstrate extraction infrastructure and the composable record rather than to validate the signals as classifiers. We agree that the current framing could be misread as implying validated detection. In revision we will rephrase the abstract and deployment claims to emphasize illustrative application, add an explicit limitations paragraph noting the absence of inter-rater or expert-comparison studies, and clarify that the signals function as reader aids derived from the schema. revision: partial

  2. Referee: [Schema derivation] Schema derivation (literature review + interviews): while the process is described, the manuscript does not provide a traceable mapping from the 52 papers/10 interviews to the exact four signals or to the reader-mode distinctions, leaving open whether the schema is comprehensive or whether alternative signals were considered and rejected.

    Authors: Section 3 outlines the review and interview protocol that generated the schema elements. To improve traceability we will add (in the main text or as supplementary material) a mapping table that connects specific themes from the 52 papers and interview notes to each of the four signals and to the reader-mode distinctions. The table will also note elements that were considered but deprioritized, thereby addressing concerns about comprehensiveness. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity identified

full rationale

The paper derives its reporting schema from an external structured review of 52 papers plus 10 stakeholder interviews, then implements four interpretive signals and applies them via a monitoring tool to independent data (5,816 models, 635 benchmarks, 101,843 results). No equations, self-definitions, fitted inputs renamed as predictions, or load-bearing self-citations appear in the derivation chain. The central composition and deployment steps remain independent of the paper's own outputs.

Assumptions & free parameters 0 free parameters · 1 assumptions · 1 invented entities

The central claim rests on the assumption that the literature review and interviews capture the necessary requirements; the main addition is the proposed framework itself rather than fitted parameters or new physical entities.

assumptions (1)
  • domain assumption A structured review of 52 papers and 10 stakeholder interviews is sufficient to derive a comprehensive reporting schema that addresses all major gaps in AI evaluation reporting.
    This premise is invoked to justify the reporting schema in the abstract.
invented entities (1)
  • EvalCards
    purpose: To act as an operational interpretive reporting layer that unifies evaluation data and renders stakeholder-specific signals.
    This is the primary new construct introduced by the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting." pith.science (2026). https://pith.science/paper/CB25XQ62

@misc{pith2026260609809,
  author       = {Pith},
  title        = {Pith review of: Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CB25XQ62}},
  note         = {Machine review of arXiv:2606.09809}
}
read the original abstract

AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers cannot reliably compare results across sources, identify what a report omits, or trace an aggregate claim to its underlying evidence. Recent efforts address isolated components but leave three gaps: they cover only narrow slices of the evaluation lifecycle and do not compose into a single interpretable record; they specify static representations that do not differentiate the questions different stakeholders bring to the same evidence; and they remain proposals on paper, lacking the extraction infrastructure required for adoption at scale. We present \EvalCards{}, an operational reporting layer that composes benchmark metadata, evaluation run data, and model metadata into a unified record. We (1) derive a reporting schema from a structured review of 52 papers and 10 stakeholder interviews, (2) implement four interpretive signals (reproducibility, documentation completeness, provenance and risk, and score comparability), rendered through reader modes calibrated to research and non-research audiences, and (3) deploy a monitoring tool that applies \EvalCards{} across 5,816 models, 635 benchmarks, and 101,843 results, surfacing systematic gaps in current reporting practice.

Figures

Figures reproduced from arXiv: 2606.09809 by the authors.

Figure 1
Figure 1. Backend canonicalization pipeline. Four sources feed four stage groups. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. The five-level rollout hierarchy. Every reported score resolves to a full path through these [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. An example EVALUATION CARD view, showing the summary view with the main information about a benchmark in plain language. More UI views are shown in Section A and Section K Summary mode foregrounds accountability and plain-language interpretation. All policy interviewees discussed the need for policy stakeholders to have clear takeaways from evaluation results, as policy stakeholders have limited time to sift through… view at source ↗
Figures from the paper (17 more)
Figure 4
Figure 4. Figure 4: Hierarchy for the composite Artificial Analysis, comprising 15 benchmarks. [PITH_FULL_IMAGE:figures/full_fig_p022_4.png]
Figure 5
Figure 5. Figure 5: Corpus-level view: The four interpretive signals aggregated across 5,816 models and [PITH_FULL_IMAGE:figures/full_fig_p023_5.png]
Figure 6
Figure 6. Figure 6: PRISMA flowchart depicting the flow of information through phases of the systematic [PITH_FULL_IMAGE:figures/full_fig_p051_6.png]
Figure 7
Figure 7. Figure 7: Summary of study characteristics by (a) Number of papers per publications year, (b) paper [PITH_FULL_IMAGE:figures/full_fig_p053_7.png]
Figure 8
Figure 8. Figure 8: Summary of extracted items by (a) their type, (b) their class, (c) the workflow stage, [PITH_FULL_IMAGE:figures/full_fig_p054_8.png]
Figure 9
Figure 9. Figure 9: Frequency of combinations of the Extracted Item (a) Workflow Stages, (b) Artifacts [PITH_FULL_IMAGE:figures/full_fig_p055_9.png]
Figure 10
Figure 10. Figure 10: Identification (left) and Who reports what (right): The [PITH_FULL_IMAGE:figures/full_fig_p084_10.png]
Figure 11
Figure 11. Figure 11: Reported metrics, Overlaps: Cross-source score comparison for GPT-5 benchmarks [PITH_FULL_IMAGE:figures/full_fig_p084_11.png]
Figure 12
Figure 12. Figure 12: Reported metrics, Category: The Category view of the reported metrics compares model [PITH_FULL_IMAGE:figures/full_fig_p085_12.png]
Figure 13
Figure 13. Figure 13: The benchmark card section shows a benchmark’s coverage tags, licensing information, [PITH_FULL_IMAGE:figures/full_fig_p085_13.png]
Figure 14
Figure 14. Figure 14: Leaderboard: (A) Frontier view showing the progression of top scores over time. (B) [PITH_FULL_IMAGE:figures/full_fig_p086_14.png]
Figure 15
Figure 15. Figure 15: Summary mode displays benchmark data in plain language, summarizing the [PITH_FULL_IMAGE:figures/full_fig_p086_15.png]
Figure 16
Figure 16. Figure 16: Interpretive signals panel displaying reproducibility, completeness, provenance, and [PITH_FULL_IMAGE:figures/full_fig_p087_16.png]
Figure 17
Figure 17. Figure 17: The comparability signal shows the threshold basis, number of models compared, and [PITH_FULL_IMAGE:figures/full_fig_p087_17.png]
Figure 18
Figure 18. Figure 18: The Model Developers section displays the list of developers with reported models and [PITH_FULL_IMAGE:figures/full_fig_p088_18.png]
Figure 19
Figure 19. Figure 19: Sankey diagram mapping from input groups, to individual items, to sources ingested by [PITH_FULL_IMAGE:figures/full_fig_p095_19.png]
Figure 20
Figure 20. Figure 20: Traceability from the literature-derived framework to E [PITH_FULL_IMAGE:figures/full_fig_p096_20.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

155 extracted references · 39 canonical work pages

  1. [1]

    Developing and maintaining an open- source repository of AI evaluations: Challenges and insights

    Alexandra Abbas, Celia Waggoner, and Justin Olive. Developing and maintaining an open- source repository of AI evaluations: Challenges and insights. InChampioning Open-source DEvelopment in ML Workshop @ ICML25, 2025

  2. [2]

    Mubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja, Pawan Sasanka Ammanamanchi, Ruchit Rawal, Vilém Zouhar, Srishti Yadav, Chenxi Whitehouse, Dayeon Ki, Jennifer Mickel, Leshem Choshen, Marek Šuppa, Jan Batzner, Jenny Chim, Jeba Sania, Yanan Long, Hossein A. Rahmani, Christina Knight, Yiyang Nan, Jyoutir Raj, Yu Fan, Shubham Singh, Subramanyam Sahoo...

  3. [3]

    Audit and assurance of AI algorithms: A framework to ensure ethical algorithmic practices in artificial intelligence, 2021

    Ramya Akula and Ivan Garibay. Audit and assurance of AI algorithms: A framework to ensure ethical algorithmic practices in artificial intelligence, 2021. URL https://arxiv.org/abs/ 2107.14046

  4. [4]

    Lessons from the trenches on evaluating machine-learning systems in materials science.Computational Materi- als Science, 2025

    Nawaf Alampara, Mara Schilling-Wilhelmi, and Kevin Maik Jablonka. Lessons from the trenches on evaluating machine-learning systems in materials science.Computational Materi- als Science, 2025

  5. [5]

    Salmanpour

    Morteza Alizadeh, Mehrdad Oveisi, Sonya Falahati, Ghazal Mousavi, Mohsen Alambardar Meybodi, Somayeh Sadat Mehrnia, Ilker Hacihaliloglu, Arman Rahmim, and Mohammad R. Salmanpour. AllMetrics: A unified Python library for standardized metric evaluation and robust data validation in machine learning, 2025. URLhttps://arxiv.org/abs/2505.15931

  6. [6]

    Arnstein

    Sherry R. Arnstein. A ladder of citizen participation.Journal of the American Institute of Planners, 1969

  7. [7]

    Comparison of AI models across intelligence, performance, and price,

    Artificial Analysis. Comparison of AI models across intelligence, performance, and price,

  8. [8]

    URLhttps://artificialanalysis.ai/models

Show all 155 references
  1. [9]

    Frank Bagehorn, Kristina Brimijoin, Elizabeth M. Daly, Jessica He, Michael Hind, Luis Garces-Erice, Christopher Giblin, Ioana Giurgiu, Jacquelyn Martino, Rahul Nair, David Piorkowski, Ambrish Rawat, John Richards, Sean Rooney, Dhaval Salwala, Seshu Tirupathi, Peter Urbanetz, K...

  2. [10]

    Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs

    Simone Balloccu, Patrícia Schmidtová, Mateusz Lango, and Ondrej Dusek. Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs. InProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volu...

  3. [11]

    When fairness isn’t statistical: The limits of machine learning in evaluating legal reasoning, 2025

    Claire Barale, Michael Rovatsos, and Nehal Bhuta. When fairness isn’t statistical: The limits of machine learning in evaluating legal reasoning, 2025. URL https://arxiv.org/abs/ 2506.03913

  4. [12]

    Every eval ever: Toward a common language for AI eval reporting

    Jan Batzner*. Every eval ever: Toward a common language for AI eval reporting. https:// evalevalai.com/infrastructure/2026/02/17/everyevalever-launch/, February

  5. [13]

    Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, María Grandury, Simeng Han, Valentin Hofma...

  6. [14]

    Silvia Beddar-Wiesing, Alice Moallemy-Oureh, Marie Kempkes, and Josephine M. Thomas. Absolute evaluation measures for machine learning: A survey, 2025. URL https://arxiv. org/abs/2507.03392

  7. [15]

    Open llm leaderboard (2023- 2024)

    Edward Beeching, Clémentine Fourrier, Nathan Habib, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf. Open llm leaderboard (2023- 2024). https://huggingface.co/spaces/open-llm-leaderboard-old/open_llm_ leaderboard, 2023

  8. [16]

    Lessons from the trenches on reproducible evaluation of language models.arXiv preprint arXiv:2405.14782, 2024

    Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, et al. Lessons from the trenches on reproducible evaluation of language models.arXiv preprint arXiv:2405.14782, 2024

  9. [17]

    A metrological framework for uncertainty evaluation in machine learning classification models.Metrologia, 2025

    Samuel Bilson, Maurice Cox, Anna Pustogvar, and Andrew Thompson. A metrological framework for uncertainty evaluation in machine learning classification models.Metrologia, 2025

  10. [18]

    The impact of standardisation and standards on innovation

    Knut Blind. The impact of standardisation and standards on innovation. InHandbook of innovation policy impact. Edward Elgar Publishing, 2016

  11. [19]

    Assessing ai: Surveying the spectrum of approaches to understanding and auditing ai systems, 2025

    Miranda Bogen, Chinmay Deshpande, Ruchika Joshi, Evani Radiya-Dixit, Amy Winecoff, and Kevin Bankston. Assessing ai: Surveying the spectrum of approaches to understanding and auditing ai systems, 2025. URL https://cdt.org/wp-content/uploads/2025/01/ 2025-01-15-CDT-AI-Gov-Lab-A...

  12. [20]

    Evaluation for change

    Rishi Bommasani. Evaluation for change. InFindings of the Association for Computational Linguistics: ACL 2023, 2023

  13. [21]

    Eval factsheets: A structured framework for documenting ai evaluations, 2025

    Florian Bordes, Candace Ross, Justine T Kao, Evangelia Spiliopoulou, and Adina Williams. Eval factsheets: A structured framework for documenting ai evaluations, 2025. URL https: //arxiv.org/abs/2512.04062

  14. [22]

    Bernice B. Brown. Delphi process: A methodology used for the elicitation of opinions of experts. Technical report, RAND Corporation, Santa Monica, CA, 1968

  15. [23]

    Cohn, and Jose Hernandez- Orallo

    María Victoria Carro, Ryan Burnell, Carlos Mougan, Anka Reuel, Wout Schellaert, Olawale Elijah Salaudeen, Lexin Zhou, Patricia Paskov, Anthony G. Cohn, and Jose Hernandez- Orallo. Prep-eval: A pre-registration and reporting protocol for ai evaluations. Manuscript under review,...

  16. [24]

    best fit

    Christopher Carroll, Andrew Booth, and Katy Cooper. A worked example of “best fit” framework synthesis: A systematic review of views concerning the taking of some potential chemopreventive agents.BMC Medical Research Methodology, 11:29, 2011. doi: 10.1186/ 1471-2288-11-29

  17. [25]

    Black-box access is insufficient for rigorous AI audits, 2024

    Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Lynn Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer, Marius Hobbhahn, Lee Sharkey, Satyapriya Krishna, Marvin V on Hagen, Silas Alberti, Alan Chan, Qinyi Sun, Michael Gerovitch, David Bau...

  18. [26]

    The problem with intelligence: Its value-laden history and the future of AI

    Stephen Cave. The problem with intelligence: Its value-laden history and the future of AI. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, pages 243–249, New York, NY , USA, 2020. Association for Computing Machinery. ISBN 9781450370615. doi: 10.1145/33756...

  19. [27]

    Managing misuse risk for dual-use foundation models

    Center for AI Standards and Innovation. Managing misuse risk for dual-use foundation models. Initial Public Draft NIST AI 800-2 IPD, National Institute of Standards and Technology, January 2026. URL https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.800-2.ipd. pdf

  20. [28]

    Test & evaluation best practices for machine learning-enabled systems, 2023

    Jaganmohan Chandrasekaran, Tyler Cody, Nicola McCarthy, Erin Lanus, and Laura Freeman. Test & evaluation best practices for machine learning-enabled systems, 2023. URL https: //arxiv.org/abs/2310.06800

  21. [29]

    Evaluating machine expertise: How graduate students develop frameworks for assessing GenAI content, 2025

    Celia Chen and Alex Leitch. Evaluating machine expertise: How graduate students develop frameworks for assessing GenAI content, 2025. URL https://arxiv.org/abs/2504. 17964

  22. [30]

    Gonzalez, and Ion Stoica

    Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. Chatbot arena: An open platform for evaluating LLMs by human preference. InForty-first International C...

  23. [31]

    Collins, Karel G

    Gary S. Collins, Karel G. M. Moons, Paula Dhiman, Richard D. Riley, Andrew L. Beam, Ben Van Calster, Marzyeh Ghassemi, Xiaoxuan Liu, Johannes B. Reitsma, Maarten van Smeden, et al. TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regressi...

  24. [32]

    Who audits the auditors? recommendations from a field scan of the algorithmic auditing ecosystem, 2023

    Sasha Costanza-Chock, Emma Harvey, Inioluwa Deborah Raji, Martha Czernuszenko, and Joy Buolamwini. Who audits the auditors? recommendations from a field scan of the algorithmic auditing ecosystem, 2023. URLhttps://arxiv.org/abs/2310.02521

  25. [33]

    Evalcards: A framework for standardized evaluation reporting, 2025

    Ruchira Dhar, Danae Sanchez Villegas, Antonia Karamolegkou, Alice Schiavone, Yifei Yuan, Xinyi Chen, Jiaang Li, Stella Frank, Laura De Grazia, Monorama Swain, et al. Evalcards: A framework for standardized evaluation reporting, 2025

  26. [34]

    Nahab, and Xiao Hu

    Cheng Ding, Zhicheng Guo, Cynthia Rudin, Ran Xiao, Fadi B. Nahab, and Xiao Hu. Reconsid- eration on evaluation of machine learning models in continuous monitoring using wearables,

  27. [35]

    URLhttps://arxiv.org/abs/2312.02300

  28. [36]

    Introducing Epoch AI’s AI benchmarking hub, 2024

    Epoch AI. Introducing Epoch AI’s AI benchmarking hub, 2024. URL https://epoch.ai/ blog/introducing-benchmarks-dashboard

  29. [37]

    Can we trust AI benchmarks? An interdisci- plinary review of current issues in AI evaluation

    Maria Eriksson, Erasmo Purificato, Arman Noroozian, Joao Vinagre, Guillaume Chaslot, Emilia Gomez, and David Fernandez-Llorca. Can we trust AI benchmarks? An interdisci- plinary review of current issues in AI evaluation. InAIES, 2025

  30. [38]

    The general-purpose AI code of practice: Safety & security chapter, July 2025

    European Commission, DG CONNECT. The general-purpose AI code of practice: Safety & security chapter, July 2025. URL https://digital-strategy.ec.europa.eu/en/ policies/ai-code-practice. European Commission policy webpage, published July 10, 2025

  31. [39]

    The general-purpose AI code of practice: Trans- parency chapter, July 2025

    European Commission, DG CONNECT. The general-purpose AI code of practice: Trans- parency chapter, July 2025. URL https://digital-strategy.ec.europa.eu/en/ policies/ai-code-practice. European Commission policy webpage, published July 10, 2025

  32. [40]

    EvalEval: Every eval ever shared task, 2024

    EvalEval Coalition. EvalEval: Every eval ever shared task, 2024. URLhttps://evalevalai. com/events/shared-task-every-eval-ever/

  33. [41]

    Good practices for evaluation of machine learning systems, 2024

    Luciana Ferrer, Odette Scharenborg, and Tom Bäckström. Good practices for evaluation of machine learning systems, 2024. URLhttps://arxiv.org/abs/2412.03700

  34. [42]

    Frontier capability assessment

    Frontier Model Forum. Frontier capability assessment. Technical report, Frontier Model Forum, April 2025

  35. [43]

    Datasheets for datasets.Commun

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. Datasheets for datasets.Commun. ACM, 2021. 13

  36. [44]

    Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text.Journal of Artificial Intelligence Research, 2023

    Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text.Journal of Artificial Intelligence Research, 2023

  37. [45]

    Shaona Ghosh, Heather Frase, Adina Williams, Sarah Luger, Paul Röttger, Fazl Barez, Sean McGregor, Kenneth Fricklas, Mala Kumar, Quentin Feuillade-Montixi, Kurt Bollacker, Felix Friedrich, Ryan Tsang, Bertie Vidgen, Alicia Parrish, Chris Knotz, Eleonora Presani, Jonathan Benni...

  38. [46]

    Stress-testing capability elicitation with password-locked models, 2024

    Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, and David Krueger. Stress-testing capability elicitation with password-locked models, 2024. URL https://arxiv.org/abs/ 2405.19550

  39. [47]

    Olmes: A standard for language model evaluations

    Yuling Gu, Oyvind Tafjord, Bailey Kuehl, Dany Haddad, Jesse Dodge, and Hannaneh Ha- jishirzi. Olmes: A standard for language model evaluations. InFindings of the Association for Computational Linguistics: NAACL 2025, pages 5005–5033, 2025

  40. [48]

    Gupta, Jessica Hullman, and Hari Subramonyam

    Neha R. Gupta, Jessica Hullman, and Hari Subramonyam. A conceptual framework for ethical evaluation of machine learning systems. InProceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society, 2025

  41. [49]

    Kakadiaris

    Furkan Gursoy and Ioannis A. Kakadiaris. System cards for AI-based decision-making for public policy, 2022. URLhttps://arxiv.org/abs/2203.04754

  42. [50]

    Empirical privacy evaluations of generative and predictive machine learning models – a review and challenges for practice, 2024

    Flavio Hafner and Chang Sun. Empirical privacy evaluations of generative and predictive machine learning models – a review and challenges for practice, 2024. URL https://arxiv. org/abs/2411.12451

  43. [51]

    Bernstein, and Mykel John Kochenderfer

    Amelia Hardy, Anka Reuel, Kiana Jafari Meimandi, Lisa Soder, Allie Griffith, Dylan M Asmar, Sanmi Koyejo, Michael S. Bernstein, and Mykel John Kochenderfer. More than marketing? on the information value of ai benchmarks for practitioners. InProceedings of the 30th Internationa...

  44. [52]

    A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility

    Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao, Samuel Albanie, Ameya Prabhu, and Matthias Bethge. A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility. InSecond Conference on Language Modeling, 2025

  45. [53]

    Auto-benchmarkcard: Automated synthesis of benchmark documentation

    Aris Hofmann, Inge Vejsbjerg, Dhaval Salwala, and Elizabeth Daly. Auto-benchmarkcard: Automated synthesis of benchmark documentation. InProceedings of the 2026 AAAI Conference on Artificial Intelligence, volume 40(48), pages 41598–41600, 2026. doi: 10.1609/aaai.v40i48.42352

  46. [54]

    Values in the wild: Discovering and analyzing values in real-world language model interactions, 2025

    Saffron Huang, Esin Durmus, Miles McCain, Kunal Handa, Alex Tamkin, Jerry Hong, Michael Stern, Arushi Somani, Xiuruo Zhang, and Deep Ganguli. Values in the wild: Discovering and analyzing values in real-world language model interactions, 2025. URL https://arxiv. org/abs/2504.15236. 14

  47. [55]

    Evaluation gaps in machine learning practice

    Ben Hutchinson, Negar Rostamzadeh, Christina Greer, Katherine Heller, and Vinodkumar Prabhakaran. Evaluation gaps in machine learning practice. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 2022

  48. [56]

    Rethinking machine learning model evaluation in pathology, 2022

    Syed Ashar Javed, Dinkar Juyal, Zahil Shanis, Shreya Chakraborty, Harsha Pokkalla, and Aaditya Prakash. Rethinking machine learning model evaluation in pathology, 2022. URL https://arxiv.org/abs/2204.05205

  49. [57]

    Deprecating benchmarks: Criteria and framework

    Ayrton San Joaquin, Rokas Gipiškis, Leon Staufer, and Ariel Gil. Deprecating benchmarks: Criteria and framework. InICML Workshop on Technical AI Governance (TAIG), 2025

  50. [58]

    Cantrell, Keiran Peng, Thanh Huy Pham, Christopher A

    Sayash Kapoor, Ethan M. Cantrell, Keiran Peng, Thanh Huy Pham, Christopher A. Bail, Odd Erik Gundersen, Jake M. Hofman, Jessica Hullman, Michael A. Lones, Meenal M. Malik, Priyanka Nanayakkara, Russell A. Poldrack, Inioluwa Deborah Raji, Mike Roberts, Matthew J. Salganik, Mart...

  51. [59]

    Benchmark profiling: Mechanistic diagnosis of LLM benchmarks, 2025

    Dongjun Kim, Gyuho Shim, Yongchan Chun, Minhyuk Kim, Chanjun Park, and Heuiseok Lim. Benchmark profiling: Mechanistic diagnosis of LLM benchmarks, 2025. URL https: //arxiv.org/abs/2510.01232

  52. [60]

    Had- field, Lukas Heim, Marianela Rodriguez, Jonas B

    Noam Kolt, Markus Anderljung, Jess Barnhart, Imogen Brass, Kevin Esvelt, Gillian K. Had- field, Lukas Heim, Marianela Rodriguez, Jonas B. Sandbrink, and Tom Woodside. Responsible reporting for frontier AI development. InProceedings of the AAAI/ACM Conference on AI, Ethics, and...

  53. [61]

    Richard Landis and Gary G

    J. Richard Landis and Gary G. Koch. The measurement of observer agreement for categorical data.Biometrics, 1977

  54. [62]

    Towards explainable evaluation metrics for machine translation.Journal of Machine Learning Research, 2024

    Christoph Leiter, Piyawat Lertvittayakumjorn, Marina Fomicheva, Wei Zhao, Yang Gao, and Steffen Eger. Towards explainable evaluation metrics for machine translation.Journal of Machine Learning Research, 2024

  55. [63]

    Frangi, Antonio R

    Karim Lekadir, Alejandro F. Frangi, Antonio R. Porras, Ben Glocker, et al. FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare.BMJ, 2025

  56. [64]

    Manning, Christopher Ré, Diana Acosta-Navas, Drew A

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas...

  57. [65]

    Manning, Christopher Ré, Diana Acosta-Navas, Drew A

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas...

  58. [66]

    Are we learning yet? a meta review of evaluation failures across machine learning

    Thomas Liao, Rohan Taori, Inioluwa Deborah Raji, and Ludwig Schmidt. Are we learning yet? a meta review of evaluation failures across machine learning. InNeurIPS 2021 Datasets and Benchmarks Track, 2021

  59. [67]

    A safe harbor for AI evaluation and red teaming

    Shayne Longpre, Sayash Kapoor, Kevin Klyman, Aviya Ramaswami, Rishi Bommasani, Borhane Blili-Hamelin, Yangsibo Huang, Aleksander Skowron, Zheng-Xin Yong, Suhas Kotha, Yi Zeng, Weiyan Shi, Xianjun Yang, Reid Southen, Alexander Robey, Patrick Chao, Diyi Yang, Robin Jia, Daniel K...

  60. [68]

    LLM cyber evaluations don’t capture real-world risk,

    Kamil˙e Lukoši¯ut˙e and Adam Swanda. LLM cyber evaluations don’t capture real-world risk,

  61. [69]

    URLhttps://arxiv.org/abs/2502.00072

  62. [70]

    Data contamination: From memorization to exploitation

    Inbal Magar and Roy Schwartz. Data contamination: From memorization to exploitation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2022

  63. [71]

    Building less-flawed metrics: Understanding and creating better measurement and incentive systems.Patterns, 2023

    David Manheim. Building less-flawed metrics: Understanding and creating better measurement and incentive systems.Patterns, 2023

  64. [72]

    Oard, Luca Soldaini, Ian Soboroff, Orion Weller, Efsun Kayi, Kate Sanders, Marc Mason, and Noah Hibbler

    James Mayfield, Eugene Yang, Dawn Lawrie, Sean MacAvaney, Paul McNamee, Douglas W. Oard, Luca Soldaini, Ian Soboroff, Orion Weller, Efsun Kayi, Kate Sanders, Marc Mason, and Noah Hibbler. On the evaluation of machine-generated reports. InProceedings of the 47th International A...

  65. [73]

    STREAM (ChemBio): A standard for transparently reporting evaluations in AI model reports, 2025

    Tegan McCaslin, Jide Alaga, Samira Nedungadi, Seth Donoughe, Tom Reed, Rishi Bommasani, Chris Painter, and Luca Righetti. STREAM (ChemBio): A standard for transparently reporting evaluations in AI model reports, 2025. URLhttps://arxiv.org/abs/2508.09853

  66. [74]

    Adding error bars to evals: A statistical approach to language model evaluations,

    Evan Miller. Adding error bars to evals: A statistical approach to language model evaluations,

  67. [75]

    URLhttps://arxiv.org/abs/2411.00640

  68. [76]

    Model cards for model reporting

    Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting. InProceedings of the Conference on Fairness, Accountability, and Transparency. Association for...

  69. [77]

    State of what art? A call for multi-prompt LLM evaluation, 2024

    Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. State of what art? A call for multi-prompt LLM evaluation, 2024. URL https://arxiv. org/abs/2401.00595

  70. [78]

    Extrinsic evaluation of machine translation metrics

    Nikita Moghe, Tom Sherborne, Mark Steedman, and Alexandra Birch. Extrinsic evaluation of machine translation metrics. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023

  71. [79]

    John Mongan, Linda Moy, and Charles E. Kahn Jr. Checklist for artificial intelligence in medical imaging (CLAIM): A guide for authors and reviewers.Radiology: Artificial Intelligence, 2020

  72. [80]

    A survey on large language model benchmarks, 2025

    Shiwen Ni, Guhong Chen, Shuaimin Li, Xuanang Chen, Siyi Li, Bingli Wang, Qiyao Wang, Xingjian Wang, Yifan Zhang, Liyang Fan, Chengming Li, Ruifeng Xu, Le Sun, and Min Yang. A survey on large language model benchmarks, 2025. URL https://arxiv.org/ abs/2508.15361

  73. [81]

    Evaluation of DeepSeek AI models

    NIST Center for AI Standards and Innovation (CAISI). Evaluation of DeepSeek AI models. Technical report, NIST, 2025. URL https://www.nist.gov/system/files/documents/ 2025/09/30/CAISI_Evaluation_of_DeepSeek_AI_Models.pdf

  74. [82]

    Patryk Orzechowski and Jason H. Moore. Generative and reproducible benchmarks for comprehensive evaluation of machine learning classifiers.Science Advances, 2022. 16

  75. [83]

    Byun, Kevin Wei, and Toby Webster

    Patricia Paskov, Michael J. Byun, Kevin Wei, and Toby Webster. Preliminary suggestions for rigorous GPAI model evaluations. Technical Report PE-A3971-1, RAND Corporation, May

  76. [84]

    URLhttps://www.rand.org/pubs/perspectives/PEA3971-1.html

  77. [85]

    Toward best practices for AI evaluation and governance: A proposal for a european union general-purpose AI model evaluation standards task force

    Patricia Paskov, Lisa Soder, and Everett Smith. Toward best practices for AI evaluation and governance: A proposal for a european union general-purpose AI model evaluation standards task force. Technical Report PE-A3624-1, RAND Corporation, June 2025. URL https://www.rand.org/...

  78. [86]

    Data cards: Purposeful and transparent dataset documentation for responsible ai

    Mahima Pushkarna, Andrew Zaldivar, and Oddur Kjartansson. Data cards: Purposeful and transparent dataset documentation for responsible ai. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 2022

  79. [87]

    Pooja S. B. Rao, Sanja Š ´cepanovi´c, Dinesh Babu Jayagopi, Mauro Cherubini, and Daniele Quercia. The ai model risk catalog: What developers and researchers miss about real-world ai harms, 2025. URLhttps://arxiv.org/abs/2508.16672

  80. [88]

    Mellor, Jonathan Uesato, Po-Sen Huang, Johannes Welbl, Laura Weidinger, Sumanth Dathathri, Amelia Glaese, Geoffrey Irving, Iason Gabriel, William Isaac, and Lisa Anne Hendricks

    Maribeth Rauh, John F.J. Mellor, Jonathan Uesato, Po-Sen Huang, Johannes Welbl, Laura Weidinger, Sumanth Dathathri, Amelia Glaese, Geoffrey Irving, Iason Gabriel, William Isaac, and Lisa Anne Hendricks. Characteristics of harmful text: Towards rigorous benchmarking of language...

  81. [89]

    Kim, Stephen Fitz, and Dan Hendrycks

    Richard Ren, Steven Basart, Adam Khoja, Alice Gatti, Long Phan, Xuwang Yin, Mantas Mazeika, Alexander Pan, Gabriel Mukobi, Ryan H. Kim, Stephen Fitz, and Dan Hendrycks. SafetyWashing: Do AI safety benchmarks actually measure safety progress?, 2024. URL https://arxiv.org/abs/2407.21792

  82. [90]

    Kochenderfer

    Anka Reuel, Amelia Hardy, Chandler Smith, Max Lamparth, Malcolm Hardy, and Mykel J. Kochenderfer. Betterbench: Assessing ai benchmarks, uncovering issues, and establishing best practices. InNeurIPS 2024 Track Datasets and Benchmarks Track, 2024

  83. [91]

    Measur- ing what matters: Connecting ai ethics evaluations to system attributes, hazards, and harms

    Shalaleh Rismani, Renee Shelby, Leah Davis, Negar Rostamzadeh, and AJung Moon. Measur- ing what matters: Connecting ai ethics evaluations to system attributes, hazards, and harms. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 2025

  84. [92]

    Measurement to meaning: A validity-centered framework for AI evaluation, 2025

    Olawale Salaudeen, Anka Reuel, Ahmed Ahmed, Suhana Bedi, Zachary Robertson, Sudharsan Sundar, Ben Domingue, Angelina Wang, and Sanmi Koyejo. Measurement to meaning: A validity-centered framework for AI evaluation, 2025. URL https://arxiv.org/abs/2505. 10573

  85. [93]

    Malik Sallam, Muna Barakat, and Maram Sallam. A preliminary checklist (METRICS) to standardize the design and reporting of studies on generative artificial intelligence-based models in health care education and practice: Development study involving a literature review. Interac...

  86. [94]

    A. G. R. Sandeepa and Sanka Mohottala. Evaluation of machine learning models in student academic performance prediction. In2025 5th International Conference on Advanced Research in Computing (ICARC), 2025

  87. [95]

    Reality check: A new evaluation ecosystem is necessary to understand ai’s real world effects, 2025

    Reva Schwartz, Rumman Chowdhury, Akash Kundu, Heather Frase, Marzieh Fadaee, Tom David, Gabriella Waters, Afaf Taik, Morgan Briggs, Patrick Hall, Shomik Jain, Kyra Yee, Spencer Thomas, Sundeep Bhandari, Paul Duncan, Andrew Thompson, Maya Carlyle, Qinghua Lu, Matthew Holmes, an...

  88. [96]

    Improving methodologies for agentic evaluations across domains: Leakage of sensitive information, fraud and cybersecurity threats, 2026

    Ee Wei Seah, Yongsen Zheng, Naga Nikshith, Mahran Morsidi, Gabriel Waikin Loh Matienzo, Nigel Gay, Akriti Vij, Benjamin Chua, En Qi Ng, Sharmini Johnson, Vanessa Wilfred, Wan Sie Lee, Anna Davidson, Catherine Devine, Erin Zorer, Gareth Holvey, Harry Coppock, James Walpole, Jer...

  89. [97]

    Model evaluation for extreme risks, 2023

    Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, Lewis Ho, Divya Siddarth, Shahar Avin, Will Hawkins, Been Kim, Iason Gabriel, Vijay Bolina, Jack Clark, Yoshua Bengio, ...

  90. [98]

    Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker

    Shivalika Singh, Yiyang Nan, Alex Wang, Daniel D’souza, Sayash Kapoor, Ahmet Üstün, Sanmi Koyejo, Yuntian Deng, Shayne Longpre, Noah A. Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker. The leaderboard illusion. InThe Thirty-ninth Annual Conference on Neural Information Pro...

  91. [99]

    Daly, Michael Hind, David Piorkowski, Xiangliang Zhang, Nuno Moniz, and Nitesh V Chawla

    Anna Sokol, Elizabeth M. Daly, Michael Hind, David Piorkowski, Xiangliang Zhang, Nuno Moniz, and Nitesh V Chawla. Benchmarkcards: Standardized documentation for large language model benchmarks. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems Datas...

  92. [100]

    Verifiable evaluations of machine learning models using zkSNARKs, 2024

    Tobin South, Alexander Camuto, Shrey Jain, Shayla Nguyen, Robert Mahari, Christian Paquin, Jason Morton, and Alex ’Sandy’ Pentland. Verifiable evaluations of machine learning models using zkSNARKs, 2024. URLhttps://arxiv.org/abs/2402.02675

  93. [101]

    Audit cards: Contextualizing AI evaluations, 2025

    Leon Staufer, Mick Yang, Anka Reuel, and Stephen Casper. Audit cards: Contextualizing AI evaluations, 2025. URLhttps://arxiv.org/abs/2504.13839

  94. [102]

    Artificial intelligence risk management framework (ai rmf 1.0)

    Elham Tabassi. Artificial intelligence risk management framework (ai rmf 1.0). NIST AI 100-1, National Institute of Standards and Technology, January 2023. URL https: //doi.org/10.6028/NIST.AI.100-1

  95. [103]

    Evaluation of machine learning algorithms for health and wellness applications: A tutorial.Computers in Biology and Medicine, 2021

    Jussi Tohka and Mark van Gils. Evaluation of machine learning algorithms for health and wellness applications: A tutorial.Computers in Biology and Medicine, 2021

  96. [104]

    Long-form tasks: A methodology for evaluating advanced AI systems, December 2024

    UK AI Safety Institute. Long-form tasks: A methodology for evaluating advanced AI systems, December 2024. URL https://www.aisi.gov.uk/work/long-form-tasks. AISI Work (report)

  97. [105]

    A structured protocol for elicitation experiments: Calibrating AI risk assessment through rigorous elicitation practices, July 2025

    UK AI Safety Institute. A structured protocol for elicitation experiments: Calibrating AI risk assessment through rigorous elicitation practices, July 2025. URL https://www.aisi.gov. uk/work/our-approach-to-ai-capability-elicitation. AISI Work (blog post)

  98. [106]

    Emerging processes for frontier AI safety

    UK Department for Science, Innovation and Technology. Emerging processes for frontier AI safety. Technical report, UK Government, October 2023. Policy paper. Open Government Licence v3.0

  99. [107]

    Ukgovernmentbeis/inspect_ai: Inspect: A framework for large lan- guage model evaluations, May 2026

    UKGovernmentBEIS. Ukgovernmentbeis/inspect_ai: Inspect: A framework for large lan- guage model evaluations, May 2026. URL https://github.com/UKGovernmentBEIS/ inspect_ai

  100. [108]

    Ullrich, Elizabeth A

    Paul A. Ullrich, Elizabeth A. Barnes, William D. Collins, Kate Dagon, Siyuan Duan, Jacqueline Elms, Jiwoo Lee, L. Ruby Leung, Dan Lu, Michael J. Molina, Travis A. O’Brien, and Forrest O. Rebassoo. Recommendations for comprehensive and independent evaluation of machine learning...

  101. [109]

    Feder Cooper, Angelina Wang, Solon Barocas, Alexandra Chouldechova, Chad Atalla, Su Lin Blodgett, Emily Corvi, P

    Hanna Wallach, Meera Desai, Nicholas Pangakis, A. Feder Cooper, Angelina Wang, Solon Barocas, Alexandra Chouldechova, Chad Atalla, Su Lin Blodgett, Emily Corvi, P. Alex Dow, Jean Garcia-Gathright, Alexandra Olteanu, Stefanie Reed, Emily Sheng, Dan Vann, Jennifer Wortman Vaugha...

  102. [110]

    Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290, 2024

    Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, et al. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95...

  103. [111]

    Wei, Patricia Paskov, Sunishchal Dev, Michael J

    Kevin L. Wei, Patricia Paskov, Sunishchal Dev, Michael J. Byun, Anka Reuel, Xavier Roberts- Gaal, Rachel Calcott, Evie Coxon, and Chinmay Deshpande. Recommendations and reporting checklist for rigorous & transparent human baselines in model evaluations, 2025. URL https://arxiv...

  104. [112]

    Toward an evaluation science for generative AI systems, 2025

    Laura Weidinger, Inioluwa Deborah Raji, Hanna Wallach, Margaret Mitchell, Angelina Wang, Olawale Salaudeen, Rishi Bommasani, Deep Ganguli, Sanmi Koyejo, and William Isaac. Toward an evaluation science for generative AI systems, 2025. URL https://arxiv.org/ abs/2503.05336

  105. [113]

    Why and how governments should monitor AI development,

    Jess Whittlestone and Jack Clark. Why and how governments should monitor AI development,

  106. [114]

    URLhttps://arxiv.org/abs/2108.12427

  107. [115]

    Chrysoula Zerva and André F. T. Martins. Conformalizing machine translation evaluation. Transactions of the Association for Computational Linguistics, 2024

  108. [116]

    SPHERE: An evaluation card for human-AI systems

    Dora Zhao, Qianou Ma, Xinran Zhao, Chenglei Si, Chenyang Yang, Ryan Louie, Ehud Reiter, Diyi Yang, and Tongshuang Wu. SPHERE: An evaluation card for human-AI systems. In Findings of the Association for Computational Linguistics: ACL 2025, pages 1340–1365. Association for Compu...

  109. [117]

    Collins, Yael Moros-Daval, Seraphina Zhang, Qinlin Zhao, Yitian Huang, Luning Sun, Jonathan E

    Lexin Zhou, Lorenzo Pacchiardi, Fernando Martínez-Plumed, Katherine M. Collins, Yael Moros-Daval, Seraphina Zhang, Qinlin Zhao, Yitian Huang, Luning Sun, Jonathan E. Prunty, Zongqian Li, Pablo Sánchez-García, Kexin Jiang-Chen, Pablo A. M. Casares, Jiyun Zu, John Burden, Behzad...

  110. [118]

    Introduce EVALUATIONCARDSas a concept and the purpose of the interview

  111. [119]

    Can you briefly describe your role and whether/how model evaluation fits into that role? Goals, Needs & Decision-Making

  112. [120]

    When you look at evaluation results, what are you trying to decide? 26 (a) How often do you need to do so?

  113. [121]

    What information would you need from an evaluation to make that decision? (a) How would you rank those dimensions relative to each other? Can you show us an example of an evaluation that served you? (b) What is the minimal level of metadata you need? (c) What information is ty...

  114. [122]

    Are evaluation results important for your stakeholders in any way? Who uses this information downstream? How so?

  115. [123]

    OR tell us about the last time you had to do this (if they have trouble remembering)

    Do you ever need to evaluate the evaluation itself? (a) If so, how do you do that currently? Do you find that method trustworthy or easy to use? (b) Tell us about the resources you use, time constraints, etc. . . OR tell us about the last time you had to do this (if they have ...

  116. [124]

    If you had a magic wand to create the tool yourself, what questions would you want it to answer? How might you design it? What would success look like?

  117. [125]

    What information is most important for you to see / be accessible to you when looking at evaluation results across models Pain Points

  118. [126]

    What’s most frustrating about evaluating models today?

  119. [127]

    When comparing models, what is the hardest element to deal with? ie, lack of standardized evaluation formats, metric definition inconsistencies, or lack of reproducibility?

  120. [128]

    We’re testing the interface, not you. Please think out loud as you explore. I may ask what you’re thinking

    Anything else you want us to know about model evaluations and how they relate to your role? Should we have asked you something that we didn’t? Usability Testing. “We’re testing the interface, not you. Please think out loud as you explore. I may ask what you’re thinking.”

  121. [129]

    Do you think as is the tool meets your needs? Rank 1-10

  122. [130]

    How does this compare to how you currently review evaluations?

  123. [131]

    What would be the most useful feature and why?

  124. [132]

    What was the least useful feature and why?

  125. [133]

    Was there anything that felt unclear or ambiguous?

  126. [134]

    How would you change the tool overall to make it more useful for you?

  127. [135]

    <metric> on <benchmark>

    Anything else you want to share with us about the tool? C.3 Limitations Ten interviews were conducted, and interviewees primarily reside in North America. Although interviews were conducted with interviewees residing in both Asia and Europe, perspectives and insights of techni...

  128. [136]

    Normalized match: both sides are reduced to a canonical surface form (lowercased, separator- stripped) before comparison

  129. [137]

    For models, this list collapses upload-format and quantization variants of the same underlying model

    Fuzzy stem match: a short, deliberately narrow list of known suffixes is stripped before comparison. For models, this list collapses upload-format and quantization variants of the same underlying model. The list is intentionally narrow to avoid false merges between genuinely d...

  130. [138]

    s c h e m a _ v e r s i o n

    No match: the raw string flows through unresolved and is surfaced as a candidate for human review. To validate performance, we focused on (models, benchmarks, metrics) and sampled 200 entities per type uniformly at random from the EEE corpus, manually labeling each prediction ...

  131. [139]

    Methodology Transparency

  132. [140]

    2.Comparability:Flags when scores change due to setup differences

    Granularity Mapping to EVALUATIONCARDS: 1.Reproducibility:Surfaces missing generation configs that block independent re-executions. 2.Comparability:Flags when scores change due to setup differences

  133. [141]

    Can I trust this claim for decision-making, and what are the caveats?

    Expands metric configuration details and highlights specific missing fields. Persona 2: Policy Actor Primary Mode:Policy Profile:A regulator, standards body member, or safety institute staffer who consults evaluation evidence to inform governance decisions, risk assessments, o...

  134. [142]

    third-party) 2.Risk Content:What risks does this benchmark actually measure?

    Interpretability of metrics / warnings about limitations Mapping to EVALUATIONCARDS: 1.Accountability:Who reported this score? (first-party vs. third-party) 2.Risk Content:What risks does this benchmark actually measure?

  135. [143]

    Did I document everything required by the current consensus framework?

    Interpretability:Plain-language summaries of what a metric means and warnings about limitations. Persona 3: Model Developer Primary Mode:Research (Self-Audit) Profile:An engineer or scientist at a lab preparing a release (e.g., a model card or technical report). They use EVALU...

  136. [144]

    Standardization (refer to EEE) Mapping to EVALUATIONCARDS:

  137. [145]

    A low completeness score indicates the developer has omitted opera- tionalizable fields identified in the framework

    Serves as a direct checklist. A low completeness score indicates the developer has omitted opera- tionalizable fields identified in the framework

  138. [146]

    Android”, “agent

    Helps the developer structure their reporting correctly, e.g., ensuring suite-level aggregates match the underlying benchmark scores. 35 G Governance EVALUATIONCARDSis maintained as a volunteer-run open-source project. The governance model below is designed to be lightweight a...

  139. [147]

    [32] proposed to document individual evaluations or evaluation instances and McCaslin et al

    and Dhar et al. [32] proposed to document individual evaluations or evaluation instances and McCaslin et al. [70] documents chemical and biological evaluations in model reports; Joaquin et al

  140. [148]

    Intelligence Index

    provides a framework for deprecation and Carro et al. [22] provide a preregistration protocol. Each documents only a slice of the evaluation pipeline, so stakeholders interpreting a single result must piece together results and context from benchmark cards, audit cards, leader...

  141. [149]

    • Normative disagreement is not an error.Items likely to provoke value disagreement should still be coded normally, not excluded

    Decision Rules and Precedence • Stated intent in the paper beats surface wording.A reporting requirement motivated by validity concerns is still Reporting. • Normative disagreement is not an error.Items likely to provoke value disagreement should still be coded normally, not excluded

  142. [150]

    Did the developer publish a verifiable preregistration for the evaluation?

    Worked Examples This section provides concrete examples of how extracted items are coded across dimensions. Ex- amples are drawn directly from items in the current dataset and v2 Delphi list, and are intended to illustrate typical—not idealized—classification decisions. Exampl...

  143. [151]

    • Disputes are resolved by lead reviewer judgment, as logged in notes

    Reviewer Workflow and Disagreement Handling • Reviewers should flag items when the current codings seem incorrect or should be changed, or if a multi-select item should be added, not if more than one possible coding seems reasonable. • Disputes are resolved by lead reviewer ju...

  144. [152]

    doc- umentation

    Versioning and Change Log Version Date Change Rationale 0.1 Dec 12 Initial draft GPT 5.2 draft, per other doc- umentation and dataset. 69 Version Date Change Rationale 1.0 Dec 22 Initial version for use David reviewed and cor- rected, regenerated some sec- tions, and made subs...

  145. [153]

    Design Goals & context

  146. [154]

    Reporting & publication AutoBenchmarkCard EEE aggregate eval EEE instance-level eval Figure 19: Sankey diagram mapping from input groups, to individual items, to sources ingested by EVALUATIONCARDS. 95

  147. [155]

    N Code and Demo Our code is available at https://huggingface.co/spaces/evaleval/general-eval-card/ tree/main, and the live demo athttps://evalcards.evalevalai.com

    Reporting & publication Goals, construct validity, task types Protocol, splits, baselines, contamination Run logs, mitigations, generation config Data access, later use, maintenance Transparency, replication, publication details Comparability Completeness Reproducibility Prove...

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.