Pith. sign in

REVIEW 3 major objections 5 minor 13 cited by

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

T0 review · 3 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read Quantitative AI benchmarks, as currently designed and used, cannot be trusted to single-handedly provide the capability and safety assurances that policymakers need.

desk verdict A useful, policy-facing synthesis of benchmark critique that is honest about its own method, but its 'fundamental fragilities' conclusion is stronger than the non-systematic sample can fully support. read the letter →

arxiv 2502.06559 v2 pith:CLRQDII6 submitted 2025-02-10 cs.AI

classification cs.AI
keywords AIbenchmarksbenchmarkcritiqueevaluationsafetyregulationdatacontaminationsandbaggingconstructvalidity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a meta-review of about 100 studies published between 2014 and 2024 that criticize quantitative AI benchmarks. It argues that benchmark scores are widely treated as reliable measures of AI capability and safety, but that they systematically overstate their own reliability. The review organizes this criticism into a taxonomy of nine issues spanning dataset construction, construct validity, cultural context, commercial incentives, gaming, weak community vetting, saturation, and unknown unknowns. Its central conclusion is that quantitative benchmarking, as currently practiced, is not equipped to be the primary basis for the safety and capability assurances policymakers are asking for. A sympathetic reader should care because benchmarks are already written into regulation, so whether they can carry that weight has immediate practical consequences.

What carries the argument

The central object is the nine-issue taxonomy: problems with data collection, annotation, and documentation; weak construct validity; sociocultural context and gap; narrow diversity and scope; economic, competitive, and commercial roots; rigging, gaming, and measure-becoming-target; dubious community vetting and path dependencies; rapid AI development and benchmark saturation; and AI complexity and unknown unknowns. This taxonomy does the argumentative work by treating the issues as deeply interlinked, so that no single fix to individual benchmarks can restore trust by itself. It also supplies a shared vocabulary for policy discussions, showing that benchmark failure is not one bug but a class of failure modes, each of which can be identified in specific published studies.

What would settle it

One concrete test would be to run a preregistered, systematic literature search of the 2014-2024 period using database keyword queries instead of citation snowballing. If that search surfaces a substantial body of studies in which benchmark scores are shown to predict real-world deployment outcomes reliably, the paper's claim of fundamental fragility would be weakened. A second test is prospective: on a cohort of newly released models, compare benchmark scores against independent red-team findings; if the scores reliably flag the same models as unsafe, the taxonomy's warning would be overstated.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is not a single failed benchmark but a pattern: the same fragilities recur across text, image, audio, and multimodal evaluation, and they are structural rather than incidental. Each issue in the taxonomy shows how a benchmark's number can come apart from what it claims to report — a dataset can encode hidden spurious cues, a safety score can track general capability instead of safety, a model can be trained to underperform on purpose, and a benchmark can become a standard through citation luck rather than through demonstrated validity. Taken together, the paper concludes, these issues point toward fundamental fragilities in current efforts to quantitatively measure and mitigate harm in AI. The specific policy claim is that quantitative benchmarking is currently ill-suited to provide, on its own or even primarily, the safety and capability assurances requested by policymakers, and that what is needed is standardized methods for assessing the trustworthiness of benchmarks rather than standardized benchmarks themselves.

Load-bearing premise

The paper assumes that following citations out from one widely shared critique of benchmarks is enough to find all the important problems; if people who rely on or defend benchmarks published their evidence outside that citation trail, the list of nine issues could be incomplete.

Editorial extensions

If this is right

  • Regulators should not treat benchmark scores as sufficient evidence for classifying a model as high-risk or systemically risky under the EU AI Act.
  • Benchmark results should be published with full documentation, reproduction scripts, multiple evaluation runs, and reported statistical significance, since few current benchmarks provide them.
  • Safety benchmark scores that correlate with general capability should not be marketed as safety progress; separate validation of the safety construct is required.
  • Static, publicly known benchmarks will keep losing value through data contamination and sandbagging, so dynamic or hidden test sets and human-in-the-loop methods need to become part of the default toolkit.
  • Trustworthy benchmarks require a standard way to evaluate the benchmarks themselves, not necessarily standardized benchmark metrics.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • We infer that the nine-issue taxonomy could be turned into a meta-evaluation scorecard: a checklist benchmark users apply before relying on a score, yielding an aggregate trustworthiness rating.
  • An implicit testable extension is to measure how much a benchmark's score moves when each of the nine issues is mitigated; if removing one issue shifts results far more than others, the fragilities do not all carry equal weight.
  • We infer that as laws hardwire benchmarks into regulatory thresholds, the incentive to game, sandbag, or contaminate those exact benchmarks will grow, so a benchmark's regulatory value is partly a function of how hard it is to optimize adversarially.
  • The review focuses on quantitative benchmarks, but its own logic suggests that qualitative methods such as red-teaming and bug-bounty programs should be treated as complementary checks whose trustworthiness is assessed with the same scrutiny.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper is an interdisciplinary meta-review of roughly 110 publications, published between 2014 and 2024, that critique quantitative AI benchmarking. The authors organize the critique into a nine-issue taxonomy covering data collection and documentation, construct validity, sociocultural context, diversity and scope, economic and competitive pressures, gaming and rigging, community vetting, benchmark saturation, and AI complexity/unknown unknowns. They argue that these issues are interlinked and that, taken together, they reveal fundamental fragilities in current efforts to quantitatively measure and mitigate AI harm. The paper concludes that quantitative benchmarking is currently ill-suited to single-handedly provide the safety and capability assurances that policy makers require, and it offers a set of policy-oriented recommendations for improving benchmark trustworthiness.

Significance. If the central claim is accepted, the paper provides a valuable and timely synthesis of a rapidly growing body of critique, with direct relevance to AI regulation and the EU AI Act. Its strengths include a genuinely interdisciplinary scope, inclusion of very recent work (more than half of the reviewed sources are from 2023 or later), and an explicit acknowledgment of the taxonomy's limitations as a 'narrative tool.' The paper also makes concrete, actionable recommendations, such as standardizing methods for assessing benchmark trustworthiness rather than standardizing benchmarks themselves. However, because the review is built on a non-systematic snowball sample and deliberately excludes benchmark-proposal papers, the strength of the general conclusion about 'fundamental fragilities' exceeds what the methodology can strictly support. The paper is useful as a high-quality qualitative synthesis, but its central claim needs to be conditioned on the nature of the sampled literature.

major comments (3)
  1. [Section 4 and Section 6] The snowball sampling procedure described in Section 4 starts from a single critique paper (Raji et al., 2021) and expands through its reference list and forward citations. This procedure is likely to over-represent works that engage with that specific critique and under-represent independent lines of benchmark critique or defense. The authors do not provide the core list of about 110 sources or a coding protocol, making it impossible to assess the completeness or bias of the corpus. Given that the conclusion in Section 6 states that 'taken together, these issues point toward fundamental fragilities in current efforts to quantitatively measure and mitigate harm in AI' and that quantitative benchmarking is 'ill-suited to single-handedly... provide the safety and capability assurances requested by policy makers,' the inference from this sample to the field as a whole is not established at the level of representativeness it asserts. I recommend that the authors provide the full corpus as supplementary material, describe the classification procedure in enough detail to be reproducible, and either soften the conclusion to explicitly refer to the reviewed literature or justify the representativeness of the sample.
  2. [Section 4, exclusion criteria] The authors exclude papers that propose new benchmarks, 'even though such articles naturally contain some level of benchmark critique.' This exclusion removes a substantial portion of the literature that might contain counterevidence or alternative perspectives on whether benchmarks are fundamentally fragile or whether the problems are corrigible. While the authors give pragmatic reasons for the exclusion, the central claim is about current benchmarking practices as a whole, not merely about the subset of papers whose primary purpose is critique. The authors should at least discuss whether the excluded literature could alter the conclusions, or restrict the conclusion to the population of critique-oriented publications.
  3. [Section 4, taxonomy development] The nine issue categories were identified after close reading and internal discussion, without a reported coding protocol or reliability checks. The authors acknowledge this and call the taxonomy 'a narrative tool,' but the conclusion treats the presence of these nine issues as evidence of 'fundamental fragilities.' Since the categories are not shown to be exhaustive or mutually exclusive, the conclusion should be presented as a synthesis of the reviewed critical literature rather than as a comprehensive map of all benchmarking problems. I recommend adding a limitations paragraph that explicitly states that the taxonomy has not been validated as a complete enumeration and that the conclusions reflect the reviewed sources.
minor comments (5)
  1. [Abstract] The sentence 'so too does concerns about how and with what effects...' should read 'so too do concerns...' (subject-verb agreement).
  2. [Section 2] 'heterogenous' should be 'heterogeneous'.
  3. [Section 6] In the sentence 'we especially identify a need for new ways of signallingwhat benchmarks to trust,' there is a missing space between 'signalling' and 'what.'
  4. [Section 5.6 and References] The in-text citation 'Weij et al. [2024]' refers to the reference 'van der Weij, Teun, Felix Hofstätter, et al. AI Sandbagging: Language Models can Strategically Underperform on Evaluations.' For consistency, the in-text citation should be 'van der Weij et al.'
  5. [Section 4] The authors refer to 'approximately 110' sources but never provide a table or appendix listing them. A supplementary list would help readers verify the coverage and would strengthen the reproducibility of the review.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the paper is a narrative meta-review whose synthesis is not derived from its inputs by construction; the only self-citations are minor and not load-bearing.

full rationale

The paper's derivation chain is a literature-based meta-review: snowball sampling (Section 4), a nine-issue taxonomy (Section 5), and a concluding synthesis (Section 6). There are no mathematical derivations, fitted parameters, or empirical predictions that could reduce to input data by construction. The central claim that benchmarks exhibit 'fundamental fragilities' is presented as a synthesis of the surveyed critique literature, not as a first-principles result. The main methodological weakness—snowball sampling seeded from Raji et al. (2021) and the non-exhaustive, unlisted corpus—is an external-validity and representativeness concern, not a circularity: the authors explicitly state the taxonomy is 'not exhaustive' and is a 'narrative tool to present our findings.' The few self-citations are minor and non-load-bearing: Noroozian (2020) is cited only as an example of benchmarking in security (Section 2), and Gomez et al. (2024) is cited as contextual support for diversity issues in the AI ecosystem (Section 5.4). Neither citation carries the paper's central argument, which rests on a broad set of external, independently published critiques. Therefore, no load-bearing step is equivalent by construction to its inputs, and the paper does not exhibit meaningful circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no free parameters or invented entities. It relies on domain assumptions about the reliability of the surveyed literature and the representativeness of its snowball sample; these are acknowledged in Section 4 but not independently verified.

assumptions (3)
  • domain assumption The critiques reported in the roughly 110 surveyed papers are accurate representations of real deficiencies in benchmarking practice.
    The review's conclusion that 'fundamental fragilities' exist is an aggregation of the surveyed critiques (Sections 5.1-5.9, 6). If many of these are unverified opinions or preprints, the synthesis inherits their biases.
  • domain assumption A snowball sample seeded from Raji et al. (2021) is representative of benchmark critique published 2014-2024.
    Section 4 describes the snowball method as a replacement for systematic keyword searches. This assumes the citation network from one critical paper reaches all relevant critique communities, including non-citation-network venues.
  • domain assumption Borrowed definitions of benchmarks, tasks, and metrics from Raji et al. (2021) and Schlangen (2020) delimit the scope of what counts as an AI benchmark.
    Section 2 adopts these definitions to restrict the review to quantitative benchmarks as operationalized by this prior work, which may exclude qualitative or alternative evaluation traditions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation." pith.science (2026). https://pith.science/paper/CLRQDII6

@misc{pith2026250206559,
  author       = {Pith},
  title        = {Pith review of: Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CLRQDII6}},
  note         = {Machine review of arXiv:2502.06559}
}
read the original abstract

Quantitative Artificial Intelligence (AI) Benchmarks have emerged as fundamental tools for evaluating the performance, capability, and safety of AI models and systems. Currently, they shape the direction of AI development and are playing an increasingly prominent role in regulatory frameworks. As their influence grows, however, so too does concerns about how and with what effects they evaluate highly sensitive topics such as capabilities, including high-impact capabilities, safety and systemic risks. This paper presents an interdisciplinary meta-review of about 100 studies that discuss shortcomings in quantitative benchmarking practices, published in the last 10 years. It brings together many fine-grained issues in the design and application of benchmarks (such as biases in dataset creation, inadequate documentation, data contamination, and failures to distinguish signal from noise) with broader sociotechnical issues (such as an over-focus on evaluating text-based AI models according to one-time testing logic that fails to account for how AI models are increasingly multimodal and interact with humans and other technical systems). Our review also highlights a series of systemic flaws in current benchmarking practices, such as misaligned incentives, construct validity issues, unknown unknowns, and problems with the gaming of benchmark results. Furthermore, it underscores how benchmark practices are fundamentally shaped by cultural, commercial and competitive dynamics that often prioritise state-of-the-art performance at the expense of broader societal concerns. By providing an overview of risks associated with existing benchmarking procedures, we problematise disproportionate trust placed in benchmarks and contribute to ongoing efforts to improve the accountability and relevance of quantitative AI benchmarks within the complexities of real-world scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems

    cs.AI 2026-08 unverdicted novelty 6.0 of 10

    MAAC defines a nine-dimension framework for process-oriented cognitive evaluation of text-based AI, shifting assessment from outputs to underlying reasoning.

  2. The Foreign Policy AI Evaluation Gap

    cs.CY 2026-07 conditional novelty 6.0 of 10

    Public technical AI governance almost never evaluates real foreign-policy AI workflows; the paper maps that gap and proposes task-scoped, human-recombined evaluation instead of model leaderboards.

  3. What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles

    cs.AI 2025-08 unverdicted novelty 6.0 of 10

    TurtleSoup-Bench is a new interactive benchmark showing that LLMs struggle with imaginative reasoning compared to humans.

  4. Deprecating Benchmarks: Criteria and Framework

    cs.CY 2025-07 conditional novelty 6.0 of 10

    A framework for deprecating outdated or flawed AI benchmarks, with seven criteria and a three-phase process of assessment, reporting, and notification.

  5. Lilith: Developmental Modular LLMs with Chemical Signaling

    q-bio.NC 2025-07 reject novelty 6.0 of 10

    A conceptual framework in which untrained modular LLMs are developed through simulated life and token-based chemical signaling, with the goal of enabling empirical study of consciousness emergence via Integrated Infor...

  6. Establishing Best Practices for Building Rigorous Agentic Benchmarks

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Agentic benchmarks frequently mis-grade agents, and the new ABC checklist helps identify and correct such errors in ten popular benchmarks.

  7. Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.

  8. VLM@school -- Evaluation of AI image understanding on German middle school knowledge

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new German middle school visual question-answering benchmark shows open-weight VLMs score below 45% overall, with especially weak results in music, math, and adversarial questions.

  9. Position: Explanation Stability Is a Property of the Model Method Pair, Not the Model

    cs.CV 2026-07 conditional novelty 5.0 of 10

    On chest X-rays, LayerCAM ranks InceptionV3 as the most stable model but Grad-CAM++ ranks DenseNet201 first, showing explanation stability is a property of the model-method pair.

  10. AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

    cs.AI 2026-02 conditional novelty 5.0 of 10

    VERA-MH, an automated safety benchmark for suicide-risk chatbot conversations, agreed with clinician ratings (IRR 0.81), though the clinical reference was not fully independent.

  11. A Conceptual Framework for AI Capability Evaluations

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A descriptive conceptual framework with seven elements (target, task, subject, inputs, instance, measurement, result analysis) for systematizing analysis of AI capability evaluations.

  12. From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation

    cs.CL 2025-06 conditional novelty 4.0 of 10

    On a new 490-question Arabic depth dataset, Claude 3.5 Sonnet answered about 30 percent correctly, while GPT-4 answered about 9 percent, showing current models are weak on culturally specialized Arabic knowledge.

  13. Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation

    cs.CR 2025-07 reject novelty 2.0 of 10

    The paper is a literature review that classifies privacy-preserving AI techniques in dataspaces using a qualitative taxonomy of privacy, performance, and compliance ratings.

Reference graph

Works this paper leans on

135 extracted references · 26 canonical work pages · cited by 13 Pith papers

  1. [1]

    Thompson

    Nur Ahmed, Muntasir Wahed, and Neil C. Thompson. The growing influence of industry in AI research. Science, 379 0 (6635): 0 884--886, March 2023. ISSN 0036-8075, 1095-9203. doi:10.1126/science.ade2420. URL https://www.science.org/doi/10.1126/science.ade2420

  2. [2]

    Field-building and the epistemic culture of AI safety

    Shazeda Ahmed, Klaudia Jaźwińska, Archana Ahlawat, Amy Winecoff, and Mona Wang. Field-building and the epistemic culture of AI safety. First Monday, April 2024. ISSN 1396-0466. doi:10.5210/fm.v29i4.13626. URL https://firstmonday.org/ojs/index.php/fm/article/view/13626

  3. [3]

    Saiful Bari, and Haidar Khan

    Norah Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed, Shaykhah Alsubaie, Yusef Almushaykeh, Faisal Mirza, Nouf Alotaibi, Nora Altwairesh, Areeb Alowisheq, M. Saiful Bari, and Haidar Khan. When Benchmarks are Targets : Revealing the Sensitivity of Large Language Model Leaderboards , July 2024. URL http://arxiv.org/abs/2402.01781

  4. [4]

    M. R. Aniba, O. Poch, and J. D. Thompson. Issues in bioinformatics benchmarking: the case study of multiple sequence alignment . Nucleic Acids Research, 38 0 (21): 0 7353--7363, 2010. URL https://doi.org/10.1093/nar/gkq625

  5. [5]

    Truth Is a Lie : Crowd Truth and the Seven Myths of Human Annotation

    Lora Aroyo and Chris Welty. Truth Is a Lie : Crowd Truth and the Seven Myths of Human Annotation . AI Magazine, 36 0 (1): 0 15--24, March 2015. ISSN 0738-4602, 2371-9621. doi:10.1609/aimag.v36i1.2564. URL https://onlinelibrary.wiley.com/doi/10.1609/aimag.v36i1.2564

  6. [6]

    Beyond the Numbers : Transparency in Relation Extraction Benchmark Creation and Leaderboards , November 2024

    Varvara Arzt and Allan Hanbury. Beyond the Numbers : Transparency in Relation Extraction Benchmark Creation and Leaderboards , November 2024. URL http://arxiv.org/abs/2411.05224

  7. [7]

    Experiences from using snowballing and database searches in systematic literature studies

    Deepika Badampudi, Claes Wohlin, and Kai Petersen. Experiences from using snowballing and database searches in systematic literature studies. In Proceedings of the 19th International Conference on Evaluation and Assessment in Software Engineering, EASE '15, New York, NY, USA, 2015. Association for Computing Machinery. ISBN 9781450333504. doi:10.1145/27458...

  8. [8]

    It's COMPASlicated : The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks , April 2022

    Michelle Bao, Angela Zhou, Samantha Zottola, Brian Brubach, Sarah Desmarais, Aaron Horowitz, Kristian Lum, and Suresh Venkatasubramanian. It's COMPASlicated : The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks , April 2022. URL http://arxiv.org/abs/2106.05498

Show all 135 references
  1. [9]

    Malan, Jason H

    Thomas Bartz-Beielstein, Carola Doerr, Daan van den Berg, Jakob Bossek, Sowmya Chandrasekaran, Tome Eftimov, Andreas Fischbach, Pascal Kerschke, William La Cava, Manuel Lopez-Ibanez, Katherine M. Malan, Jason H. Moore, Boris Naujoks, Patryk Orzechowski, Vanessa Volz, Markus Wa...

  2. [10]

    The Death of the Static AI Benchmark , March 2024

    Sandi Besen. The Death of the Static AI Benchmark , March 2024. URL https://towardsdatascience.com/the-death-of-the-static-ai-benchmark-88b5ff437086

  3. [11]

    Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick, Jason Phang, Aviya Skowron, Samson Tan, Xiangru Tang, Kevin A

    Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Jeffrey Hsu, Mimansa Jaiswal...

  4. [12]

    AI auditing: The Broken Bus on the Road to AI Accountability , January 2024

    Abeba Birhane, Ryan Steed, Victor Ojewale, Briana Vecchione, and Inioluwa Deborah Raji. AI auditing: The Broken Bus on the Road to AI Accountability , January 2024. URL http://arxiv.org/abs/2401.14462

  5. [13]

    Benchmark datasets driving artificial intelligence development fail to capture the needs of medical professionals

    Kathrin Blagec, Jakob Kraiger, Wolfgang Frühwirt, and Matthias Samwald. Benchmark datasets driving artificial intelligence development fail to capture the needs of medical professionals. Journal of Biomedical Informatics, 137: 0 104274, January 2023. ISSN 15320464. doi:10.1016...

  6. [14]

    Making Intelligence : Ethical Values in IQ and ML Benchmarks

    Borhane Blili-Hamelin and Leif Hancox-Li. Making Intelligence : Ethical Values in IQ and ML Benchmarks . In 2023 ACM Conference on Fairness , Accountability , and Transparency , pages 271--284, Chicago IL USA, June 2023. ACM. ISBN 9798400701924. doi:10.1145/3593013.3593996. UR...

  7. [15]

    Stereotyping Norwegian Salmon : An Inventory of Pitfalls in Fairness Benchmark Datasets

    Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. Stereotyping Norwegian Salmon : An Inventory of Pitfalls in Fairness Benchmark Datasets . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th ...

  8. [16]

    Bowman and George Dahl

    Samuel R. Bowman and George Dahl. What Will it Take to Fix Benchmarking in Natural Language Understanding ? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics : Human Language Technologies , pages 4843--4855, On...

  9. [17]

    Benchmarking, pages 363--368

    Isabelle Bruno. Benchmarking, pages 363--368. Springer Netherlands, Dordrecht, 2014. ISBN 978-94-007-0753-5. doi:10.1007/978-94-007-0753-5_170. URL https://doi.org/10.1007/978-94-007-0753-5_170

  10. [18]

    Evaluating AI Evaluation : Perils and Prospects , July 2024

    John Burden. Evaluating AI Evaluation : Perils and Prospects , July 2024. URL https://arxiv.org/abs/2407.09221v1

  11. [19]

    Ullman, Fernando Martinez-Plumed, Joshua B

    Ryan Burnell, Wout Schellaert, John Burden, Tomer D. Ullman, Fernando Martinez-Plumed, Joshua B. Tenenbaum, Danaja Rutar, Lucy G. Cheke, Jascha Sohl-Dickstein, Melanie Mitchell, Douwe Kiela, Murray Shanahan, Ellen M. Voorhees, Anthony G. Cohn, Joel Z. Leibo, and Jose Hernandez...

  12. [20]

    Robert C. Camp. Benchmarking : The Search for Industry Best Practices That Lead to Superior Performance. Quality Press , the University of Michigan , 1989

  13. [22]

    Yu, Qiang Yang, and Xing Xie

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. A Survey on Evaluation of Large Language Models , December 2023. URL http://arxiv.org/a...

  14. [23]

    Cheng, CS

    Z. Cheng, CS. Pang, P. Wang, and et al. How to report and benchmark emerging field-effect transistors. Nature Electronics, 5: 0 416--423, 2020. URL https://doi.org/10.1038/s41928-022-00798-8

  15. [24]

    On the Measure of Intelligence , November 2019

    François Chollet. On the Measure of Intelligence , November 2019. URL http://arxiv.org/abs/1911.01547

  16. [25]

    A survey of 25 years of evaluation

    Kenneth Ward Church and Joel Hestness. A survey of 25 years of evaluation. Natural Language Engineering, 25 0 (06): 0 753--767, November 2019. ISSN 1351-3249, 1469-8110. doi:10.1017/S1351324919000275. URL https://www.cambridge.org/core/product/identifier/S1351324919000275/type...

  17. [26]

    Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals

    Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals. The Benchmark Lottery , July 2021. URL http://arxiv.org/abs/2107.07002

  18. [27]

    On the genealogy of machine learning datasets: A critical history of ImageNet

    Emily Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, and Hilary Nicole. On the genealogy of machine learning datasets: A critical history of ImageNet . Big Data & Society, 8 0 (2): 0 1--14, July 2021. ISSN 2053-9517, 2053-9517. doi:10.1177/20539517211035955. URL http://j...

  19. [28]

    Bringing the People Back In : Contesting Benchmark Machine Learning Datasets , July 2020

    Remi Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, Hilary Nicole, and Morgan Klaus Scheuerman. Bringing the People Back In : Contesting Benchmark Machine Learning Datasets , July 2020. URL http://arxiv.org/abs/2007.07399

  20. [30]

    Agreements ‘in the wild’: Standards and alignment in machine learning benchmark dataset construction

    Isak Engdahl. Agreements ‘in the wild’: Standards and alignment in machine learning benchmark dataset construction. Big Data & Society, 11 0 (2): 0 20539517241242457, June 2024. ISSN 2053-9517, 2053-9517. doi:10.1177/20539517241242457. URL https://journals.sagepub.com/doi/10.1...

  21. [31]

    Utility is in the Eye of the User : A Critique of NLP Leaderboards , March 2021

    Kawin Ethayarajh and Dan Jurafsky. Utility is in the Eye of the User : A Critique of NLP Leaderboards , March 2021. URL http://arxiv.org/abs/2009.13888

  22. [32]

    European Union . Regulation (EU) 2022/2065 of the European Parliament and of the Council of 19 October 2022 on a Single Market For Digital Services and amending Directive 2000/31/EC (Digital Services Act) , 2022

  23. [33]

    First Draft of the General-Purpose AI Code of Practice published, written by independent experts , 2024 a

    European Union . First Draft of the General-Purpose AI Code of Practice published, written by independent experts , 2024 a

  24. [34]

    Second Draft of the General-Purpose AI Code of Practice published, written by independent experts , 2024 b

    European Union . Second Draft of the General-Purpose AI Code of Practice published, written by independent experts , 2024 b

  25. [35]

    European Union . Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence and amending Regulations (Artificial Intelligence Act) , 2024 c

  26. [36]

    Simon Frieder, Jonas Bayer, Katherine M. Collins, Julius Berner, Jacob Loader, András Juhász, Fabian Ruehle, Sean Welleck, Gabriel Poesia, Ryan-Rhys Griffiths, Adrian Weller, Anirudh Goyal, Thomas Lukasiewicz, and Timothy Gowers. Data for Mathematical Copilots : Better Ways of...

  27. [37]

    Datasheets for Datasets , December 2021

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. Datasheets for Datasets , December 2021. URL http://arxiv.org/abs/1803.09010

  28. [38]

    Repairing the Cracked Foundation : A Survey of Obstacles in Evaluation Practices for Generated Text

    Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. Repairing the Cracked Foundation : A Survey of Obstacles in Evaluation Practices for Generated Text . Journal of Artificial Intelligence Research, 77: 0 103--166, May 2023. ISSN 1076-9757. doi:10.1613/jair.1.13715. URL ...

  29. [39]

    Wichmann

    Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. Shortcut Learning in Deep Neural Networks . Nature Machine Intelligence, 2 0 (11): 0 665--673, November 2020. ISSN 2522-5839. doi:10.1038/s42256-020...

  30. [40]

    Are We Done with MMLU ?, June 2024

    Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile van Krieken, and Pasquale Minervini...

  31. [41]

    Diversity in artificial intelligence conferences

    Emilia Gomez, Porcaro Lorenzo, Pedro Frau Amar, and Joao Vinagre. Diversity in artificial intelligence conferences. Publications Office of the European Union JRC137550, Publications Office, 2024. URL https://data.europa.eu/doi/10.2760/796551

  32. [42]

    Bowman, and Evan Hubinger

    Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, Sören Mindermann, Ethan Perez, Linda Petrini, Jonathan Uesato, Jared Kaplan, Buck Shlegeris, Sam...

  33. [44]

    COMPL - AI Framework : A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act , October 2024

    Philipp Guldimann, Alexander Spiridonov, Robin Staab, Nikola Jovanović, Mark Vero, Velko Vechev, Anna Gueorguieva, Mislav Balunović, Nikola Konstantinov, Pavol Bielik, Petar Tsankov, and Martin Vechev. COMPL - AI Framework : A Technical Interpretation and LLM Benchmarking Suit...

  34. [45]

    Ai hype is built on high test scores

    Douglas Heaven. Ai hype is built on high test scores. those tests are flawed. Report, MIT Technology Review, 2023. URL https://www.technologyreview.com/2023/08/30/1078670/large-language-models-arent-people-lets-stop-testing-them-like-they-were/

  35. [46]

    Measuring Massive Multitask Language Understanding , January 2021

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring Massive Multitask Language Understanding , January 2021. URL http://arxiv.org/abs/2009.03300

  36. [47]

    J.L. Henning. Spec cpu2000: measuring cpu performance in the new millennium. Computer, 33 0 (7): 0 28--35, 2000. doi:10.1109/2.869367

  37. [48]

    On the Limitations of Compute Thresholds as a Governance Strategy , July 2024

    Sara Hooker. On the Limitations of Compute Thresholds as a Governance Strategy , July 2024. URL http://arxiv.org/abs/2407.05694

  38. [49]

    Evaluation Gaps in Machine Learning Practice

    Ben Hutchinson, Negar Rostamzadeh, Christina Greer, Katherine Heller, and Vinodkumar Prabhakaran. Evaluation Gaps in Machine Learning Practice . In 2022 ACM Conference on Fairness , Accountability , and Transparency , pages 1859--1876, Seoul Republic of Korea, June 2022. ACM. ...

  39. [50]

    Systematic literature studies: database searches vs

    Samireh Jalali and Claes Wohlin. Systematic literature studies: database searches vs. backward snowballing. In Proceedings of the ACM-IEEE International Symposium on Empirical Software Engineering and Measurement, ESEM '12, page 29–38, New York, NY, USA, 2012. Association for ...

  40. [51]

    Escaping the McNamara Fallacy : Toward More Impactful Recommender Systems Research

    Dietmar Jannach and Christine Bauer. Escaping the McNamara Fallacy : Toward More Impactful Recommender Systems Research . AI Magazine, 41 0 (4): 0 79--95, December 2020. ISSN 0738-4602, 2371-9621. doi:10.1609/aimag.v41i4.5312. URL https://onlinelibrary.wiley.com/doi/10.1609/ai...

  41. [52]

    The Constitution of Algorithms : Ground - Truthing , Programming , Formulating

    Florian Jaton. The Constitution of Algorithms : Ground - Truthing , Programming , Formulating . Inside Technology . The MIT Press, Cambridge, 2021. ISBN 978-0-262-54214-2 978-0-262-36323-5

  42. [53]

    Under the radar? examining the evaluation of foundation models

    Elliot Jones, Mahi Hardalupas, and William Agrew. Under the radar? examining the evaluation of foundation models. Report, Ada Lovelace Institute, 2024. URL https://www.adalovelaceinstitute.org/report/under-the-radar/

  43. [54]

    Ground truth tracings ( GTT ): On the epistemic limits of machine learning

    Edward B Kang. Ground truth tracings ( GTT ): On the epistemic limits of machine learning. Big Data & Society, 10 0 (1): 0 20539517221146122, January 2023. ISSN 2053-9517, 2053-9517. doi:10.1177/20539517221146122. URL https://journals.sagepub.com/doi/10.1177/20539517221146122

  44. [55]

    Siegel, Nitya Nadgir, and Arvind Narayanan

    Sayash Kapoor, Benedikt Stroebl, Zachary S. Siegel, Nitya Nadgir, and Arvind Narayanan. AI Agents That Matter , July 2024. URL http://arxiv.org/abs/2407.01502

  45. [56]

    Leakage in data mining: Formulation, detection, and avoidance

    Shachar Kaufman, Saharon Rosset, Claudia Perlich, and Ori Stitelman. Leakage in data mining: Formulation, detection, and avoidance. ACM Trans. Knowl. Discov. Data, 6 0 (4), December 2012. ISSN 1556-4681. doi:10.1145/2382577.2382579. URL https://doi.org/10.1145/2382577.2382579

  46. [57]

    Everyone Is Judging AI by These Tests

    Jon Keegan. Everyone Is Judging AI by These Tests . But Experts Say They ’re Close to Meaningless . The Markup, July 2024. URL https://themarkup.org/artificial-intelligence/2024/07/17/everyone-is-judging-ai-by-these-tests-but-experts-say-theyre-close-to-meaningless

  47. [58]

    Mulvehill, and Deborah L

    Mayank Kejriwal, Henrique Santos, Ke Shen, Alice M. Mulvehill, and Deborah L. McGuinness. A noise audit of human-labeled benchmarks for machine commonsense reasoning. Scientific Reports, 14 0 (1): 0 8609, April 2024. ISSN 2045-2322. doi:10.1038/s41598-024-58937-4. URL https://...

  48. [59]

    Feeling fixes: Mess and emotion in algorithmic audits

    Os Keyes and Jeanie Austin. Feeling fixes: Mess and emotion in algorithmic audits. Big Data & Society, 9 0 (2): 0 20539517221113772, July 2022. ISSN 2053-9517, 2053-9517. doi:10.1177/20539517221113772. URL https://journals.sagepub.com/doi/10.1177/20539517221113772

  49. [60]

    Bernard Koch, Emily Denton, Alex Hanna, and Jacob G. Foster. Reduced, Reused and Recycled : The Life of a Dataset in Machine Learning Research , December 2021. URL http://arxiv.org/abs/2112.01716

  50. [61]

    Koch and David Peterson

    Bernard J. Koch and David Peterson. From Protoscience to Epistemic Monoculture : How Benchmarking Set the Stage for the Deep Learning Revolution , April 2024. URL http://arxiv.org/abs/2404.06647

  51. [62]

    Metaethical Perspectives on ' Benchmarking ' AI Ethics , April 2022

    Travis LaCroix and Alexandra Sasha Luccioni. Metaethical Perspectives on ' Benchmarking ' AI Ethics , April 2022. URL http://arxiv.org/abs/2204.05151

  52. [63]

    Vazquez, Niclas Kupper, Misha Yagudin, and Laurence Aitchison

    Gavin Leech, Juan J. Vazquez, Niclas Kupper, Misha Yagudin, and Laurence Aitchison. Questionable practices in machine learning, July 2024. URL https://arxiv.org/abs/2407.12220v2

  53. [64]

    Question and answer test-train overlap in open-domain question answering datasets

    Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel. Question and answer test-train overlap in open-domain question answering datasets. In Paola Merlo, Jorg Tiedemann, and Reut Tsarfaty, editors, Proceedings of the 16th Conference of the European Chapter of the Association f...

  54. [65]

    Manning, Christopher Ré, Diana Acosta-Navas, Drew A

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas...

  55. [66]

    Vera Liao and Ziang Xiao

    Q. Vera Liao and Ziang Xiao. Rethinking Model Evaluation as Narrowing the Socio - Technical Gap , June 2023. URL http://arxiv.org/abs/2306.03100

  56. [67]

    Are we learning yet? a meta review of evaluation failures across machine learning

    Thomas Liao, Rohan Taori, Inioluwa Deborah Raji, and Ludwig Schmidt. Are we learning yet? a meta review of evaluation failures across machine learning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https:...

  57. [68]

    ExplainaBoard : An Explainable Leaderboard for NLP , July 2021

    Pengfei Liu, Jinlan Fu, Yang Xiao, Weizhe Yuan, Shuaicheng Chang, Junqi Dai, Yixin Liu, Zihuiwen Ye, Zi-Yi Dou, and Graham Neubig. ExplainaBoard : An Explainable Leaderboard for NLP , July 2021. URL http://arxiv.org/abs/2104.06387

  58. [69]

    AI competitions as infrastructures of power in medical imaging

    Dieuwertje Luitse, Tobias Blanke, and Thomas Poell. AI competitions as infrastructures of power in medical imaging. Information, Communication & Society, pages 1--22, March 2024. ISSN 1369-118X, 1468-4462. doi:10.1080/1369118X.2024.2334393. URL https://www.tandfonline.com/doi/...

  59. [70]

    Bias in Language Models : Beyond Trick Tests and Toward RUTEd Evaluation , February 2024

    Kristian Lum, Jacy Reese Anthis, Chirag Nagpal, and Alexander D'Amour. Bias in Language Models : Beyond Trick Tests and Toward RUTEd Evaluation , February 2024. URL https://arxiv.org/abs/2402.12649v1

  60. [71]

    Data contamination: From memorization to exploitation

    Inbal Magar and Roy Schwartz. Data contamination: From memorization to exploitation. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 15...

  61. [72]

    Practices of Benchmarking : Vulnerability in the Computer Vision Pipeline

    Nicolas Malevé. Practices of Benchmarking : Vulnerability in the Computer Vision Pipeline . photographies, 16 0 (2): 0 173--189, May 2023. ISSN 1754-0763, 1754-0771. doi:10.1080/17540763.2023.2189159. URL https://www.tandfonline.com/doi/full/10.1080/17540763.2023.2189159

  62. [73]

    Put to the test: For a new sociology of testing

    Noortje Marres and David Stark. Put to the test: For a new sociology of testing. The British Journal of Sociology, 71 0 (3): 0 423--443, June 2020. ISSN 0007-1315, 1468-4446. doi:10.1111/1468-4446.12746. URL https://onlinelibrary.wiley.com/doi/10.1111/1468-4446.12746

  63. [74]

    Mlperf training benchmark

    Peter Mattson, Christine Cheng, Gregory Diamos, Cody Coleman, Paulius Micikevicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, David Brooks, Dehao Chen, Debo Dutta, Udit Gupta, Kim Hazelwood, Andy Hock, Xinyuan Huang, Daniel Kang, David Kanter, Na...

  64. [75]

    Mlperf: An industry standard benchmark suite for machine learning performance

    Peter Mattson, Vijay Janapa Reddi, Christine Cheng, Cody Coleman, Greg Diamos, David Kanter, Paulius Micikevicius, David Patterson, Guenther Schmuelling, Hanlin Tang, Gu-Yeon Wei, and Carole-Jean Wu. Mlperf: An industry standard benchmark suite for machine learning performance...

  65. [76]

    McIntosh, Teo Susnjak, Nalin Arachchilage, Tong Liu, Paul Watters, and Malka N

    Timothy R. McIntosh, Teo Susnjak, Nalin Arachchilage, Tong Liu, Paul Watters, and Malka N. Halgamuge. Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence , October 2024. URL http://arxiv.org/abs/2402.09880

  66. [77]

    Frontier Models are Capable of In -context Scheming , December 2024

    Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn. Frontier Models are Capable of In -context Scheming , December 2024. URL https://arxiv.org/abs/2412.04984v1

  67. [78]

    Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, and Samuel R. Bowman. What Do NLP Researchers Believe ? Results of the NLP Community Metasurvey , August 2022. URL http://arx...

  68. [79]

    Benchmarking the Benchmarks

    Marc Miltenberger, Steven Arzt, Philipp Holzinger, and Julius Näumann. Benchmarking the Benchmarks . In Proceedings of the ACM Asia Conference on Computer and Communications Security , pages 387--400, Melbourne VIC Australia, July 2023. ACM. ISBN 9798400700989. doi:10.1145/357...

  69. [80]

    How do we know how smart AI systems are? Science, 381 0 (6654), July 2023

    Margaret Mitchell. How do we know how smart AI systems are? Science, 381 0 (6654), July 2023. doi:10.1126/science.adj595. URL https://www.science.org/doi/10.1126/science.adj5957

  70. [82]

    State of What Art ? A Call for Multi - Prompt LLM Evaluation , May 2024

    Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. State of What Art ? A Call for Multi - Prompt LLM Evaluation , May 2024. URL http://arxiv.org/abs/2401.00595

  71. [83]

    Proxies: The Cultural Work of Standing In

    Dylan Mulvin. Proxies: The Cultural Work of Standing In . Infrastructures. The MIT Press, Cambridge, 2021. ISBN 978-0-262-04514-8 978-0-262-36624-3

  72. [84]

    GPT -4 and professional benchmarks: the wrong answer to the wrong question, March 2023 a

    Arvind Narayanan and Sayash Kapoor. GPT -4 and professional benchmarks: the wrong answer to the wrong question, March 2023 a . URL https://www.aisnakeoil.com/p/gpt-4-and-professional-benchmarks

  73. [85]

    Evaluating LLMs is a minefield, 2023 b

    Arvind Narayanan and Sayash Kapoor. Evaluating LLMs is a minefield, 2023 b . URL https://www.cs.princeton.edu/ arvindn/talks/evaluating_llms_minefield/

  74. [86]

    Feder Cooper, Daphne Ippolito, Christopher A

    Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, and Katherine Lee. Extracting Training Data from ChatGPT , November 2023 a . URL https://not-just-memorization.github.io/extracting-...

  75. [87]

    Feder Cooper, Daphne Ippolito, Christopher A

    Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee. Scalable Extraction of Training Data from ( Production ) Language Models , November 2023 b . URL ...

  76. [88]

    Evaluating Hosting Provider Security Through Abuse Data and the Creation of Metrics

    Arman Noroozian. Evaluating Hosting Provider Security Through Abuse Data and the Creation of Metrics . Dissertation ( TU Delft ), 2020. ISBN: 9789065624451

  77. [89]

    Hidden Stratification Causes Clinically Meaningful Failures in Machine Learning for Medical Imaging , November 2019

    Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro, and Christopher Ré. Hidden Stratification Causes Clinically Meaningful Failures in Machine Learning for Medical Imaging , November 2019. URL http://arxiv.org/abs/1909.12475

  78. [90]

    Towards AI Accountability Infrastructure : Gaps and Opportunities in AI Audit Tooling , March 2024

    Victor Ojewale, Ryan Steed, Briana Vecchione, Abeba Birhane, and Inioluwa Deborah Raji. Towards AI Accountability Infrastructure : Gaps and Opportunities in AI Audit Tooling , March 2024. URL http://arxiv.org/abs/2402.17861

  79. [91]

    The social construction of datasets: On the practices, processes, and challenges of dataset creation for machine learning

    Will Orr and Kate Crawford. The social construction of datasets: On the practices, processes, and challenges of dataset creation for machine learning. New Media & Society, 26 0 (9): 0 4955--4972, 2024 a

  80. [92]

    Building Better Datasets : Seven Recommendations for Responsible Design from Dataset Creators , August 2024 b

    Will Orr and Kate Crawford. Building Better Datasets : Seven Recommendations for Responsible Design from Dataset Creators , August 2024 b . URL http://arxiv.org/abs/2409.00252

  81. [93]

    Will Orr and Edward B. Kang. AI as a Sport : On the Competitive Epistemologies of Benchmarking . In The 2024 ACM Conference on Fairness , Accountability , and Transparency , pages 1875--1884, Rio de Janeiro Brazil, June 2024. ACM. ISBN 9798400704505. doi:10.1145/3630106.365901...

  82. [94]

    Mapping global dynamics of benchmark creation and saturation in artificial intelligence

    Simon Ott, Adriano Barbosa-Silva, Kathrin Blagec, Jan Brauner, and Matthias Samwald. Mapping global dynamics of benchmark creation and saturation in artificial intelligence. Nature Communications, 13 0 (1): 0 6793, November 2022. ISSN 2041-1723. doi:10.1038/s41467-022-34591-0....

  83. [95]

    Benchmark

    Oxford English Dictionary . Benchmark. meaning and use, 2017. URL https://www.oed.com/dictionary/benchmark_n?tab=meaning_and_use&tl=true

  84. [96]

    Cheke, and José Hernández-Orallo

    Lorenzo Pacchiardi, Marko Tesic, Lucy G. Cheke, and José Hernández-Orallo. Leaving the barn door open for Clever Hans : Simple features predict LLM benchmark answers, October 2024. URL http://arxiv.org/abs/2410.11672

  85. [97]

    Raison d’être of the benchmark dataset: A Survey of Current Practices of Benchmark Dataset Sharing Platforms

    Jaihyun Park and Sullam Jeoung. Raison d’être of the benchmark dataset: A Survey of Current Practices of Benchmark Dataset Sharing Platforms . In Proceedings of NLP Power ! The First Workshop on Efficient Benchmarking in NLP , pages 1--10, Dublin, Ireland, 2022. Association fo...

  86. [98]

    Bender, Emily Denton, and Alex Hanna

    Amandalynne Paullada, Inioluwa Deborah Raji, Emily M. Bender, Emily Denton, and Alex Hanna. Data and its (dis)contents: A survey of dataset development and use in machine learning research. Patterns, 2 0 (11): 0 100336, November 2021. ISSN 26663899. doi:10.1016/j.patter.2021.1...

  87. [99]

    Understanding and Benchmarking Artificial Intelligence : OpenAI 's o3 Is Not AGI , January 2025

    Rolf Pfister and Hansueli Jud. Understanding and Benchmarking Artificial Intelligence : OpenAI 's o3 Is Not AGI , January 2025. URL http://arxiv.org/abs/2501.07458

  88. [100]

    Testing - One , Two , Three ... Testing !

    Trevor Pinch. " Testing - One , Two , Three ... Testing !": Toward a Sociology of Testing . Science, Technology, & Human Values, 18 0 (1): 0 25--41, January 1993. ISSN 0162-2439, 1552-8251. doi:10.1177/016224399301800103. URL https://journals.sagepub.com/doi/10.1177/016224399301800103

  89. [101]

    The Roles of English in Evaluating Multilingual Language Models , December 2024

    Wessel Poelman and Miryam de Lhoneux. The Roles of English in Evaluating Multilingual Language Models , December 2024. URL http://arxiv.org/abs/2412.08392

  90. [102]

    Fine-tuning Aligned Language Models Compromises Safety , Even When Users Do Not Intend To !, October 2023

    Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. Fine-tuning Aligned Language Models Compromises Safety , Even When Users Do Not Intend To !, October 2023. URL http://arxiv.org/abs/2310.03693

  91. [103]

    Bender, Alex Hanna, and Amandalynne Paullada

    Inioluwa Deborah Raji, Emily Denton, Emily M. Bender, Alex Hanna, and Amandalynne Paullada. AI and the everything in the whole wide world benchmark. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https://o...

  92. [104]

    Gaps in the Safety Evaluation of Generative AI

    Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Ramona Comanescu, Canfer Akbulut, Tom Stepleton, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, William Isaac, and Laura Weidinger. Gaps in the Safet...

  93. [105]

    Safetywashing: Do AI safety benchmarks actually measure safety progress? In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024

    Richard Ren, Steven Basart, Adam Khoja, Alice Gatti, Long Phan, Xuwang Yin, Mantas Mazeika, Alexander Pan, Gabriel Mukobi, Ryan Hwang Kim, Stephen Fitz, and Dan Hendrycks. Safetywashing: Do AI safety benchmarks actually measure safety progress? In The Thirty-eight Conference o...

  94. [106]

    Betterbench: Assessing AI benchmarks, uncovering issues, and establishing best practices

    Anka Reuel, Amelia Hardy, Chandler Smith, Max Lamparth, Malcolm Hardy, and Mykel Kochenderfer. Betterbench: Assessing AI benchmarks, uncovering issues, and establishing best practices. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchma...

  95. [107]

    Data Contamination Through the Lens of Time , October 2023

    Manley Roberts, Himanshu Thakur, Christine Herlihy, Colin White, and Samuel Dooley. Data Contamination Through the Lens of Time , October 2023. URL http://arxiv.org/abs/2310.10628

  96. [108]

    Lalor, Robin Jia, and Jordan Boyd-Graber

    Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, and Jordan Boyd-Graber. Evaluation Examples are not Equally Informative : How should that change NLP Leaderboards ? In Proceedings of the 59th Annual Meeting of the Association for Computational L...

  97. [109]

    Kevin Roose. A.i. has a measurement problem. Report, New York Times, 2024. URL https://www.nytimes.com/2024/04/15/technology/ai-models-measurement.html

  98. [110]

    SafetyPrompts : a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety , April 2024

    Paul Röttger, Fabio Pernisi, Bertie Vidgen, and Dirk Hovy. SafetyPrompts : a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety , April 2024. URL http://arxiv.org/abs/2404.05399

  99. [111]

    Everyone wants to do the model work, not the data work

    Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. “ Everyone wants to do the model work, not the data work”: Data Cascades in High - Stakes AI . In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems...

  100. [112]

    Do Datasets Have Politics ? Disciplinary Values in Computer Vision Dataset Development

    Morgan Klaus Scheuerman, Alex Hanna, and Emily Denton. Do Datasets Have Politics ? Disciplinary Values in Computer Vision Dataset Development . Proceedings of the ACM on Human-Computer Interaction, 5 0 (CSCW2): 0 1--37, October 2021. ISSN 2573-0142. doi:10.1145/3476058. URL ht...

  101. [113]

    Targeting the Benchmark : On Methodology in Current Natural Language Processing Research , July 2020

    David Schlangen. Targeting the Benchmark : On Methodology in Current Natural Language Processing Research , July 2020. URL http://arxiv.org/abs/2007.04792

  102. [114]

    Winner's Curse ? On Pace , Progress , and Empirical Rigor

    D Sculley, Jasper Snoek, Ali Rahimi, and Alex Wiltschko. Winner's Curse ? On Pace , Progress , and Empirical Rigor . Vancouver, BC, Canada, 2018. URL https://openreview.net/pdf?id=rJWF0Fywf

  103. [115]

    Selbst, Danah Boyd, Sorelle A

    Andrew D. Selbst, Danah Boyd, Sorelle A. Friedler, Suresh Venkatasubramanian, and Janet Vertesi. Fairness and Abstraction in Sociotechnical Systems . In Proceedings of the Conference on Fairness , Accountability , and Transparency , pages 59--68, Atlanta GA USA, January 2019. ...

  104. [116]

    Arafat

    Shilad Sen, Margaret E. Giesel, Rebecca Gold, Benjamin Hillmann, Matt Lesicko, Samuel Naden, Jesse Russell, Zixiao (Ken) Wang, and Brent Hecht. Turkers, Scholars , " Arafat " and " Peace ": Cultural Communities and Algorithmic Gold Standards . In Proceedings of the 18th ACM Co...

  105. [117]

    Lazy Data Practices Harm Fairness Research

    Jan Simson, Alessandro Fabris, and Christoph Kern. Lazy Data Practices Harm Fairness Research . In The 2024 ACM Conference on Fairness , Accountability , and Transparency , pages 642--659, Rio de Janeiro Brazil, June 2024. ACM. ISBN 9798400704505. doi:10.1145/3630106.3658931. ...

  106. [118]

    Smith, Saleema Amershi, Solon Barocas, Hanna Wallach, and Jennifer Wortman Vaughan

    Jessie J. Smith, Saleema Amershi, Solon Barocas, Hanna Wallach, and Jennifer Wortman Vaughan. REAL ML : Recognizing , Exploring , and Articulating Limitations of Machine Learning Research . In 2022 ACM Conference on Fairness , Accountability , and Transparency , pages 587--597...

  107. [119]

    Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, et al

    Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, et al. Beyond the Imitati...

  108. [120]

    Another science is possible: a manifesto for slow science

    Isabelle Stengers. Another science is possible: a manifesto for slow science. Polity press, Cambridge, 2018. ISBN 978-1-5095-2180-7

  109. [121]

    The olympics of ai: Benchmarking machine learning systems, 2023

    Matthew Stewart. The olympics of ai: Benchmarking machine learning systems, 2023. URL https://towardsdatascience.com/the-olympics-of-ai-benchmarking-machine-learning-systems-c4b2051fbd2b

  110. [122]

    ‘ Improving ratings’: audit in the British University system

    Marilyn Strathern. ‘ Improving ratings’: audit in the British University system. European Review, 5 0 (3): 0 305--321, July 1997. ISSN 10627987, 1234981X. doi:10.1002/(SICI)1234-981X(199707)5:3<305::AID-EURO184>3.0.CO;2-4. URL https://www.cambridge.org/core/product/identifier/...

  111. [123]

    It Takes Two to Tango : Navigating Conceptualizations of NLP Tasks and Measurements of Performance , May 2023

    Arjun Subramonian, Xingdi Yuan, Hal Daumé III, and Su Lin Blodgett. It Takes Two to Tango : Navigating Conceptualizations of NLP Tasks and Measurements of Performance , May 2023. URL http://arxiv.org/abs/2305.09022

  112. [124]

    BEIR : A heterogeneous benchmark for zero-shot evaluation of information retrieval models

    Nandan Thakur, Nils Reimers, Andreas R \"u ckl \'e , Abhishek Srivastava, and Iryna Gurevych. BEIR : A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks ...

  113. [125]

    Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence , 2023

    The White House . Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence , 2023

  114. [126]

    Politics of data reuse in machine learning systems: Theorizing reuse entanglements

    Nanna Bonde Thylstrup, Kristian Bondo Hansen, Mikkel Flyverbom, and Louise Amoore. Politics of data reuse in machine learning systems: Theorizing reuse entanglements. Big Data & Society, 9 0 (2): 0 20539517221139785, July 2022. ISSN 2053-9517, 2053-9517. doi:10.1177/2053951722...

  115. [127]

    Memorization Without Overfitting : Analyzing the Training Dynamics of Large Language Models

    Kushal Tirumala, Aram Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. Memorization Without Overfitting : Analyzing the Training Dynamics of Large Language Models . In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information ...

  116. [128]

    From ImageNet to Image Classification : Contextualizing Progress on Benchmarks , May 2020

    Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas, and Aleksander Madry. From ImageNet to Image Classification : Contextualizing Progress on Benchmarks , May 2020. URL http://arxiv.org/abs/2005.11295

  117. [129]

    Online Safety Act 2023 , 2023

    UK Parliament . Online Safety Act 2023 , 2023

  118. [130]

    Framework for Artificial Intelligence Diffusion , 2025

    US Department of Commerce . Framework for Artificial Intelligence Diffusion , 2025

  119. [131]

    Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan

    Keyon Vafa, Justin Y. Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan. Evaluating the World Model Implicit in a Generative Model , November 2024. URL http://arxiv.org/abs/2406.03689

  120. [132]

    Sociotechnical Safety Evaluation of Generative AI Systems , October 2023

    Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, and William Isaac. Sociotechnical Safety Evaluation of Generative AI Systems , Octobe...

  121. [133]

    Brown, and Francis Rhys Ward

    Teun van der Weij, Felix Hofstätter, Ollie Jaffe, Samuel F. Brown, and Francis Rhys Ward. AI Sandbagging : Language Models can Strategically Underperform on Evaluations , June 2024. URL http://arxiv.org/abs/2406.07358

  122. [134]

    Benchmark Data Contamination of Large Language Models : A Survey , June 2024 a

    Cheng Xu, Shuhao Guan, Derek Greene, and M.-Tahar Kechadi. Benchmark Data Contamination of Large Language Models : A Survey , June 2024 a . URL http://arxiv.org/abs/2406.04244

  123. [135]

    Benchmarking Benchmark Leakage in Large Language Models , April 2024 b

    Ruijie Xu, Zengzhi Wang, Run-Ze Fan, and Pengfei Liu. Benchmarking Benchmark Leakage in Large Language Models , April 2024 b . URL http://arxiv.org/abs/2404.18824

  124. [136]

    Gonzalez, and Ion Stoica

    Shuo Yang, Wei-Lin Chiang, Lianmin Zheng, Joseph E. Gonzalez, and Ion Stoica. Rethinking Benchmark and Contamination for Language Models with Rephrased Samples , November 2023. URL http://arxiv.org/abs/2311.04850

  125. [137]

    Revisiting Out -of-distribution Robustness in NLP : Benchmark , Analysis , and LLMs Evaluations

    Lifan Yuan, Yangyi Chen, Ganqu Cui, Hongcheng Gao, Fangyuan Zou, Xingyi Cheng, Heng Ji, Zhiyuan Liu, and Maosong Sun. Revisiting Out -of-distribution Robustness in NLP : Benchmark , Analysis , and LLMs Evaluations . 37th Conference on Neural Information Processing Systems (Neu...

  126. [138]

    Zhang, Kevin Klyman, Yifan Mai, Yoav Levine, Yian Zhang, Rishi Bommasani, and Percy Liang

    Andy K. Zhang, Kevin Klyman, Yifan Mai, Yoav Levine, Yian Zhang, Rishi Bommasani, and Percy Liang. Language model developers should report train-test overlap, October 2024. URL http://arxiv.org/abs/2410.08385

  127. [139]

    Top llms in china and the u.s

    Lin Zhijia. Top llms in china and the u.s. only 5 months apart: Kai-fu lee, 2024. URL https://en.tmtpost.com/post/7289212

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.