REVIEW 3 major objections 5 minor 13 cited by
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
T0 review · 3 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Quantitative AI benchmarks, as currently designed and used, cannot be trusted to single-handedly provide the capability and safety assurances that policymakers need.
desk verdict A useful, policy-facing synthesis of benchmark critique that is honest about its own method, but its 'fundamental fragilities' conclusion is stronger than the non-systematic sample can fully support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the nine-issue taxonomy: problems with data collection, annotation, and documentation; weak construct validity; sociocultural context and gap; narrow diversity and scope; economic, competitive, and commercial roots; rigging, gaming, and measure-becoming-target; dubious community vetting and path dependencies; rapid AI development and benchmark saturation; and AI complexity and unknown unknowns. This taxonomy does the argumentative work by treating the issues as deeply interlinked, so that no single fix to individual benchmarks can restore trust by itself. It also supplies a shared vocabulary for policy discussions, showing that benchmark failure is not one bug but a class of failure modes, each of which can be identified in specific published studies.
What would settle it
One concrete test would be to run a preregistered, systematic literature search of the 2014-2024 period using database keyword queries instead of citation snowballing. If that search surfaces a substantial body of studies in which benchmark scores are shown to predict real-world deployment outcomes reliably, the paper's claim of fundamental fragility would be weakened. A second test is prospective: on a cohort of newly released models, compare benchmark scores against independent red-team findings; if the scores reliably flag the same models as unsafe, the taxonomy's warning would be overstated.
Extended reading notes
Core claim
On the paper's own terms, the discovery is not a single failed benchmark but a pattern: the same fragilities recur across text, image, audio, and multimodal evaluation, and they are structural rather than incidental. Each issue in the taxonomy shows how a benchmark's number can come apart from what it claims to report — a dataset can encode hidden spurious cues, a safety score can track general capability instead of safety, a model can be trained to underperform on purpose, and a benchmark can become a standard through citation luck rather than through demonstrated validity. Taken together, the paper concludes, these issues point toward fundamental fragilities in current efforts to quantitatively measure and mitigate harm in AI. The specific policy claim is that quantitative benchmarking is currently ill-suited to provide, on its own or even primarily, the safety and capability assurances requested by policymakers, and that what is needed is standardized methods for assessing the trustworthiness of benchmarks rather than standardized benchmarks themselves.
Load-bearing premise
The paper assumes that following citations out from one widely shared critique of benchmarks is enough to find all the important problems; if people who rely on or defend benchmarks published their evidence outside that citation trail, the list of nine issues could be incomplete.
Editorial extensions
If this is right
- Regulators should not treat benchmark scores as sufficient evidence for classifying a model as high-risk or systemically risky under the EU AI Act.
- Benchmark results should be published with full documentation, reproduction scripts, multiple evaluation runs, and reported statistical significance, since few current benchmarks provide them.
- Safety benchmark scores that correlate with general capability should not be marketed as safety progress; separate validation of the safety construct is required.
- Static, publicly known benchmarks will keep losing value through data contamination and sandbagging, so dynamic or hidden test sets and human-in-the-loop methods need to become part of the default toolkit.
- Trustworthy benchmarks require a standard way to evaluate the benchmarks themselves, not necessarily standardized benchmark metrics.
Reading between the lines
- We infer that the nine-issue taxonomy could be turned into a meta-evaluation scorecard: a checklist benchmark users apply before relying on a score, yielding an aggregate trustworthiness rating.
- An implicit testable extension is to measure how much a benchmark's score moves when each of the nine issues is mitigated; if removing one issue shifts results far more than others, the fragilities do not all carry equal weight.
- We infer that as laws hardwire benchmarks into regulatory thresholds, the incentive to game, sandbag, or contaminate those exact benchmarks will grow, so a benchmark's regulatory value is partly a function of how hard it is to optimize adversarially.
- The review focuses on quantitative benchmarks, but its own logic suggests that qualitative methods such as red-teaming and bug-bounty programs should be treated as complementary checks whose trustworthiness is assessed with the same scrutiny.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is an interdisciplinary meta-review of roughly 110 publications, published between 2014 and 2024, that critique quantitative AI benchmarking. The authors organize the critique into a nine-issue taxonomy covering data collection and documentation, construct validity, sociocultural context, diversity and scope, economic and competitive pressures, gaming and rigging, community vetting, benchmark saturation, and AI complexity/unknown unknowns. They argue that these issues are interlinked and that, taken together, they reveal fundamental fragilities in current efforts to quantitatively measure and mitigate AI harm. The paper concludes that quantitative benchmarking is currently ill-suited to single-handedly provide the safety and capability assurances that policy makers require, and it offers a set of policy-oriented recommendations for improving benchmark trustworthiness.
Significance. If the central claim is accepted, the paper provides a valuable and timely synthesis of a rapidly growing body of critique, with direct relevance to AI regulation and the EU AI Act. Its strengths include a genuinely interdisciplinary scope, inclusion of very recent work (more than half of the reviewed sources are from 2023 or later), and an explicit acknowledgment of the taxonomy's limitations as a 'narrative tool.' The paper also makes concrete, actionable recommendations, such as standardizing methods for assessing benchmark trustworthiness rather than standardizing benchmarks themselves. However, because the review is built on a non-systematic snowball sample and deliberately excludes benchmark-proposal papers, the strength of the general conclusion about 'fundamental fragilities' exceeds what the methodology can strictly support. The paper is useful as a high-quality qualitative synthesis, but its central claim needs to be conditioned on the nature of the sampled literature.
major comments (3)
- [Section 4 and Section 6] The snowball sampling procedure described in Section 4 starts from a single critique paper (Raji et al., 2021) and expands through its reference list and forward citations. This procedure is likely to over-represent works that engage with that specific critique and under-represent independent lines of benchmark critique or defense. The authors do not provide the core list of about 110 sources or a coding protocol, making it impossible to assess the completeness or bias of the corpus. Given that the conclusion in Section 6 states that 'taken together, these issues point toward fundamental fragilities in current efforts to quantitatively measure and mitigate harm in AI' and that quantitative benchmarking is 'ill-suited to single-handedly... provide the safety and capability assurances requested by policy makers,' the inference from this sample to the field as a whole is not established at the level of representativeness it asserts. I recommend that the authors provide the full corpus as supplementary material, describe the classification procedure in enough detail to be reproducible, and either soften the conclusion to explicitly refer to the reviewed literature or justify the representativeness of the sample.
- [Section 4, exclusion criteria] The authors exclude papers that propose new benchmarks, 'even though such articles naturally contain some level of benchmark critique.' This exclusion removes a substantial portion of the literature that might contain counterevidence or alternative perspectives on whether benchmarks are fundamentally fragile or whether the problems are corrigible. While the authors give pragmatic reasons for the exclusion, the central claim is about current benchmarking practices as a whole, not merely about the subset of papers whose primary purpose is critique. The authors should at least discuss whether the excluded literature could alter the conclusions, or restrict the conclusion to the population of critique-oriented publications.
- [Section 4, taxonomy development] The nine issue categories were identified after close reading and internal discussion, without a reported coding protocol or reliability checks. The authors acknowledge this and call the taxonomy 'a narrative tool,' but the conclusion treats the presence of these nine issues as evidence of 'fundamental fragilities.' Since the categories are not shown to be exhaustive or mutually exclusive, the conclusion should be presented as a synthesis of the reviewed critical literature rather than as a comprehensive map of all benchmarking problems. I recommend adding a limitations paragraph that explicitly states that the taxonomy has not been validated as a complete enumeration and that the conclusions reflect the reviewed sources.
minor comments (5)
- [Abstract] The sentence 'so too does concerns about how and with what effects...' should read 'so too do concerns...' (subject-verb agreement).
- [Section 2] 'heterogenous' should be 'heterogeneous'.
- [Section 6] In the sentence 'we especially identify a need for new ways of signallingwhat benchmarks to trust,' there is a missing space between 'signalling' and 'what.'
- [Section 5.6 and References] The in-text citation 'Weij et al. [2024]' refers to the reference 'van der Weij, Teun, Felix Hofstätter, et al. AI Sandbagging: Language Models can Strategically Underperform on Evaluations.' For consistency, the in-text citation should be 'van der Weij et al.'
- [Section 4] The authors refer to 'approximately 110' sources but never provide a table or appendix listing them. A supplementary list would help readers verify the coverage and would strengthen the reproducibility of the review.
Circularity Check
No significant circularity: the paper is a narrative meta-review whose synthesis is not derived from its inputs by construction; the only self-citations are minor and not load-bearing.
full rationale
The paper's derivation chain is a literature-based meta-review: snowball sampling (Section 4), a nine-issue taxonomy (Section 5), and a concluding synthesis (Section 6). There are no mathematical derivations, fitted parameters, or empirical predictions that could reduce to input data by construction. The central claim that benchmarks exhibit 'fundamental fragilities' is presented as a synthesis of the surveyed critique literature, not as a first-principles result. The main methodological weakness—snowball sampling seeded from Raji et al. (2021) and the non-exhaustive, unlisted corpus—is an external-validity and representativeness concern, not a circularity: the authors explicitly state the taxonomy is 'not exhaustive' and is a 'narrative tool to present our findings.' The few self-citations are minor and non-load-bearing: Noroozian (2020) is cited only as an example of benchmarking in security (Section 2), and Gomez et al. (2024) is cited as contextual support for diversity issues in the AI ecosystem (Section 5.4). Neither citation carries the paper's central argument, which rests on a broad set of external, independently published critiques. Therefore, no load-bearing step is equivalent by construction to its inputs, and the paper does not exhibit meaningful circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption The critiques reported in the roughly 110 surveyed papers are accurate representations of real deficiencies in benchmarking practice.
- domain assumption A snowball sample seeded from Raji et al. (2021) is representative of benchmark critique published 2014-2024.
- domain assumption Borrowed definitions of benchmarks, tasks, and metrics from Raji et al. (2021) and Schlangen (2020) delimit the scope of what counts as an AI benchmark.
Cite this review
Pith. "Pith review of Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation." pith.science (2026). https://pith.science/paper/CLRQDII6
@misc{pith2026250206559,
author = {Pith},
title = {Pith review of: Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/CLRQDII6}},
note = {Machine review of arXiv:2502.06559}
}
read the original abstract
Quantitative Artificial Intelligence (AI) Benchmarks have emerged as fundamental tools for evaluating the performance, capability, and safety of AI models and systems. Currently, they shape the direction of AI development and are playing an increasingly prominent role in regulatory frameworks. As their influence grows, however, so too does concerns about how and with what effects they evaluate highly sensitive topics such as capabilities, including high-impact capabilities, safety and systemic risks. This paper presents an interdisciplinary meta-review of about 100 studies that discuss shortcomings in quantitative benchmarking practices, published in the last 10 years. It brings together many fine-grained issues in the design and application of benchmarks (such as biases in dataset creation, inadequate documentation, data contamination, and failures to distinguish signal from noise) with broader sociotechnical issues (such as an over-focus on evaluating text-based AI models according to one-time testing logic that fails to account for how AI models are increasingly multimodal and interact with humans and other technical systems). Our review also highlights a series of systemic flaws in current benchmarking practices, such as misaligned incentives, construct validity issues, unknown unknowns, and problems with the gaming of benchmark results. Furthermore, it underscores how benchmark practices are fundamentally shaped by cultural, commercial and competitive dynamics that often prioritise state-of-the-art performance at the expense of broader societal concerns. By providing an overview of risks associated with existing benchmarking procedures, we problematise disproportionate trust placed in benchmarks and contribute to ongoing efforts to improve the accountability and relevance of quantitative AI benchmarks within the complexities of real-world scenarios.
Forward citations
Cited by 13 Pith papers
-
Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems
MAAC defines a nine-dimension framework for process-oriented cognitive evaluation of text-based AI, shifting assessment from outputs to underlying reasoning.
-
The Foreign Policy AI Evaluation Gap
Public technical AI governance almost never evaluates real foreign-policy AI workflows; the paper maps that gap and proposes task-scoped, human-recombined evaluation instead of model leaderboards.
-
What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles
TurtleSoup-Bench is a new interactive benchmark showing that LLMs struggle with imaginative reasoning compared to humans.
-
Deprecating Benchmarks: Criteria and Framework
A framework for deprecating outdated or flawed AI benchmarks, with seven criteria and a three-phase process of assessment, reporting, and notification.
-
Lilith: Developmental Modular LLMs with Chemical Signaling
A conceptual framework in which untrained modular LLMs are developed through simulated life and token-based chemical signaling, with the goal of enabling empirical study of consciousness emergence via Integrated Infor...
-
Establishing Best Practices for Building Rigorous Agentic Benchmarks
Agentic benchmarks frequently mis-grade agents, and the new ABC checklist helps identify and correct such errors in ten popular benchmarks.
-
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.
-
VLM@school -- Evaluation of AI image understanding on German middle school knowledge
A new German middle school visual question-answering benchmark shows open-weight VLMs score below 45% overall, with especially weak results in music, math, and adversarial questions.
-
Position: Explanation Stability Is a Property of the Model Method Pair, Not the Model
On chest X-rays, LayerCAM ranks InceptionV3 as the most stable model but Grad-CAM++ ranks DenseNet201 first, showing explanation stability is a property of the model-method pair.
-
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
VERA-MH, an automated safety benchmark for suicide-risk chatbot conversations, agreed with clinician ratings (IRR 0.81), though the clinical reference was not fully independent.
-
A Conceptual Framework for AI Capability Evaluations
A descriptive conceptual framework with seven elements (target, task, subject, inputs, instance, measurement, result analysis) for systematizing analysis of AI capability evaluations.
-
From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation
On a new 490-question Arabic depth dataset, Claude 3.5 Sonnet answered about 30 percent correctly, while GPT-4 answered about 9 percent, showing current models are weak on culturally specialized Arabic knowledge.
-
Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation
The paper is a literature review that classifies privacy-preserving AI techniques in dataspaces using a qualitative taxonomy of privacy, performance, and compliance ratings.
Reference graph
Works this paper leans on
-
[1]
Nur Ahmed, Muntasir Wahed, and Neil C. Thompson. The growing influence of industry in AI research. Science, 379 0 (6635): 0 884--886, March 2023. ISSN 0036-8075, 1095-9203. doi:10.1126/science.ade2420. URL https://www.science.org/doi/10.1126/science.ade2420
-
[2]
Field-building and the epistemic culture of AI safety
Shazeda Ahmed, Klaudia Jaźwińska, Archana Ahlawat, Amy Winecoff, and Mona Wang. Field-building and the epistemic culture of AI safety. First Monday, April 2024. ISSN 1396-0466. doi:10.5210/fm.v29i4.13626. URL https://firstmonday.org/ojs/index.php/fm/article/view/13626
-
[3]
Norah Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed, Shaykhah Alsubaie, Yusef Almushaykeh, Faisal Mirza, Nouf Alotaibi, Nora Altwairesh, Areeb Alowisheq, M. Saiful Bari, and Haidar Khan. When Benchmarks are Targets : Revealing the Sensitivity of Large Language Model Leaderboards , July 2024. URL http://arxiv.org/abs/2402.01781
arXiv 2024
-
[4]
M. R. Aniba, O. Poch, and J. D. Thompson. Issues in bioinformatics benchmarking: the case study of multiple sequence alignment . Nucleic Acids Research, 38 0 (21): 0 7353--7363, 2010. URL https://doi.org/10.1093/nar/gkq625
-
[5]
Truth Is a Lie : Crowd Truth and the Seven Myths of Human Annotation
Lora Aroyo and Chris Welty. Truth Is a Lie : Crowd Truth and the Seven Myths of Human Annotation . AI Magazine, 36 0 (1): 0 15--24, March 2015. ISSN 0738-4602, 2371-9621. doi:10.1609/aimag.v36i1.2564. URL https://onlinelibrary.wiley.com/doi/10.1609/aimag.v36i1.2564
-
[6]
Varvara Arzt and Allan Hanbury. Beyond the Numbers : Transparency in Relation Extraction Benchmark Creation and Leaderboards , November 2024. URL http://arxiv.org/abs/2411.05224
arXiv 2024
-
[7]
Experiences from using snowballing and database searches in systematic literature studies
Deepika Badampudi, Claes Wohlin, and Kai Petersen. Experiences from using snowballing and database searches in systematic literature studies. In Proceedings of the 19th International Conference on Evaluation and Assessment in Software Engineering, EASE '15, New York, NY, USA, 2015. Association for Computing Machinery. ISBN 9781450333504. doi:10.1145/27458...
arXiv 2015
-
[8]
Michelle Bao, Angela Zhou, Samantha Zottola, Brian Brubach, Sarah Desmarais, Aaron Horowitz, Kristian Lum, and Suresh Venkatasubramanian. It's COMPASlicated : The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks , April 2022. URL http://arxiv.org/abs/2106.05498
arXiv 2022
Show all 135 references
-
[9]
Malan, Jason H
Thomas Bartz-Beielstein, Carola Doerr, Daan van den Berg, Jakob Bossek, Sowmya Chandrasekaran, Tome Eftimov, Andreas Fischbach, Pascal Kerschke, William La Cava, Manuel Lopez-Ibanez, Katherine M. Malan, Jason H. Moore, Boris Naujoks, Patryk Orzechowski, Vanessa Volz, Markus Wa...
2020 arXiv
-
[10]
The Death of the Static AI Benchmark , March 2024
Sandi Besen. The Death of the Static AI Benchmark , March 2024. URL https://towardsdatascience.com/the-death-of-the-static-ai-benchmark-88b5ff437086
2024
-
[11]
Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick, Jason Phang, Aviya Skowron, Samson Tan, Xiangru Tang, Kevin A
Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Jeffrey Hsu, Mimansa Jaiswal...
2024 arXiv
-
[12]
AI auditing: The Broken Bus on the Road to AI Accountability , January 2024
Abeba Birhane, Ryan Steed, Victor Ojewale, Briana Vecchione, and Inioluwa Deborah Raji. AI auditing: The Broken Bus on the Road to AI Accountability , January 2024. URL http://arxiv.org/abs/2401.14462
2024 arXiv
-
[13]
Benchmark datasets driving artificial intelligence development fail to capture the needs of medical professionals
Kathrin Blagec, Jakob Kraiger, Wolfgang Frühwirt, and Matthias Samwald. Benchmark datasets driving artificial intelligence development fail to capture the needs of medical professionals. Journal of Biomedical Informatics, 137: 0 104274, January 2023. ISSN 15320464. doi:10.1016...
2023
-
[14]
Making Intelligence : Ethical Values in IQ and ML Benchmarks
Borhane Blili-Hamelin and Leif Hancox-Li. Making Intelligence : Ethical Values in IQ and ML Benchmarks . In 2023 ACM Conference on Fairness , Accountability , and Transparency , pages 271--284, Chicago IL USA, June 2023. ACM. ISBN 9798400701924. doi:10.1145/3593013.3593996. UR...
2023
-
[15]
Stereotyping Norwegian Salmon : An Inventory of Pitfalls in Fairness Benchmark Datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. Stereotyping Norwegian Salmon : An Inventory of Pitfalls in Fairness Benchmark Datasets . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th ...
2021 doi
-
[16]
Bowman and George Dahl
Samuel R. Bowman and George Dahl. What Will it Take to Fix Benchmarking in Natural Language Understanding ? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics : Human Language Technologies , pages 4843--4855, On...
2021 doi
-
[17]
Benchmarking, pages 363--368
Isabelle Bruno. Benchmarking, pages 363--368. Springer Netherlands, Dordrecht, 2014. ISBN 978-94-007-0753-5. doi:10.1007/978-94-007-0753-5_170. URL https://doi.org/10.1007/978-94-007-0753-5_170
2014 doi
-
[18]
Evaluating AI Evaluation : Perils and Prospects , July 2024
John Burden. Evaluating AI Evaluation : Perils and Prospects , July 2024. URL https://arxiv.org/abs/2407.09221v1
2024 arXiv
-
[19]
Ullman, Fernando Martinez-Plumed, Joshua B
Ryan Burnell, Wout Schellaert, John Burden, Tomer D. Ullman, Fernando Martinez-Plumed, Joshua B. Tenenbaum, Danaja Rutar, Lucy G. Cheke, Jascha Sohl-Dickstein, Melanie Mitchell, Douwe Kiela, Murray Shanahan, Ellen M. Voorhees, Anthony G. Cohn, Joel Z. Leibo, and Jose Hernandez...
2023 doi
-
[20]
Robert C. Camp. Benchmarking : The Search for Industry Best Practices That Lead to Superior Performance. Quality Press , the University of Michigan , 1989
1989
-
[22]
Yu, Qiang Yang, and Xing Xie
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. A Survey on Evaluation of Large Language Models , December 2023. URL http://arxiv.org/a...
2023 arXiv
-
[23]
Cheng, CS
Z. Cheng, CS. Pang, P. Wang, and et al. How to report and benchmark emerging field-effect transistors. Nature Electronics, 5: 0 416--423, 2020. URL https://doi.org/10.1038/s41928-022-00798-8
2020 doi
-
[24]
On the Measure of Intelligence , November 2019
François Chollet. On the Measure of Intelligence , November 2019. URL http://arxiv.org/abs/1911.01547
2019 arXiv
-
[25]
A survey of 25 years of evaluation
Kenneth Ward Church and Joel Hestness. A survey of 25 years of evaluation. Natural Language Engineering, 25 0 (06): 0 753--767, November 2019. ISSN 1351-3249, 1469-8110. doi:10.1017/S1351324919000275. URL https://www.cambridge.org/core/product/identifier/S1351324919000275/type...
2019 doi
-
[26]
Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals
Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals. The Benchmark Lottery , July 2021. URL http://arxiv.org/abs/2107.07002
2021 arXiv
-
[27]
On the genealogy of machine learning datasets: A critical history of ImageNet
Emily Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, and Hilary Nicole. On the genealogy of machine learning datasets: A critical history of ImageNet . Big Data & Society, 8 0 (2): 0 1--14, July 2021. ISSN 2053-9517, 2053-9517. doi:10.1177/20539517211035955. URL http://j...
2021 doi
-
[28]
Bringing the People Back In : Contesting Benchmark Machine Learning Datasets , July 2020
Remi Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, Hilary Nicole, and Morgan Klaus Scheuerman. Bringing the People Back In : Contesting Benchmark Machine Learning Datasets , July 2020. URL http://arxiv.org/abs/2007.07399
2020 arXiv
-
[30]
Agreements ‘in the wild’: Standards and alignment in machine learning benchmark dataset construction
Isak Engdahl. Agreements ‘in the wild’: Standards and alignment in machine learning benchmark dataset construction. Big Data & Society, 11 0 (2): 0 20539517241242457, June 2024. ISSN 2053-9517, 2053-9517. doi:10.1177/20539517241242457. URL https://journals.sagepub.com/doi/10.1...
2024 doi
-
[31]
Utility is in the Eye of the User : A Critique of NLP Leaderboards , March 2021
Kawin Ethayarajh and Dan Jurafsky. Utility is in the Eye of the User : A Critique of NLP Leaderboards , March 2021. URL http://arxiv.org/abs/2009.13888
2021 arXiv
-
[32]
European Union . Regulation (EU) 2022/2065 of the European Parliament and of the Council of 19 October 2022 on a Single Market For Digital Services and amending Directive 2000/31/EC (Digital Services Act) , 2022
2022
-
[33]
First Draft of the General-Purpose AI Code of Practice published, written by independent experts , 2024 a
European Union . First Draft of the General-Purpose AI Code of Practice published, written by independent experts , 2024 a
2024
-
[34]
Second Draft of the General-Purpose AI Code of Practice published, written by independent experts , 2024 b
European Union . Second Draft of the General-Purpose AI Code of Practice published, written by independent experts , 2024 b
2024
-
[35]
European Union . Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence and amending Regulations (Artificial Intelligence Act) , 2024 c
2024
-
[36]
Simon Frieder, Jonas Bayer, Katherine M. Collins, Julius Berner, Jacob Loader, András Juhász, Fabian Ruehle, Sean Welleck, Gabriel Poesia, Ryan-Rhys Griffiths, Adrian Weller, Anirudh Goyal, Thomas Lukasiewicz, and Timothy Gowers. Data for Mathematical Copilots : Better Ways of...
2024
-
[37]
Datasheets for Datasets , December 2021
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. Datasheets for Datasets , December 2021. URL http://arxiv.org/abs/1803.09010
2021 arXiv
-
[38]
Repairing the Cracked Foundation : A Survey of Obstacles in Evaluation Practices for Generated Text
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. Repairing the Cracked Foundation : A Survey of Obstacles in Evaluation Practices for Generated Text . Journal of Artificial Intelligence Research, 77: 0 103--166, May 2023. ISSN 1076-9757. doi:10.1613/jair.1.13715. URL ...
2023 doi
-
[39]
Wichmann
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. Shortcut Learning in Deep Neural Networks . Nature Machine Intelligence, 2 0 (11): 0 665--673, November 2020. ISSN 2522-5839. doi:10.1038/s42256-020...
2020 arXiv
-
[40]
Are We Done with MMLU ?, June 2024
Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile van Krieken, and Pasquale Minervini...
2024 arXiv
-
[41]
Diversity in artificial intelligence conferences
Emilia Gomez, Porcaro Lorenzo, Pedro Frau Amar, and Joao Vinagre. Diversity in artificial intelligence conferences. Publications Office of the European Union JRC137550, Publications Office, 2024. URL https://data.europa.eu/doi/10.2760/796551
2024 doi
-
[42]
Bowman, and Evan Hubinger
Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, Sören Mindermann, Ethan Perez, Linda Petrini, Jonathan Uesato, Jared Kaplan, Buck Shlegeris, Sam...
2024 arXiv
-
[44]
COMPL - AI Framework : A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act , October 2024
Philipp Guldimann, Alexander Spiridonov, Robin Staab, Nikola Jovanović, Mark Vero, Velko Vechev, Anna Gueorguieva, Mislav Balunović, Nikola Konstantinov, Pavol Bielik, Petar Tsankov, and Martin Vechev. COMPL - AI Framework : A Technical Interpretation and LLM Benchmarking Suit...
2024 arXiv
-
[45]
Ai hype is built on high test scores
Douglas Heaven. Ai hype is built on high test scores. those tests are flawed. Report, MIT Technology Review, 2023. URL https://www.technologyreview.com/2023/08/30/1078670/large-language-models-arent-people-lets-stop-testing-them-like-they-were/
2023
-
[46]
Measuring Massive Multitask Language Understanding , January 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring Massive Multitask Language Understanding , January 2021. URL http://arxiv.org/abs/2009.03300
2021 arXiv
-
[47]
J.L. Henning. Spec cpu2000: measuring cpu performance in the new millennium. Computer, 33 0 (7): 0 28--35, 2000. doi:10.1109/2.869367
2000 doi
-
[48]
On the Limitations of Compute Thresholds as a Governance Strategy , July 2024
Sara Hooker. On the Limitations of Compute Thresholds as a Governance Strategy , July 2024. URL http://arxiv.org/abs/2407.05694
2024 arXiv
-
[49]
Evaluation Gaps in Machine Learning Practice
Ben Hutchinson, Negar Rostamzadeh, Christina Greer, Katherine Heller, and Vinodkumar Prabhakaran. Evaluation Gaps in Machine Learning Practice . In 2022 ACM Conference on Fairness , Accountability , and Transparency , pages 1859--1876, Seoul Republic of Korea, June 2022. ACM. ...
2022
-
[50]
Systematic literature studies: database searches vs
Samireh Jalali and Claes Wohlin. Systematic literature studies: database searches vs. backward snowballing. In Proceedings of the ACM-IEEE International Symposium on Empirical Software Engineering and Measurement, ESEM '12, page 29–38, New York, NY, USA, 2012. Association for ...
2012
-
[51]
Escaping the McNamara Fallacy : Toward More Impactful Recommender Systems Research
Dietmar Jannach and Christine Bauer. Escaping the McNamara Fallacy : Toward More Impactful Recommender Systems Research . AI Magazine, 41 0 (4): 0 79--95, December 2020. ISSN 0738-4602, 2371-9621. doi:10.1609/aimag.v41i4.5312. URL https://onlinelibrary.wiley.com/doi/10.1609/ai...
2020 doi
-
[52]
The Constitution of Algorithms : Ground - Truthing , Programming , Formulating
Florian Jaton. The Constitution of Algorithms : Ground - Truthing , Programming , Formulating . Inside Technology . The MIT Press, Cambridge, 2021. ISBN 978-0-262-54214-2 978-0-262-36323-5
2021
-
[53]
Under the radar? examining the evaluation of foundation models
Elliot Jones, Mahi Hardalupas, and William Agrew. Under the radar? examining the evaluation of foundation models. Report, Ada Lovelace Institute, 2024. URL https://www.adalovelaceinstitute.org/report/under-the-radar/
2024
-
[54]
Ground truth tracings ( GTT ): On the epistemic limits of machine learning
Edward B Kang. Ground truth tracings ( GTT ): On the epistemic limits of machine learning. Big Data & Society, 10 0 (1): 0 20539517221146122, January 2023. ISSN 2053-9517, 2053-9517. doi:10.1177/20539517221146122. URL https://journals.sagepub.com/doi/10.1177/20539517221146122
2023 doi
-
[55]
Siegel, Nitya Nadgir, and Arvind Narayanan
Sayash Kapoor, Benedikt Stroebl, Zachary S. Siegel, Nitya Nadgir, and Arvind Narayanan. AI Agents That Matter , July 2024. URL http://arxiv.org/abs/2407.01502
2024 arXiv
-
[56]
Leakage in data mining: Formulation, detection, and avoidance
Shachar Kaufman, Saharon Rosset, Claudia Perlich, and Ori Stitelman. Leakage in data mining: Formulation, detection, and avoidance. ACM Trans. Knowl. Discov. Data, 6 0 (4), December 2012. ISSN 1556-4681. doi:10.1145/2382577.2382579. URL https://doi.org/10.1145/2382577.2382579
2012
-
[57]
Everyone Is Judging AI by These Tests
Jon Keegan. Everyone Is Judging AI by These Tests . But Experts Say They ’re Close to Meaningless . The Markup, July 2024. URL https://themarkup.org/artificial-intelligence/2024/07/17/everyone-is-judging-ai-by-these-tests-but-experts-say-theyre-close-to-meaningless
2024
-
[58]
Mulvehill, and Deborah L
Mayank Kejriwal, Henrique Santos, Ke Shen, Alice M. Mulvehill, and Deborah L. McGuinness. A noise audit of human-labeled benchmarks for machine commonsense reasoning. Scientific Reports, 14 0 (1): 0 8609, April 2024. ISSN 2045-2322. doi:10.1038/s41598-024-58937-4. URL https://...
2024 doi
-
[59]
Feeling fixes: Mess and emotion in algorithmic audits
Os Keyes and Jeanie Austin. Feeling fixes: Mess and emotion in algorithmic audits. Big Data & Society, 9 0 (2): 0 20539517221113772, July 2022. ISSN 2053-9517, 2053-9517. doi:10.1177/20539517221113772. URL https://journals.sagepub.com/doi/10.1177/20539517221113772
2022 doi
-
[60]
Bernard Koch, Emily Denton, Alex Hanna, and Jacob G. Foster. Reduced, Reused and Recycled : The Life of a Dataset in Machine Learning Research , December 2021. URL http://arxiv.org/abs/2112.01716
2021 arXiv
-
[61]
Koch and David Peterson
Bernard J. Koch and David Peterson. From Protoscience to Epistemic Monoculture : How Benchmarking Set the Stage for the Deep Learning Revolution , April 2024. URL http://arxiv.org/abs/2404.06647
2024 arXiv
-
[62]
Metaethical Perspectives on ' Benchmarking ' AI Ethics , April 2022
Travis LaCroix and Alexandra Sasha Luccioni. Metaethical Perspectives on ' Benchmarking ' AI Ethics , April 2022. URL http://arxiv.org/abs/2204.05151
2022 arXiv
-
[63]
Vazquez, Niclas Kupper, Misha Yagudin, and Laurence Aitchison
Gavin Leech, Juan J. Vazquez, Niclas Kupper, Misha Yagudin, and Laurence Aitchison. Questionable practices in machine learning, July 2024. URL https://arxiv.org/abs/2407.12220v2
2024 arXiv
-
[64]
Question and answer test-train overlap in open-domain question answering datasets
Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel. Question and answer test-train overlap in open-domain question answering datasets. In Paola Merlo, Jorg Tiedemann, and Reut Tsarfaty, editors, Proceedings of the 16th Conference of the European Chapter of the Association f...
2021 doi
-
[65]
Manning, Christopher Ré, Diana Acosta-Navas, Drew A
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas...
2023 arXiv
-
[66]
Vera Liao and Ziang Xiao
Q. Vera Liao and Ziang Xiao. Rethinking Model Evaluation as Narrowing the Socio - Technical Gap , June 2023. URL http://arxiv.org/abs/2306.03100
2023 arXiv
-
[67]
Are we learning yet? a meta review of evaluation failures across machine learning
Thomas Liao, Rohan Taori, Inioluwa Deborah Raji, and Ludwig Schmidt. Are we learning yet? a meta review of evaluation failures across machine learning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https:...
2021
-
[68]
ExplainaBoard : An Explainable Leaderboard for NLP , July 2021
Pengfei Liu, Jinlan Fu, Yang Xiao, Weizhe Yuan, Shuaicheng Chang, Junqi Dai, Yixin Liu, Zihuiwen Ye, Zi-Yi Dou, and Graham Neubig. ExplainaBoard : An Explainable Leaderboard for NLP , July 2021. URL http://arxiv.org/abs/2104.06387
2021 arXiv
-
[69]
AI competitions as infrastructures of power in medical imaging
Dieuwertje Luitse, Tobias Blanke, and Thomas Poell. AI competitions as infrastructures of power in medical imaging. Information, Communication & Society, pages 1--22, March 2024. ISSN 1369-118X, 1468-4462. doi:10.1080/1369118X.2024.2334393. URL https://www.tandfonline.com/doi/...
2024
-
[70]
Bias in Language Models : Beyond Trick Tests and Toward RUTEd Evaluation , February 2024
Kristian Lum, Jacy Reese Anthis, Chirag Nagpal, and Alexander D'Amour. Bias in Language Models : Beyond Trick Tests and Toward RUTEd Evaluation , February 2024. URL https://arxiv.org/abs/2402.12649v1
2024 arXiv
-
[71]
Data contamination: From memorization to exploitation
Inbal Magar and Roy Schwartz. Data contamination: From memorization to exploitation. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 15...
2022 doi
-
[72]
Practices of Benchmarking : Vulnerability in the Computer Vision Pipeline
Nicolas Malevé. Practices of Benchmarking : Vulnerability in the Computer Vision Pipeline . photographies, 16 0 (2): 0 173--189, May 2023. ISSN 1754-0763, 1754-0771. doi:10.1080/17540763.2023.2189159. URL https://www.tandfonline.com/doi/full/10.1080/17540763.2023.2189159
2023
-
[73]
Put to the test: For a new sociology of testing
Noortje Marres and David Stark. Put to the test: For a new sociology of testing. The British Journal of Sociology, 71 0 (3): 0 423--443, June 2020. ISSN 0007-1315, 1468-4446. doi:10.1111/1468-4446.12746. URL https://onlinelibrary.wiley.com/doi/10.1111/1468-4446.12746
2020
-
[74]
Mlperf training benchmark
Peter Mattson, Christine Cheng, Gregory Diamos, Cody Coleman, Paulius Micikevicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, David Brooks, Dehao Chen, Debo Dutta, Udit Gupta, Kim Hazelwood, Andy Hock, Xinyuan Huang, Daniel Kang, David Kanter, Na...
2020
-
[75]
Mlperf: An industry standard benchmark suite for machine learning performance
Peter Mattson, Vijay Janapa Reddi, Christine Cheng, Cody Coleman, Greg Diamos, David Kanter, Paulius Micikevicius, David Patterson, Guenther Schmuelling, Hanlin Tang, Gu-Yeon Wei, and Carole-Jean Wu. Mlperf: An industry standard benchmark suite for machine learning performance...
2020
-
[76]
McIntosh, Teo Susnjak, Nalin Arachchilage, Tong Liu, Paul Watters, and Malka N
Timothy R. McIntosh, Teo Susnjak, Nalin Arachchilage, Tong Liu, Paul Watters, and Malka N. Halgamuge. Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence , October 2024. URL http://arxiv.org/abs/2402.09880
2024 arXiv
-
[77]
Frontier Models are Capable of In -context Scheming , December 2024
Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn. Frontier Models are Capable of In -context Scheming , December 2024. URL https://arxiv.org/abs/2412.04984v1
2024 arXiv
-
[78]
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, and Samuel R. Bowman. What Do NLP Researchers Believe ? Results of the NLP Community Metasurvey , August 2022. URL http://arx...
2022 arXiv
-
[79]
Benchmarking the Benchmarks
Marc Miltenberger, Steven Arzt, Philipp Holzinger, and Julius Näumann. Benchmarking the Benchmarks . In Proceedings of the ACM Asia Conference on Computer and Communications Security , pages 387--400, Melbourne VIC Australia, July 2023. ACM. ISBN 9798400700989. doi:10.1145/357...
2023
-
[80]
How do we know how smart AI systems are? Science, 381 0 (6654), July 2023
Margaret Mitchell. How do we know how smart AI systems are? Science, 381 0 (6654), July 2023. doi:10.1126/science.adj595. URL https://www.science.org/doi/10.1126/science.adj5957
2023 doi
-
[82]
State of What Art ? A Call for Multi - Prompt LLM Evaluation , May 2024
Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. State of What Art ? A Call for Multi - Prompt LLM Evaluation , May 2024. URL http://arxiv.org/abs/2401.00595
2024 arXiv
-
[83]
Proxies: The Cultural Work of Standing In
Dylan Mulvin. Proxies: The Cultural Work of Standing In . Infrastructures. The MIT Press, Cambridge, 2021. ISBN 978-0-262-04514-8 978-0-262-36624-3
2021
-
[84]
GPT -4 and professional benchmarks: the wrong answer to the wrong question, March 2023 a
Arvind Narayanan and Sayash Kapoor. GPT -4 and professional benchmarks: the wrong answer to the wrong question, March 2023 a . URL https://www.aisnakeoil.com/p/gpt-4-and-professional-benchmarks
2023
-
[85]
Evaluating LLMs is a minefield, 2023 b
Arvind Narayanan and Sayash Kapoor. Evaluating LLMs is a minefield, 2023 b . URL https://www.cs.princeton.edu/ arvindn/talks/evaluating_llms_minefield/
2023
-
[86]
Feder Cooper, Daphne Ippolito, Christopher A
Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, and Katherine Lee. Extracting Training Data from ChatGPT , November 2023 a . URL https://not-just-memorization.github.io/extracting-...
2023
-
[87]
Feder Cooper, Daphne Ippolito, Christopher A
Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee. Scalable Extraction of Training Data from ( Production ) Language Models , November 2023 b . URL ...
2023 arXiv
-
[88]
Evaluating Hosting Provider Security Through Abuse Data and the Creation of Metrics
Arman Noroozian. Evaluating Hosting Provider Security Through Abuse Data and the Creation of Metrics . Dissertation ( TU Delft ), 2020. ISBN: 9789065624451
2020
-
[89]
Hidden Stratification Causes Clinically Meaningful Failures in Machine Learning for Medical Imaging , November 2019
Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro, and Christopher Ré. Hidden Stratification Causes Clinically Meaningful Failures in Machine Learning for Medical Imaging , November 2019. URL http://arxiv.org/abs/1909.12475
2019 arXiv
-
[90]
Towards AI Accountability Infrastructure : Gaps and Opportunities in AI Audit Tooling , March 2024
Victor Ojewale, Ryan Steed, Briana Vecchione, Abeba Birhane, and Inioluwa Deborah Raji. Towards AI Accountability Infrastructure : Gaps and Opportunities in AI Audit Tooling , March 2024. URL http://arxiv.org/abs/2402.17861
2024 arXiv
-
[91]
The social construction of datasets: On the practices, processes, and challenges of dataset creation for machine learning
Will Orr and Kate Crawford. The social construction of datasets: On the practices, processes, and challenges of dataset creation for machine learning. New Media & Society, 26 0 (9): 0 4955--4972, 2024 a
2024
-
[92]
Building Better Datasets : Seven Recommendations for Responsible Design from Dataset Creators , August 2024 b
Will Orr and Kate Crawford. Building Better Datasets : Seven Recommendations for Responsible Design from Dataset Creators , August 2024 b . URL http://arxiv.org/abs/2409.00252
2024 arXiv
-
[93]
Will Orr and Edward B. Kang. AI as a Sport : On the Competitive Epistemologies of Benchmarking . In The 2024 ACM Conference on Fairness , Accountability , and Transparency , pages 1875--1884, Rio de Janeiro Brazil, June 2024. ACM. ISBN 9798400704505. doi:10.1145/3630106.365901...
2024
-
[94]
Mapping global dynamics of benchmark creation and saturation in artificial intelligence
Simon Ott, Adriano Barbosa-Silva, Kathrin Blagec, Jan Brauner, and Matthias Samwald. Mapping global dynamics of benchmark creation and saturation in artificial intelligence. Nature Communications, 13 0 (1): 0 6793, November 2022. ISSN 2041-1723. doi:10.1038/s41467-022-34591-0....
2022 doi
-
[95]
Benchmark
Oxford English Dictionary . Benchmark. meaning and use, 2017. URL https://www.oed.com/dictionary/benchmark_n?tab=meaning_and_use&tl=true
2017
-
[96]
Cheke, and José Hernández-Orallo
Lorenzo Pacchiardi, Marko Tesic, Lucy G. Cheke, and José Hernández-Orallo. Leaving the barn door open for Clever Hans : Simple features predict LLM benchmark answers, October 2024. URL http://arxiv.org/abs/2410.11672
2024 arXiv
-
[97]
Raison d’être of the benchmark dataset: A Survey of Current Practices of Benchmark Dataset Sharing Platforms
Jaihyun Park and Sullam Jeoung. Raison d’être of the benchmark dataset: A Survey of Current Practices of Benchmark Dataset Sharing Platforms . In Proceedings of NLP Power ! The First Workshop on Efficient Benchmarking in NLP , pages 1--10, Dublin, Ireland, 2022. Association fo...
2022 doi
-
[98]
Bender, Emily Denton, and Alex Hanna
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M. Bender, Emily Denton, and Alex Hanna. Data and its (dis)contents: A survey of dataset development and use in machine learning research. Patterns, 2 0 (11): 0 100336, November 2021. ISSN 26663899. doi:10.1016/j.patter.2021.1...
2021
-
[99]
Understanding and Benchmarking Artificial Intelligence : OpenAI 's o3 Is Not AGI , January 2025
Rolf Pfister and Hansueli Jud. Understanding and Benchmarking Artificial Intelligence : OpenAI 's o3 Is Not AGI , January 2025. URL http://arxiv.org/abs/2501.07458
2025 arXiv
-
[100]
Testing - One , Two , Three ... Testing !
Trevor Pinch. " Testing - One , Two , Three ... Testing !": Toward a Sociology of Testing . Science, Technology, & Human Values, 18 0 (1): 0 25--41, January 1993. ISSN 0162-2439, 1552-8251. doi:10.1177/016224399301800103. URL https://journals.sagepub.com/doi/10.1177/016224399301800103
1993 doi
-
[101]
The Roles of English in Evaluating Multilingual Language Models , December 2024
Wessel Poelman and Miryam de Lhoneux. The Roles of English in Evaluating Multilingual Language Models , December 2024. URL http://arxiv.org/abs/2412.08392
2024 arXiv
-
[102]
Fine-tuning Aligned Language Models Compromises Safety , Even When Users Do Not Intend To !, October 2023
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. Fine-tuning Aligned Language Models Compromises Safety , Even When Users Do Not Intend To !, October 2023. URL http://arxiv.org/abs/2310.03693
2023 arXiv
-
[103]
Bender, Alex Hanna, and Amandalynne Paullada
Inioluwa Deborah Raji, Emily Denton, Emily M. Bender, Alex Hanna, and Amandalynne Paullada. AI and the everything in the whole wide world benchmark. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https://o...
2021
-
[104]
Gaps in the Safety Evaluation of Generative AI
Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Ramona Comanescu, Canfer Akbulut, Tom Stepleton, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, William Isaac, and Laura Weidinger. Gaps in the Safet...
2024 doi
-
[105]
Safetywashing: Do AI safety benchmarks actually measure safety progress? In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024
Richard Ren, Steven Basart, Adam Khoja, Alice Gatti, Long Phan, Xuwang Yin, Mantas Mazeika, Alexander Pan, Gabriel Mukobi, Ryan Hwang Kim, Stephen Fitz, and Dan Hendrycks. Safetywashing: Do AI safety benchmarks actually measure safety progress? In The Thirty-eight Conference o...
2024
-
[106]
Betterbench: Assessing AI benchmarks, uncovering issues, and establishing best practices
Anka Reuel, Amelia Hardy, Chandler Smith, Max Lamparth, Malcolm Hardy, and Mykel Kochenderfer. Betterbench: Assessing AI benchmarks, uncovering issues, and establishing best practices. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchma...
2024
-
[107]
Data Contamination Through the Lens of Time , October 2023
Manley Roberts, Himanshu Thakur, Christine Herlihy, Colin White, and Samuel Dooley. Data Contamination Through the Lens of Time , October 2023. URL http://arxiv.org/abs/2310.10628
2023 arXiv
-
[108]
Lalor, Robin Jia, and Jordan Boyd-Graber
Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, and Jordan Boyd-Graber. Evaluation Examples are not Equally Informative : How should that change NLP Leaderboards ? In Proceedings of the 59th Annual Meeting of the Association for Computational L...
2021 doi
-
[109]
Kevin Roose. A.i. has a measurement problem. Report, New York Times, 2024. URL https://www.nytimes.com/2024/04/15/technology/ai-models-measurement.html
2024
-
[110]
SafetyPrompts : a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety , April 2024
Paul Röttger, Fabio Pernisi, Bertie Vidgen, and Dirk Hovy. SafetyPrompts : a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety , April 2024. URL http://arxiv.org/abs/2404.05399
2024 arXiv
-
[111]
Everyone wants to do the model work, not the data work
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. “ Everyone wants to do the model work, not the data work”: Data Cascades in High - Stakes AI . In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems...
2021
-
[112]
Do Datasets Have Politics ? Disciplinary Values in Computer Vision Dataset Development
Morgan Klaus Scheuerman, Alex Hanna, and Emily Denton. Do Datasets Have Politics ? Disciplinary Values in Computer Vision Dataset Development . Proceedings of the ACM on Human-Computer Interaction, 5 0 (CSCW2): 0 1--37, October 2021. ISSN 2573-0142. doi:10.1145/3476058. URL ht...
2021 doi
-
[113]
Targeting the Benchmark : On Methodology in Current Natural Language Processing Research , July 2020
David Schlangen. Targeting the Benchmark : On Methodology in Current Natural Language Processing Research , July 2020. URL http://arxiv.org/abs/2007.04792
2020 arXiv
-
[114]
Winner's Curse ? On Pace , Progress , and Empirical Rigor
D Sculley, Jasper Snoek, Ali Rahimi, and Alex Wiltschko. Winner's Curse ? On Pace , Progress , and Empirical Rigor . Vancouver, BC, Canada, 2018. URL https://openreview.net/pdf?id=rJWF0Fywf
2018
-
[115]
Selbst, Danah Boyd, Sorelle A
Andrew D. Selbst, Danah Boyd, Sorelle A. Friedler, Suresh Venkatasubramanian, and Janet Vertesi. Fairness and Abstraction in Sociotechnical Systems . In Proceedings of the Conference on Fairness , Accountability , and Transparency , pages 59--68, Atlanta GA USA, January 2019. ...
2019
-
[116]
Arafat
Shilad Sen, Margaret E. Giesel, Rebecca Gold, Benjamin Hillmann, Matt Lesicko, Samuel Naden, Jesse Russell, Zixiao (Ken) Wang, and Brent Hecht. Turkers, Scholars , " Arafat " and " Peace ": Cultural Communities and Algorithmic Gold Standards . In Proceedings of the 18th ACM Co...
2015
-
[117]
Lazy Data Practices Harm Fairness Research
Jan Simson, Alessandro Fabris, and Christoph Kern. Lazy Data Practices Harm Fairness Research . In The 2024 ACM Conference on Fairness , Accountability , and Transparency , pages 642--659, Rio de Janeiro Brazil, June 2024. ACM. ISBN 9798400704505. doi:10.1145/3630106.3658931. ...
2024
-
[118]
Smith, Saleema Amershi, Solon Barocas, Hanna Wallach, and Jennifer Wortman Vaughan
Jessie J. Smith, Saleema Amershi, Solon Barocas, Hanna Wallach, and Jennifer Wortman Vaughan. REAL ML : Recognizing , Exploring , and Articulating Limitations of Machine Learning Research . In 2022 ACM Conference on Fairness , Accountability , and Transparency , pages 587--597...
2022
-
[119]
Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, et al
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, et al. Beyond the Imitati...
2023 arXiv
-
[120]
Another science is possible: a manifesto for slow science
Isabelle Stengers. Another science is possible: a manifesto for slow science. Polity press, Cambridge, 2018. ISBN 978-1-5095-2180-7
2018
-
[121]
The olympics of ai: Benchmarking machine learning systems, 2023
Matthew Stewart. The olympics of ai: Benchmarking machine learning systems, 2023. URL https://towardsdatascience.com/the-olympics-of-ai-benchmarking-machine-learning-systems-c4b2051fbd2b
2023
-
[122]
‘ Improving ratings’: audit in the British University system
Marilyn Strathern. ‘ Improving ratings’: audit in the British University system. European Review, 5 0 (3): 0 305--321, July 1997. ISSN 10627987, 1234981X. doi:10.1002/(SICI)1234-981X(199707)5:3<305::AID-EURO184>3.0.CO;2-4. URL https://www.cambridge.org/core/product/identifier/...
1997 doi
-
[123]
It Takes Two to Tango : Navigating Conceptualizations of NLP Tasks and Measurements of Performance , May 2023
Arjun Subramonian, Xingdi Yuan, Hal Daumé III, and Su Lin Blodgett. It Takes Two to Tango : Navigating Conceptualizations of NLP Tasks and Measurements of Performance , May 2023. URL http://arxiv.org/abs/2305.09022
2023 arXiv
-
[124]
BEIR : A heterogeneous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas R \"u ckl \'e , Abhishek Srivastava, and Iryna Gurevych. BEIR : A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks ...
2021
-
[125]
Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence , 2023
The White House . Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence , 2023
2023
-
[126]
Politics of data reuse in machine learning systems: Theorizing reuse entanglements
Nanna Bonde Thylstrup, Kristian Bondo Hansen, Mikkel Flyverbom, and Louise Amoore. Politics of data reuse in machine learning systems: Theorizing reuse entanglements. Big Data & Society, 9 0 (2): 0 20539517221139785, July 2022. ISSN 2053-9517, 2053-9517. doi:10.1177/2053951722...
2022 doi
-
[127]
Memorization Without Overfitting : Analyzing the Training Dynamics of Large Language Models
Kushal Tirumala, Aram Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. Memorization Without Overfitting : Analyzing the Training Dynamics of Large Language Models . In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information ...
2022
-
[128]
From ImageNet to Image Classification : Contextualizing Progress on Benchmarks , May 2020
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas, and Aleksander Madry. From ImageNet to Image Classification : Contextualizing Progress on Benchmarks , May 2020. URL http://arxiv.org/abs/2005.11295
2020 arXiv
-
[129]
Online Safety Act 2023 , 2023
UK Parliament . Online Safety Act 2023 , 2023
2023
-
[130]
Framework for Artificial Intelligence Diffusion , 2025
US Department of Commerce . Framework for Artificial Intelligence Diffusion , 2025
2025
-
[131]
Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan
Keyon Vafa, Justin Y. Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan. Evaluating the World Model Implicit in a Generative Model , November 2024. URL http://arxiv.org/abs/2406.03689
2024 arXiv
-
[132]
Sociotechnical Safety Evaluation of Generative AI Systems , October 2023
Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, and William Isaac. Sociotechnical Safety Evaluation of Generative AI Systems , Octobe...
2023 arXiv
-
[133]
Brown, and Francis Rhys Ward
Teun van der Weij, Felix Hofstätter, Ollie Jaffe, Samuel F. Brown, and Francis Rhys Ward. AI Sandbagging : Language Models can Strategically Underperform on Evaluations , June 2024. URL http://arxiv.org/abs/2406.07358
2024 arXiv
-
[134]
Benchmark Data Contamination of Large Language Models : A Survey , June 2024 a
Cheng Xu, Shuhao Guan, Derek Greene, and M.-Tahar Kechadi. Benchmark Data Contamination of Large Language Models : A Survey , June 2024 a . URL http://arxiv.org/abs/2406.04244
2024 arXiv
-
[135]
Benchmarking Benchmark Leakage in Large Language Models , April 2024 b
Ruijie Xu, Zengzhi Wang, Run-Ze Fan, and Pengfei Liu. Benchmarking Benchmark Leakage in Large Language Models , April 2024 b . URL http://arxiv.org/abs/2404.18824
2024 arXiv
-
[136]
Gonzalez, and Ion Stoica
Shuo Yang, Wei-Lin Chiang, Lianmin Zheng, Joseph E. Gonzalez, and Ion Stoica. Rethinking Benchmark and Contamination for Language Models with Rephrased Samples , November 2023. URL http://arxiv.org/abs/2311.04850
2023 arXiv
-
[137]
Revisiting Out -of-distribution Robustness in NLP : Benchmark , Analysis , and LLMs Evaluations
Lifan Yuan, Yangyi Chen, Ganqu Cui, Hongcheng Gao, Fangyuan Zou, Xingyi Cheng, Heng Ji, Zhiyuan Liu, and Maosong Sun. Revisiting Out -of-distribution Robustness in NLP : Benchmark , Analysis , and LLMs Evaluations . 37th Conference on Neural Information Processing Systems (Neu...
2023
-
[138]
Zhang, Kevin Klyman, Yifan Mai, Yoav Levine, Yian Zhang, Rishi Bommasani, and Percy Liang
Andy K. Zhang, Kevin Klyman, Yifan Mai, Yoav Levine, Yian Zhang, Rishi Bommasani, and Percy Liang. Language model developers should report train-test overlap, October 2024. URL http://arxiv.org/abs/2410.08385
2024 arXiv
-
[139]
Top llms in china and the u.s
Lin Zhijia. Top llms in china and the u.s. only 5 months apart: Kai-fu lee, 2024. URL https://en.tmtpost.com/post/7289212
2024
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.