REVIEW 12 cited by
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Quantitative Artificial Intelligence (AI) Benchmarks have emerged as fundamental tools for evaluating the performance, capability, and safety of AI models and systems. Currently, they shape the direction of AI development and are playing an increasingly prominent role in regulatory frameworks. As their influence grows, however, so too does concerns about how and with what effects they evaluate highly sensitive topics such as capabilities, including high-impact capabilities, safety and systemic risks. This paper presents an interdisciplinary meta-review of about 100 studies that discuss shortcomings in quantitative benchmarking practices, published in the last 10 years. It brings together many fine-grained issues in the design and application of benchmarks (such as biases in dataset creation, inadequate documentation, data contamination, and failures to distinguish signal from noise) with broader sociotechnical issues (such as an over-focus on evaluating text-based AI models according to one-time testing logic that fails to account for how AI models are increasingly multimodal and interact with humans and other technical systems). Our review also highlights a series of systemic flaws in current benchmarking practices, such as misaligned incentives, construct validity issues, unknown unknowns, and problems with the gaming of benchmark results. Furthermore, it underscores how benchmark practices are fundamentally shaped by cultural, commercial and competitive dynamics that often prioritise state-of-the-art performance at the expense of broader societal concerns. By providing an overview of risks associated with existing benchmarking procedures, we problematise disproportionate trust placed in benchmarks and contribute to ongoing efforts to improve the accountability and relevance of quantitative AI benchmarks within the complexities of real-world scenarios.
Forward citations
Cited by 12 Pith papers
-
Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems
MAAC defines a nine-dimension framework for process-oriented cognitive evaluation of text-based AI, shifting assessment from outputs to underlying reasoning.
-
The Foreign Policy AI Evaluation Gap
Public technical AI governance almost never evaluates real foreign-policy AI workflows; the paper maps that gap and proposes task-scoped, human-recombined evaluation instead of model leaderboards.
-
What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles
TurtleSoup-Bench is a new interactive benchmark showing that LLMs struggle with imaginative reasoning compared to humans.
-
Deprecating Benchmarks: Criteria and Framework
A framework for deprecating outdated or flawed AI benchmarks, with seven criteria and a three-phase process of assessment, reporting, and notification.
-
Lilith: Developmental Modular LLMs with Chemical Signaling
A conceptual framework in which untrained modular LLMs are developed through simulated life and token-based chemical signaling, with the goal of enabling empirical study of consciousness emergence via Integrated Infor...
-
Establishing Best Practices for Building Rigorous Agentic Benchmarks
Agentic benchmarks frequently mis-grade agents, and the new ABC checklist helps identify and correct such errors in ten popular benchmarks.
-
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.
-
VLM@school -- Evaluation of AI image understanding on German middle school knowledge
A new German middle school visual question-answering benchmark shows open-weight VLMs score below 45% overall, with especially weak results in music, math, and adversarial questions.
-
Position: Explanation Stability Is a Property of the Model Method Pair, Not the Model
On chest X-rays, LayerCAM ranks InceptionV3 as the most stable model but Grad-CAM++ ranks DenseNet201 first, showing explanation stability is a property of the model-method pair.
-
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
VERA-MH, an automated safety benchmark for suicide-risk chatbot conversations, agreed with clinician ratings (IRR 0.81), though the clinical reference was not fully independent.
-
A Conceptual Framework for AI Capability Evaluations
A descriptive conceptual framework with seven elements (target, task, subject, inputs, instance, measurement, result analysis) for systematizing analysis of AI capability evaluations.
-
Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation
The paper is a literature review that classifies privacy-preserving AI techniques in dataspaces using a qualitative taxonomy of privacy, performance, and compliance ratings.
Discussion (0). Sign in to comment.