REVIEW 18 cited by
AI and the Everything in the Whole Wide World Benchmark
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
There is a tendency across different subfields in AI to valorize a small collection of influential benchmarks. These benchmarks operate as stand-ins for a range of anointed common problems that are frequently framed as foundational milestones on the path towards flexible and generalizable AI systems. State-of-the-art performance on these benchmarks is widely understood as indicative of progress towards these long-term goals. In this position paper, we explore the limits of such benchmarks in order to reveal the construct validity issues in their framing as the functionally "general" broad measures of progress they are set up to be.
Forward citations
Cited by 18 Pith papers
-
On the Fitness Landscape in the $NK$ Model
For the NK fitness landscape with K/N tending to alpha, exact limits for free energy and maximum fitness are identified, together with the geometry of near-fittest peaks.
-
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
In a randomized trial of 246 real open-source tasks, experienced developers took 19% longer when AI tools were allowed, despite forecasting 24% faster completion.
-
Localized Adaptation Reveals Distinct Learning Signatures in Transformers
Adaptation site in transformers (early/middle/late layers) systematically changes acquisition, transfer, and boundedness, with distinct profiles across five learning objectives.
-
Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing
Safety governance designed for narrow AI presupposes bounded tasks; general-purpose AI's unbounded, fluent outputs break those assumptions, so policing frameworks risk failing unless rebuilt around GPAI-specific evidence.
-
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.
-
The Benchmark Ceiling: Human Judgment, Evaluation Scarcity, and the Political Economy of AI Capability Measurement
Frontier AI benchmarks lose discriminatory power as models saturate easy items, concentrating valid signal in scarce expert-authored hard-tail items whose replacement cost rises convexly with capability.
-
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
A practical evaluation protocol for AI pentesting agents that uses validated vulnerability discovery, LLM semantic matching, and bipartite scoring to assess performance in realistic, complex targets.
-
Intuitions of Machine Learning Researchers about Transfer Learning for Medical Image Classification
Source-dataset selection for medical transfer learning is driven by community practice and perceived similarity, and 'more similar is better' does not consistently hold.
-
Countering Privacy Nihilism
Privacy nihilism, the claim that AI's inferential power makes data categories useless, is unjustified because many AI inference claims rest on conceptually overfitted models.
-
Deprecating Benchmarks: Criteria and Framework
A framework for deprecating outdated or flawed AI benchmarks, with seven criteria and a three-phase process of assessment, reporting, and notification.
-
Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks
Data scientists construct prediction targets through bricolage, applying five reformulation strategies (piggybacking, composing, swapping, bridging, refining) to balance five criteria: validity, simplicity, predictabi...
-
Potemkin Understanding in Large Language Models
LLMs frequently pass definition questions yet fail to use the same concepts in classification, generation, and editing tasks, a gap the authors call potemkin understanding.
-
VLM@school -- Evaluation of AI image understanding on German middle school knowledge
A new German middle school visual question-answering benchmark shows open-weight VLMs score below 45% overall, with especially weak results in music, math, and adversarial questions.
-
Evidence-Grounded Constraint Checking in Construction Documents
A controlled four-image comparison shows region-focused crops beat page overviews on six projects but lose on 23 projects, indicating a resolution–breadth tradeoff rather than a dominant strategy.
-
Red Teaming AI Policy: A Taxonomy of Avoision and the EU AI Act
A taxonomy of avoision under the EU AI Act, with strategies to escape scope, exploit exemptions, and manipulate risk or operator categories.
-
Multi-modal video data-pipelines for machine learning with minimal human supervision
An open-source video pipeline automatically extracts 13+ visual modalities from raw video with no human annotation, and a sub-1M-parameter distilled model reaches near-Mask2Former accuracy on an aerial scene benchmark.
-
Private, Verifiable, and Auditable AI Systems
A thesis demonstrating partial prototypes for zk-verifiable model evaluation and privacy-preserving retrieval, and arguing these pieces can compose into end-to-end auditable AI systems.
-
Cognitive Castes: Artificial Intelligence, Epistemic Stratification, and the Dissolution of Democratic Discourse
The paper claims AI entrenches cognitive castes by amplifying the reasoning of the trained while pacifying everyone else, thereby eroding deliberative democracy.
Discussion (0). Sign in to comment.