REVIEW 20 cited by
Do Large Language Model Benchmarks Test Reliability?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Do Large Language Model Benchmarks Test Reliability?
read the original abstract
When deploying large language models (LLMs), it is important to ensure that these models are not only capable, but also reliable. Many benchmarks have been created to track LLMs' growing capabilities, however there has been no similar focus on measuring their reliability. To understand the potential ramifications of this gap, we investigate how well current benchmarks quantify model reliability. We find that pervasive label errors can compromise these evaluations, obscuring lingering model failures and hiding unreliable behavior. Motivated by this gap in the evaluation of reliability, we then propose the concept of so-called platinum benchmarks, i.e., benchmarks carefully curated to minimize label errors and ambiguity. As a first attempt at constructing such benchmarks, we revise examples from fifteen existing popular benchmarks. We evaluate a wide range of models on these platinum benchmarks and find that, indeed, frontier LLMs still exhibit failures on simple tasks such as elementary-level math word problems. Analyzing these failures further reveals previously unidentified patterns of problems on which frontier models consistently struggle. We provide code at https://github.com/MadryLab/platinum-benchmarks
Forward citations
Cited by 20 Pith papers
-
Fluid Language Model Benchmarking
Fluid Benchmarking, combining IRT-based ability estimation with Fisher-information-based adaptive item selection, improves LM evaluation across efficiency, validity, variance, and saturation in pretraining settings.
-
Flaws in the LLM Automation Narrative
A new code-writing data analysis benchmark shows human experts outperforming a frontier LLM on average with lower performance variance.
-
Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility
SGRE extracts a reasoning skeleton from a teacher trace, coarsens its graph, and verbalizes it densely; the final answer is preserved verbatim, and students distilled on the edited traces show large accuracy drops.
-
Auditing LLM Benchmarks with Item Response Theory
An IRT-based detector identifies mislabeled examples in LLM benchmarks at 95% precision in the top 200 cases, outperforming supervised classifiers and revealing reward-model specialization on style over facts.
-
Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation
Item-level Reliable Change Index analysis shows that LLM version upgrades result in bidirectional performance shifts on individual questions, making aggregate accuracy gains the net residual of improvements and deteri...
-
Position: Evaluation of ECG Representations Must Be Fixed
Current ECG representation benchmarks overstate the benefits of pretraining and produce unstable method rankings; a random encoder with linear probing is competitive on many tasks.
-
Weight Decay Improves Language Model Plasticity
Pretrained models trained with larger weight decay fine-tune better on downstream tasks, so the best pretraining checkpoint by loss is not always the best starting point for later training.
-
Model soups need only one ingredient
A single checkpoint, edited by splitting each layer's update with SVD and reweighting the high- and low-energy parts, reaches soup-level OOD robustness without multi-model training.
-
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
An IRT-based adaptive testing framework, ATLAS, estimates LLM ability with 30-89 items per benchmark, matching whole-bank ability estimates and re-ranking 23-31% of models relative to accuracy.
-
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models
League of LLMs organizes LLMs into a self-governed mutual evaluation league using dynamic, transparent, objective, and professional criteria to distinguish model capabilities with 70.7% top-k ranking stability.
-
Benchmarking Misuse Mitigation Against Covert Adversaries
Develops the BSD data generation pipeline and two new datasets to evaluate decomposition attacks as effective misuse enablers and stateful defenses as a countermeasure in language model safety.
-
A Sovereign, Open-Source Foundation Model for German and English
A fully documented German–English hybrid MoE base model matches dense 14–27B peers, leads open code scores, and sustains high long-context throughput at 3B active parameters.
-
A Sovereign, Open-Source Foundation Model for German and English
Soofi S 30B-A3B, a fully documented German–English hybrid Mamba-MoE base model, matches dense 14–27B peers on bilingual aggregates while delivering 8–9× long-context decode throughput.
-
A Sovereign, Open-Source Foundation Model for German and English
Soofi S 30B-A3B, a hybrid Mamba-MoE model pretrained on ~27T tokens with deliberately up-weighted German, reports the highest English and German aggregate scores among fully open base models in its comparison while ma...
-
Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity
An LLM-as-a-judge evaluation framework for math reasoning outperforms symbolic methods by accurately assessing diverse answer representations and formats.
-
LLM Reasoning Is Latent, Not the Chain of Thought
LLM reasoning is primarily mediated by latent-state trajectories rather than by explicit surface chain-of-thought outputs.
-
Simple Policy Gradients for Reasoning with Diffusion Language Models
AGRPO makes GRPO-style policy gradients tractable for diffusion LLMs by Monte-Carlo sampling denoising timesteps, but the unbiasedness claim only holds for a step-level objective, not the token-level GRPO objective.
-
Position: AI Evaluations Should be Grounded on a Theory of Capability
AI evaluations should be reframed as inference tasks grounded in an explicit theory of capability, with an empirical demonstration that results depend on modeling assumptions and a proposed Evaluation Card for transparency.
-
Kimi K2: Open Agentic Intelligence
Kimi K2 is a 1-trillion-parameter MoE model that leads open-source non-thinking models on agentic benchmarks including 65.8 on SWE-Bench Verified and 66.1 on Tau2-Bench.
-
Measuring AI Reasoning: A Guide for Researchers
Reasoning in language models should be measured by the faithfulness and validity of their multi-step search processes and intermediate traces, not final-answer accuracy.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.