Pith. sign in

Quantifying Variance in Evaluation Benchmarks

25 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.

25 Pith papers citing it
2 external citations · Pith
abstract

Evaluation benchmarks are the cornerstone of measuring capabilities of large language models (LLMs), as well as driving progress in said capabilities. Originally designed to make claims about capabilities (or lack thereof) in fully pretrained models, evaluation benchmarks are now also extensively used to decide between various training choices. Despite this widespread usage, we rarely quantify the variance in our evaluation benchmarks, which dictates whether differences in performance are meaningful. Here, we define and measure a range of metrics geared towards measuring variance in evaluation benchmarks, including seed variance across initialisations, and monotonicity during training. By studying a large number of models -- both openly available and pretrained from scratch -- we provide empirical estimates for a variety of variance metrics, with considerations and recommendations for practitioners. We also evaluate the utility and tradeoffs of continuous versus discrete performance measures and explore options for better understanding and reducing this variance. We find that simple changes, such as framing choice tasks (like MMLU) as completion tasks, can often reduce variance for smaller scale ($\sim$7B) models, while more involved methods inspired from human testing literature (such as item analysis and item response theory) struggle to meaningfully reduce variance. Overall, our work provides insights into variance in evaluation benchmarks, suggests LM-specific techniques to reduce variance, and more generally encourages practitioners to carefully factor in variance when comparing models.

representative citing papers

The Art of Scaling Reinforcement Learning Compute for LLMs

cs.LG · 2025-10-15 · unverdicted · novelty 7.0

A 400k+ GPU-hour study shows RL scaling in LLMs follows predictable sigmoidal trajectories, with most design choices affecting efficiency rather than the performance asymptote, enabling accurate large-scale predictions via the ScaleRL recipe.

How Benchmark Prediction from Fewer Data Misses the Mark

cs.LG · 2025-06-09 · conditional · novelty 6.0

Benchmark prediction methods mostly work by interpolation among similar models and fail on better, unfamiliar models, where random sampling with an AIPW-style correction is the only consistent improvement.

HARP: A challenging human-annotated math reasoning benchmark

cs.LG · 2024-12-11 · conditional · novelty 6.0

HARP, a new benchmark of 5,409 US math competition problems with human solutions and choices, keeps frontier LLMs far from saturation: the best model scores 75.9% overall and only 41.1% on the hardest 197 problems.

Loss-to-Loss Prediction: Scaling Laws for All Datasets

cs.LG · 2024-11-19 · conditional · novelty 6.0

Losses of models trained on different datasets are related by shifted power laws, enabling translation of scaling laws and prediction of downstream performance from a few runs.

Validity Threats for Foundation Model Research

cs.LG · 2026-06-03 · accept · novelty 6.0

Maps common low-compute research strategies for foundation models onto statistical, internal, external, and construct validity threats via a causal-inference lens.

Resolution Diagnostics for Paired LLM Evaluation

cs.CL · 2026-05-28 · unverdicted · novelty 6.0

Paired LLM leaderboard comparisons frequently lack resolution at conventional (alpha=0.05, power=0.8) levels, with a new per-pair ratio q=N/N* showing that common unpaired shortcuts underestimate required samples by roughly a factor of two.

Statistical Multicriteria Evaluation of LLM-Generated Text

cs.CL · 2025-06-22 · conditional · novelty 4.0

Using generalized stochastic dominance, the authors find that human-written text completions are not significantly outperformed by five LLM decoding strategies across mixed cardinal and ordinal quality metrics.

citing papers explorer

Showing 25 of 25 citing papers.