Pith. sign in

REVIEW 1 cited by

When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.01781 v2 pith:DDECFACT submitted 2024-02-01 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords benchmarkbenchmarksleaderboardsmodelrankingsselectionanswerexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Model (LLM) leaderboards based on benchmark rankings are regularly used to guide practitioners in model selection. Often, the published leaderboard rankings are taken at face value - we show this is a (potentially costly) mistake. Under existing leaderboards, the relative performance of LLMs is highly sensitive to (often minute) details. We show that for popular multiple-choice question benchmarks (e.g., MMLU), minor perturbations to the benchmark, such as changing the order of choices or the method of answer selection, result in changes in rankings up to 8 positions. We explain this phenomenon by conducting systematic experiments over three broad categories of benchmark perturbations and identifying the sources of this behavior. Our analysis results in several best-practice recommendations, including the advantage of a hybrid scoring method for answer selection. Our study highlights the dangers of relying on simple benchmark evaluations and charts the path for more robust evaluation schemes on the existing benchmarks. The code for this paper is available at https://github.com/National-Center-for-AI-Saudi-Arabia/lm-evaluation-harness.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Holding one player model fixed, changing the verdict grammar, criterion disclosure, and budget rendering moved measured honesty verdicts dramatically, so eval findings can reflect the instrument.

Pith tools