Pith. sign in

REVIEW 13 cited by

Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.14749 v4 pith:REZHTCQR submitted 2021-03-26 stat.ML cs.AIcs.LG

Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

classification stat.ML cs.AIcs.LG
keywords errorstestdatasetslabelsetsacrosslabeledlearning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We identify label errors in the test sets of 10 of the most commonly-used computer vision, natural language, and audio datasets, and subsequently study the potential for these label errors to affect benchmark results. Errors in test sets are numerous and widespread: we estimate an average of at least 3.3% errors across the 10 datasets, where for example label errors comprise at least 6% of the ImageNet validation set. Putative label errors are identified using confident learning algorithms and then human-validated via crowdsourcing (51% of the algorithmically-flagged candidates are indeed erroneously labeled, on average across the datasets). Traditionally, machine learning practitioners choose which model to deploy based on test accuracy - our findings advise caution here, proposing that judging models over correctly labeled test sets may be more useful, especially for noisy real-world datasets. Surprisingly, we find that lower capacity models may be practically more useful than higher capacity models in real-world datasets with high proportions of erroneously labeled data. For example, on ImageNet with corrected labels: ResNet-18 outperforms ResNet-50 if the prevalence of originally mislabeled test examples increases by just 6%. On CIFAR-10 with corrected labels: VGG-11 outperforms VGG-19 if the prevalence of originally mislabeled test examples increases by just 5%. Test set errors across the 10 datasets can be viewed at https://labelerrors.com and all label errors can be reproduced by https://github.com/cleanlab/label-errors.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks

    cs.CL 2026-04 unverdicted novelty 7.0

    A schema-level diagnostic uses multi-annotator criterion judgments to separate unstable criteria from systematic category overlaps in subjective NLP annotation prior to gold-label creation.

  2. DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning

    cs.AI 2025-11 unverdicted novelty 7.0

    DecompSR is a large, symbolically verified benchmark dataset and generation framework that independently varies productivity, substitutivity, overgeneralisation, and systematicity to probe compositional multihop spati...

  3. FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

    cs.AI 2024-11 unverdicted novelty 7.0

    FrontierMath is a new benchmark of hundreds of original hard math problems that current AI models solve less than 2% of.

  4. Wake Vision: A Tailored Dataset and Benchmark Suite for TinyML Computer Vision Applications

    cs.CV 2024-05 unverdicted novelty 7.0

    Wake Vision pipeline produces a 6M-image person detection dataset for TinyML with 2.2% label error, improving model accuracy up to 6.6% over prior VWW benchmark across architectures and subsets.

  5. A Novel Method to Evaluate Models on Unreliable, Noisy and Inconsistent Labels: Adaptive Resolution Label Aggregation (ARLA)

    cs.CV 2026-07 conditional novelty 6.5

    Adaptive Resolution Label Aggregation (ARLA) coarsens label and prediction together at chosen subpatch size and sensitivity so evaluation metrics better reflect true model error on noisy segmentation labels.

  6. MMGist: A Comprehensive Multimodal Benchmark for 2027

    cs.CV 2026-06 unverdicted novelty 6.0

    MMGist filters 23,250 items from 18 benchmarks down to 7,262 using three-stage pipeline, preserving model rankings (Spearman ρ=0.98) while cutting items 69% and raising discrimination 78%.

  7. Learning to Annotate Delayed and False AEB Events: A Practical System for Extreme Class Imbalance and Asymmetric Label Noise

    cs.RO 2026-06 unverdicted novelty 6.0

    An automated AEB annotation framework uses data augmentation and noise suppression to achieve 80% recall improvement and 50% workload reduction for rare delayed/false triggers under class imbalance and asymmetric label noise.

  8. Signal-to-Noise Ratio and Sample Size Govern Representational Alignment in Neural Networks

    stat.ML 2026-05 unverdicted novelty 6.0

    Representational alignment varies monotonically with SNR and non-monotonically with sample size (minimized near interpolation threshold) across linear and nonlinear networks, and is decoupled from generalization error.

  9. Semantic Trimming and Auxiliary Multi-step Prediction for Generative Recommendation

    cs.IR 2026-04 unverdicted novelty 6.0

    STAMP mitigates semantic dilution in SID-based generative recommendation via adaptive input pruning and densified output supervision, delivering 1.23-1.38x speedup and 17-55% VRAM savings with maintained or improved accuracy.

  10. Noise is not always detrimental: the capacity of quantum batteries is enhanced in black holes

    quant-ph 2026-04 unverdicted novelty 5.0

    Hawking radiation is claimed to enhance quantum battery capacity for bipartite mixed states, while environmental noise generally degrades it in type-dependent ways.

  11. Noise is not always detrimental: the capacity of quantum batteries is enhanced in black holes

    quant-ph 2026-04 unverdicted novelty 5.0

    Hawking radiation enhances quantum battery capacity in black hole spacetimes, counter to typical noise effects, with degradation patterns depending on noise type.

  12. Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets

    cs.CV 2026-05 unverdicted novelty 4.0

    A validation-free metric combining neighbor-consistency and effective rank to estimate face recognition dataset quality for downstream model performance.

  13. Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing

    cs.CL 2026-04 unverdicted novelty 4.0

    A scoping review organizes decades of NLP evaluation debates into a taxonomy of recurring concerns and trade-offs with a structured checklist for better evaluation design.