REVIEW 23 cited by
Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We identify label errors in the test sets of 10 of the most commonly-used computer vision, natural language, and audio datasets, and subsequently study the potential for these label errors to affect benchmark results. Errors in test sets are numerous and widespread: we estimate an average of at least 3.3% errors across the 10 datasets, where for example label errors comprise at least 6% of the ImageNet validation set. Putative label errors are identified using confident learning algorithms and then human-validated via crowdsourcing (51% of the algorithmically-flagged candidates are indeed erroneously labeled, on average across the datasets). Traditionally, machine learning practitioners choose which model to deploy based on test accuracy - our findings advise caution here, proposing that judging models over correctly labeled test sets may be more useful, especially for noisy real-world datasets. Surprisingly, we find that lower capacity models may be practically more useful than higher capacity models in real-world datasets with high proportions of erroneously labeled data. For example, on ImageNet with corrected labels: ResNet-18 outperforms ResNet-50 if the prevalence of originally mislabeled test examples increases by just 6%. On CIFAR-10 with corrected labels: VGG-11 outperforms VGG-19 if the prevalence of originally mislabeled test examples increases by just 5%. Test set errors across the 10 datasets can be viewed at https://labelerrors.com and all label errors can be reproduced by https://github.com/cleanlab/label-errors.
Forward citations
Cited by 23 Pith papers
-
DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning
DecompSR is a large, symbolically verified benchmark dataset and generation framework that independently varies productivity, substitutivity, overgeneralisation, and systematicity to probe compositional multihop spati...
-
Do Large Language Model Benchmarks Test Reliability?
After removing label errors from fifteen standard benchmarks, frontier LLMs still fail simple tasks, so current benchmarks measure capability but not reliability.
-
A Novel Method to Evaluate Models on Unreliable, Noisy and Inconsistent Labels: Adaptive Resolution Label Aggregation (ARLA)
Adaptive Resolution Label Aggregation (ARLA) coarsens label and prediction together at chosen subpatch size and sensitivity so evaluation metrics better reflect true model error on noisy segmentation labels.
-
Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings
A ground-truth-first synthetic memory benchmark shows that agent-memory architecture rankings invert with history length: short-horizon leaders lose at nine weeks.
-
Representation Unlearning: Forgetting through Information Compression
Representation Unlearning removes the influence of specific training samples by learning a lightweight transformation over the model's penultimate-layer representations, guided by information-bottleneck variational bounds.
-
GFLC: Graph-based Fairness-aware Label Correction for Fair Classification
GFLC is a new label-correction method that uses confidence scores, graph curvature, and demographic parity to improve both accuracy and fairness under group-dependent label noise.
-
The Pitfalls of Benchmarking in Algorithm Selection: What We Are Getting Wrong
Non-informative features and a scale-only feature achieve strong prediction errors under common algorithm-selection evaluation schemes, without actually selecting better algorithms.
-
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
ZeroBench is a hand-built 100-question visual reasoning benchmark, adversarially filtered so every evaluated frontier LMM scored 0% at release.
-
Fundamental Challenges in Evaluating Text2SQL Solutions and Detecting Their Limitations
Aggregate Text2SQL benchmark numbers are distorted by ambiguous single labels and by the SQL-equivalence match functions, a problem the paper organizes into a taxonomy with concrete Spider examples.
-
NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries
NoisyEQA benchmarks embodied QA agents against four noise types and claims a self-correcting prompt (NACoT) markedly improves noisy-question accuracy as scored by GPT-4.
-
BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges
BACON calibrates multiple AI judges against a small human-labeled sample, then uses cross-fitted outcome models and augmented estimating equations to produce calibrated summary estimates and item-level surrogate scores.
-
Noise is not always detrimental: the capacity of quantum batteries is enhanced in black holes
Hawking radiation is claimed to enhance quantum battery capacity for bipartite mixed states, while environmental noise generally degrades it in type-dependent ways.
-
PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic
PaTAS propagates Subjective Logic trust opinions through every neuron of a network and updates parameter trust from gradient evidence, yielding per-prediction trust scores intended to flag poisoned or low-reliability inputs.
-
When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification
REVEAL ensembles four VLMs and label-noise detectors to detect and correct noisy and missing labels in six image classification test sets, reporting high agreement with human annotations.
-
When Dynamic Data Selection Meets Data Augmentation
A training framework that selects low-density, semantically consistent samples for light augmentation, claiming 50% cost reduction on ImageNet-1k with lossless performance.
-
Network Intrusion Datasets: A Survey, Limitations, and Recommendations
A systematic review of 89 public NIDS datasets with 13 extracted properties, popularity and trend analysis, and best practices for dataset selection, creation, and usage.
-
Image Recognition with Vision and Language Embeddings of VLMs
A benchmark of dual-encoder VLMs finds text and image embeddings give complementary class accuracy, and a per-class precision fusion rule adds about 0.4% accuracy over either alone on ImageNet.
-
Advancing Mental Disorder Detection: A Comparative Evaluation of Transformer and LSTM Architectures on Social Media
On Reddit posts labeled by subreddit membership, transformer models, led by RoBERTa, reach 99.5% F1, far above LSTM baselines, but the proxy labels may inflate the apparent detection ability.
-
Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning
A dynamic pruning method scores each sample by combining task loss with CLIP image-text similarity and selects samples near the median score each epoch.
-
First-of-its-kind AI model for bioacoustic detection using a lightweight associative memory Hopfield neural network
A Hopfield neural network trained on two bat calls classifies 10,384 recordings in 5.4 seconds with claimed accuracy up to 80%, though the headline numbers hinge on removing ambiguous calls.
-
Machine Unlearning for Robust DNNs: Attribution-Guided Partitioning and Neuron Pruning in Noisy Environments
The paper proposes attribution-guided data partitioning plus regression-based neuron pruning and fine-tuning for noisy training data, but the headline label-noise result is contradicted by the feature-noise-only experiments.
-
TMLC-Net: Transferable Meta Label Correction for Noisy Label Learning
TMLC-Net trains an LSTM to output softened versions of the noisy labels, then claims this transfers as label correction to new datasets without ever seeing a true clean label.
-
The Achilles Heel of AI: Fundamentals of Risk-Aware Training Data for High-Consequence Models
The paper claims that curated 20 to 40 percent subsets of training data can match full-data models and that shared label errors can inflate validation scores.
Discussion (0). Continue with ORCID to comment.