REVIEW 10 cited by
Leakage and the Reproducibility Crisis in ML-based Science
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The use of machine learning (ML) methods for prediction and forecasting has become widespread across the quantitative sciences. However, there are many known methodological pitfalls, including data leakage, in ML-based science. In this paper, we systematically investigate reproducibility issues in ML-based science. We show that data leakage is indeed a widespread problem and has led to severe reproducibility failures. Specifically, through a survey of literature in research communities that adopted ML methods, we find 17 fields where errors have been found, collectively affecting 329 papers and in some cases leading to wildly overoptimistic conclusions. Based on our survey, we present a fine-grained taxonomy of 8 types of leakage that range from textbook errors to open research problems. We argue for fundamental methodological changes to ML-based science so that cases of leakage can be caught before publication. To that end, we propose model info sheets for reporting scientific claims based on ML models that would address all types of leakage identified in our survey. To investigate the impact of reproducibility errors and the efficacy of model info sheets, we undertake a reproducibility study in a field where complex ML models are believed to vastly outperform older statistical models such as Logistic Regression (LR): civil war prediction. We find that all papers claiming the superior performance of complex ML models compared to LR models fail to reproduce due to data leakage, and complex ML models don't perform substantively better than decades-old LR models. While none of these errors could have been caught by reading the papers, model info sheets would enable the detection of leakage in each case.
Forward citations
Cited by 10 Pith papers
-
The Generalization Gap in Machine Learning EoS Inference from Core-Collapse Supernova Gravitational Waves
LightGBM and other regressors achieve R^{2}≈0.6–0.7 under random CV on CCSN GW catalogues but collapse to worse-than-mean performance under Leave-One-EoS-Out validation, exposing a generalisation gap for unseen EoS families.
-
U-Net based particle localization in granular experiments: Accuracy limits and optimization
A U-Net with optimized anti-aliased masks and multi-labeler training achieves 97.7% recall and 3.7%-diameter accuracy in granular particle localization, with human-labeler bias defining the accuracy floor.
-
A Reproducible, Scalable Pipeline for Synthesizing Autoregressive Model Literature
A scalable literature-synthesis pipeline that retrieves, filters, extracts, summarizes, and converts AR-model papers into runnable training scripts, with F1 > 0.85 extraction and three reproduction case studies.
-
Explanatory Debiasing: Involving Domain Experts in the Data Generation Process to Mitigate Representation Bias in AI Systems
Involving domain experts in AI data generation can reduce representation bias while maintaining or slightly improving model accuracy, according to a 35-participant healthcare user study.
-
Bridging the Data Provenance Gap Across Text, Speech and Video
A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with...
-
Wavelet Scattering Transform for Interpretable Schizophrenia Biomarker Discovery and Classification from Resting-State EEG
A single sentence stating the discovery directly. ≤ 300 chars.
-
Are Large Language Models Memorizing Bug Benchmarks?
Some base LLMs, particularly codegen-multi, show strong memorization of the Defects4J bug benchmark, while newer models like LLaMa 3.1 show weaker leakage signals.
-
Provenance Tracking in Large-Scale Machine Learning Systems
yProv4ML is a new provenance-tracking library for ML workflows that logs experiments as W3C PROV-compliant provenance graphs and demonstrates its use in large-scale distributed training studies.
-
Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
A position paper argues that understanding AI's second-order effects requires moving from static benchmarks to an ecosystem of field testing, red teaming, and contextual evaluation.
-
autrainer: A Modular and Extensible Deep Learning Toolkit for Computer Audition Tasks
autrainer is a config-driven PyTorch toolkit for computer audition that supports low-code training, preprocessing pipelines, augmentation, and release of pretrained audio models.
Discussion (0). Continue with ORCID to comment.