Pith. sign in

REVIEW 10 cited by

Leakage and the Reproducibility Crisis in ML-based Science

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.07048 v1 pith:CVGV3CCR submitted 2022-07-14 cs.LG cs.AIstat.ME

classification cs.LGcs.AIstat.ME
keywords leakagemodelsreproducibilityerrorsml-basedsciencecomplexdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The use of machine learning (ML) methods for prediction and forecasting has become widespread across the quantitative sciences. However, there are many known methodological pitfalls, including data leakage, in ML-based science. In this paper, we systematically investigate reproducibility issues in ML-based science. We show that data leakage is indeed a widespread problem and has led to severe reproducibility failures. Specifically, through a survey of literature in research communities that adopted ML methods, we find 17 fields where errors have been found, collectively affecting 329 papers and in some cases leading to wildly overoptimistic conclusions. Based on our survey, we present a fine-grained taxonomy of 8 types of leakage that range from textbook errors to open research problems. We argue for fundamental methodological changes to ML-based science so that cases of leakage can be caught before publication. To that end, we propose model info sheets for reporting scientific claims based on ML models that would address all types of leakage identified in our survey. To investigate the impact of reproducibility errors and the efficacy of model info sheets, we undertake a reproducibility study in a field where complex ML models are believed to vastly outperform older statistical models such as Logistic Regression (LR): civil war prediction. We find that all papers claiming the superior performance of complex ML models compared to LR models fail to reproduce due to data leakage, and complex ML models don't perform substantively better than decades-old LR models. While none of these errors could have been caught by reading the papers, model info sheets would enable the detection of leakage in each case.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 139 citations worldwide. Full citation record

  1. The Generalization Gap in Machine Learning EoS Inference from Core-Collapse Supernova Gravitational Waves

    astro-ph.HE 2026-07 conditional novelty 6.5 of 10

    LightGBM and other regressors achieve R^{2}≈0.6–0.7 under random CV on CCSN GW catalogues but collapse to worse-than-mean performance under Leave-One-EoS-Out validation, exposing a generalisation gap for unseen EoS families.

  2. U-Net based particle localization in granular experiments: Accuracy limits and optimization

    cond-mat.stat-mech 2026-03 conditional novelty 6.0 of 10

    A U-Net with optimized anti-aliased masks and multi-labeler training achieves 97.7% recall and 3.7%-diameter accuracy in granular particle localization, with human-labeler bias defining the accuracy floor.

  3. A Reproducible, Scalable Pipeline for Synthesizing Autoregressive Model Literature

    cs.IR 2025-08 conditional novelty 6.0 of 10

    A scalable literature-synthesis pipeline that retrieves, filters, extracts, summarizes, and converts AR-model papers into runnable training scripts, with F1 > 0.85 extraction and three reproduction case studies.

  4. Explanatory Debiasing: Involving Domain Experts in the Data Generation Process to Mitigate Representation Bias in AI Systems

    cs.HC 2024-12 conditional novelty 6.0 of 10

    Involving domain experts in AI data generation can reduce representation bias while maintaining or slightly improving model accuracy, according to a 35-participant healthcare user study.

  5. Bridging the Data Provenance Gap Across Text, Speech and Video

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with...

  6. Wavelet Scattering Transform for Interpretable Schizophrenia Biomarker Discovery and Classification from Resting-State EEG

    eess.SP 2026-07 conditional novelty 5.0 of 10

    A single sentence stating the discovery directly. ≤ 300 chars.

  7. Are Large Language Models Memorizing Bug Benchmarks?

    cs.SE 2024-11 conditional novelty 5.0 of 10

    Some base LLMs, particularly codegen-multi, show strong memorization of the Defects4J bug benchmark, while newer models like LLaMa 3.1 show weaker leakage signals.

  8. Provenance Tracking in Large-Scale Machine Learning Systems

    cs.LG 2025-07 conditional novelty 4.0 of 10

    yProv4ML is a new provenance-tracking library for ML workflows that logs experiments as W3C PROV-compliant provenance graphs and demonstrates its use in large-scale distributed training studies.

  9. Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects

    cs.CY 2025-05 conditional novelty 4.0 of 10

    A position paper argues that understanding AI's second-order effects requires moving from static benchmarks to an ecosystem of field testing, red teaming, and contextual evaluation.

  10. autrainer: A Modular and Extensible Deep Learning Toolkit for Computer Audition Tasks

    cs.SD 2024-12 conditional novelty 4.0 of 10

    autrainer is a config-driven PyTorch toolkit for computer audition that supports low-code training, preprocessing pipelines, augmentation, and release of pretrained audio models.

Pith tools