Pith. sign in

REVIEW 3 cited by

Sources of Irreproducibility in Machine Learning: A Review

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.07610 v2 pith:YGZ7LHQJ submitted 2022-04-15 cs.LG cs.AI

classification cs.LGcs.AI
keywords frameworkconclusionsexperimentslearningmachinedesignexperimentfactors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Background: Many published machine learning studies are irreproducible. Issues with methodology and not properly accounting for variation introduced by the algorithm themselves or their implementations are attributed as the main contributors to the irreproducibility.Problem: There exist no theoretical framework that relates experiment design choices to potential effects on the conclusions. Without such a framework, it is much harder for practitioners and researchers to evaluate experiment results and describe the limitations of experiments. The lack of such a framework also makes it harder for independent researchers to systematically attribute the causes of failed reproducibility experiments. Objective: The objective of this paper is to develop a framework that enable applied data science practitioners and researchers to understand which experiment design choices can lead to false findings and how and by this help in analyzing the conclusions of reproducibility experiments. Method: We have compiled an extensive list of factors reported in the literature that can lead to machine learning studies being irreproducible. These factors are organized and categorized in a reproducibility framework motivated by the stages of the scientific method. The factors are analyzed for how they can affect the conclusions drawn from experiments. A model comparison study is used as an example. Conclusion: We provide a framework that describes machine learning methodology from experimental design decisions to the conclusions inferred from them.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 31 citations worldwide. Full citation record

  1. Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Across 35 models and two languages, language models perform at or below chance at telling possible-but-unlikely events from impossible ones when semantic relatedness conflicts with possibility.

  2. Statistical Quality and Reproducibility of Pseudorandom Number Generators in Machine Learning technologies

    cs.OH 2025-07 conditional novelty 4.0 of 10

    Comparing PRNGs in PyTorch, TensorFlow, and NumPy against reference C implementations with TestU01 BigCrush reveals implementation-specific test failures, though most vanish under a strict p-value filter.

  3. Reducing Variability of Multiple Instance Learning Methods for Digital Pathology

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Briefly training ten MIL models, merging the top three validation performers, and then fully training the merged model reduces run-to-run test AUC variability across five MIL methods on two pathology datasets.

Pith tools