REVIEW 3 cited by
Sources of Irreproducibility in Machine Learning: A Review
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Background: Many published machine learning studies are irreproducible. Issues with methodology and not properly accounting for variation introduced by the algorithm themselves or their implementations are attributed as the main contributors to the irreproducibility.Problem: There exist no theoretical framework that relates experiment design choices to potential effects on the conclusions. Without such a framework, it is much harder for practitioners and researchers to evaluate experiment results and describe the limitations of experiments. The lack of such a framework also makes it harder for independent researchers to systematically attribute the causes of failed reproducibility experiments. Objective: The objective of this paper is to develop a framework that enable applied data science practitioners and researchers to understand which experiment design choices can lead to false findings and how and by this help in analyzing the conclusions of reproducibility experiments. Method: We have compiled an extensive list of factors reported in the literature that can lead to machine learning studies being irreproducible. These factors are organized and categorized in a reproducibility framework motivated by the stages of the scientific method. The factors are analyzed for how they can affect the conclusions drawn from experiments. A model comparison study is used as an example. Conclusion: We provide a framework that describes machine learning methodology from experimental design decisions to the conclusions inferred from them.
Forward citations
Cited by 3 Pith papers
-
Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events
Across 35 models and two languages, language models perform at or below chance at telling possible-but-unlikely events from impossible ones when semantic relatedness conflicts with possibility.
-
Statistical Quality and Reproducibility of Pseudorandom Number Generators in Machine Learning technologies
Comparing PRNGs in PyTorch, TensorFlow, and NumPy against reference C implementations with TestU01 BigCrush reveals implementation-specific test failures, though most vanish under a strict p-value filter.
-
Reducing Variability of Multiple Instance Learning Methods for Digital Pathology
Briefly training ten MIL models, merging the top three validation performers, and then fully training the merged model reduces run-to-run test AUC variability across five MIL methods on two pathology datasets.
Discussion (0). Sign in to comment.