Pith. sign in

REVIEW 1 cited by

Reproducibility in Machine Learning-based Research: Overview, Barriers and Drivers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14325 v3 pith:Y7KCR3YP submitted 2024-06-20 cs.SE cs.IRcs.LG

classification cs.SEcs.IRcs.LG
keywords reproducibilityresearchissuesbarriersdriverspoorareascode
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many research fields are currently reckoning with issues of poor levels of reproducibility. Some label it a "crisis", and research employing or building Machine Learning (ML) models is no exception. Issues including lack of transparency, data or code, poor adherence to standards, and the sensitivity of ML training conditions mean that many papers are not even reproducible in principle. Where they are, though, reproducibility experiments have found worryingly low degrees of similarity with original results. Despite previous appeals from ML researchers on this topic and various initiatives from conference reproducibility tracks to the ACM's new Emerging Interest Group on Reproducibility and Replicability, we contend that the general community continues to take this issue too lightly. Poor reproducibility threatens trust in and integrity of research results. Therefore, in this article, we lay out a new perspective on the key barriers and drivers (both procedural and technical) to increased reproducibility at various levels (methods, code, data, and experiments). We then map the drivers to the barriers to give concrete advice for strategies for researchers to mitigate reproducibility issues in their own work, to lay out key areas where further research is needed in specific areas, and to further ignite discussion on the threat presented by these urgent issues.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics

    cs.SE 2025-05 conditional novelty 6.0 of 10

    Neural metrics for evaluating code comments are unreliable for multilingual output, often scoring random noise as high as real generated comments.

Pith tools