Pith. sign in

REVIEW 5 cited by

Improving Reproducibility in Machine Learning Research (A Report from the NeurIPS 2019 Reproducibility Program)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.12206 v4 pith:OZ55IPEK submitted 2020-03-27 cs.LG stat.ML

classification cs.LGstat.ML
keywords reproducibilityresearchlearningmachineprogramcodecommunitycomponents
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

One of the challenges in machine learning research is to ensure that presented and published results are sound and reliable. Reproducibility, that is obtaining similar results as presented in a paper or talk, using the same code and data (when available), is a necessary step to verify the reliability of research findings. Reproducibility is also an important step to promote open and accessible research, thereby allowing the scientific community to quickly integrate new findings and convert ideas to practice. Reproducibility also promotes the use of robust experimental workflows, which potentially reduce unintentional errors. In 2019, the Neural Information Processing Systems (NeurIPS) conference, the premier international conference for research in machine learning, introduced a reproducibility program, designed to improve the standards across the community for how we conduct, communicate, and evaluate machine learning research. The program contained three components: a code submission policy, a community-wide reproducibility challenge, and the inclusion of the Machine Learning Reproducibility checklist as part of the paper submission process. In this paper, we describe each of these components, how it was deployed, as well as what we were able to learn from this initiative.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports

    cs.AI 2026-07 accept novelty 6.0 of 10

    A prospective benchmark of 8,736 World Cup match forecasts shows that seven frontier LLMs perform comparably, web access adds a small Brier-score improvement, and prompting order and forecast horizon have little effect.

  2. The C-index illusion: discrimination without calibration in published survival models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Published survival models with high C-index scores can assign systematically wrong probabilities, and discrimination-only evaluation hides this failure.

  3. Improving the Reproducibility of Deep Learning Software: An Initial Investigation through a Case Study Analysis

    cs.LG 2025-05 conditional novelty 5.0 of 10

    The authors reproduce the TRUNK neural network across three datasets, find that missing training details cause large accuracy gaps, and extend existing reproducibility guidelines with sensitivity analysis and minimal ...

  4. On the missing benchmarks layer and a potential solution

    cs.AI 2026-08 unverdicted novelty 4.0 of 10

    An open regional EvalsHub, starting with LatamBoard, is proposed as the answer to Latin America's missing benchmark layer, but no benchmark or evaluation is presented.

  5. Prompt-Driven Image Analysis with Multimodal Generative AI: Detection, Segmentation, Inpainting, and Interpretation

    cs.CV 2025-09 conditional novelty 4.0 of 10

    LSID composes GroundingDINO, SAM, stable-diffusion inpainting, and LLaVA behind one prompt, with artifact logging and guardrails, and reports a small n=40 success slice.

Pith tools