Pith. sign in

REVIEW 7 cited by

Model Evaluation, Model Selection, and Algorithm Selection in Machine Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1811.12808 v3 pith:ATCQDFZX submitted 2018-11-13 cs.LG stat.ML

classification cs.LGstat.ML
keywords selectioncross-validationmodelalgorithmlearningmachinedifferentevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The correct use of model evaluation, model selection, and algorithm selection techniques is vital in academic machine learning research as well as in many industrial settings. This article reviews different techniques that can be used for each of these three subtasks and discusses the main advantages and disadvantages of each technique with references to theoretical and empirical studies. Further, recommendations are given to encourage best yet feasible practices in research and applications of machine learning. Common methods such as the holdout method for model evaluation and selection are covered, which are not recommended when working with small datasets. Different flavors of the bootstrap technique are introduced for estimating the uncertainty of performance estimates, as an alternative to confidence intervals via normal approximation if bootstrapping is computationally feasible. Common cross-validation techniques such as leave-one-out cross-validation and k-fold cross-validation are reviewed, the bias-variance trade-off for choosing k is discussed, and practical tips for the optimal choice of k are given based on empirical evidence. Different statistical tests for algorithm comparisons are presented, and strategies for dealing with multiple comparisons such as omnibus tests and multiple-comparison corrections are discussed. Finally, alternative methods for algorithm selection, such as the combined F-test 5x2 cross-validation and nested cross-validation, are recommended for comparing machine learning algorithms when datasets are small.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 256 citations worldwide. Full citation record

  1. A Statistical Test for the Benefits of Personalizing Interventions

    stat.ME 2026-07 accept novelty 6.5 of 10

    KPT is a cross-fit doubly-robust test that controls Type I error and attains semiparametric efficiency for the personalization benefit of multi-action policies over the best single action.

  2. Physics-Informed Machine Learning Under Small-Data Constraints: Lessons from Abrasive Waterjet Milling

    cs.LG 2026-07 conditional novelty 6.0 of 10

    For small AWJM process data, treating statistical curation as competing hypotheses, using multi-fold evaluation, and residual physics with GPs yields more stable rankings and calibrated uncertainty than single-split pure ML.

  3. Automated identification of Ichneumonoidea wasps via YOLO-based deep learning: Integrating HiresCam for Explainable AI

    cs.CV 2026-03 conditional novelty 5.0 of 10

    A YOLOv26 classifier reaches 96% accuracy on 11 Hymenoptera families, and HiResCAM maps link its decisions to traditional taxonomic traits.

  4. Can human clinical rationales improve the performance and explainability of clinical text classification models?

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Adding 96,679 human rationale highlights improves cancer-site classification less than adding the same number of full pathology reports, and the explainability gain is small.

  5. Wirelessly transmitted subthalamic nucleus signals predict endogenous pain levels in Parkinson's disease patients

    q-bio.NC 2025-06 conditional novelty 5.0 of 10

    STN local field potentials, transmitted wirelessly, predicted binary pain ratings in four of six PD-related pain reports, with bilateral frequency-band contributions.

  6. Machine Learning Methods for Small Data and Upstream Bioprocessing Applications: A Comprehensive Review

    cs.LG 2025-06 accept novelty 4.0 of 10

    The paper's new contribution is a taxonomy that organizes small-data ML methods by ML lifecycle stage rather than by technique.

  7. Interpretable Machine Learning for Air Pollution and Respiratory Health Prediction: A Socioeconomic Subgroup Analysis

    cs.LG 2026-07 conditional novelty 2.0 of 10

    On a synthetic country-level climate-health dataset, PM2.5 is the dominant predictor of respiratory disease rate and AQI-derived air quality, and removing it sharply lowers classification accuracy.

Pith tools