REVIEW 7 cited by
Model Evaluation, Model Selection, and Algorithm Selection in Machine Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The correct use of model evaluation, model selection, and algorithm selection techniques is vital in academic machine learning research as well as in many industrial settings. This article reviews different techniques that can be used for each of these three subtasks and discusses the main advantages and disadvantages of each technique with references to theoretical and empirical studies. Further, recommendations are given to encourage best yet feasible practices in research and applications of machine learning. Common methods such as the holdout method for model evaluation and selection are covered, which are not recommended when working with small datasets. Different flavors of the bootstrap technique are introduced for estimating the uncertainty of performance estimates, as an alternative to confidence intervals via normal approximation if bootstrapping is computationally feasible. Common cross-validation techniques such as leave-one-out cross-validation and k-fold cross-validation are reviewed, the bias-variance trade-off for choosing k is discussed, and practical tips for the optimal choice of k are given based on empirical evidence. Different statistical tests for algorithm comparisons are presented, and strategies for dealing with multiple comparisons such as omnibus tests and multiple-comparison corrections are discussed. Finally, alternative methods for algorithm selection, such as the combined F-test 5x2 cross-validation and nested cross-validation, are recommended for comparing machine learning algorithms when datasets are small.
Forward citations
Cited by 7 Pith papers
-
A Statistical Test for the Benefits of Personalizing Interventions
KPT is a cross-fit doubly-robust test that controls Type I error and attains semiparametric efficiency for the personalization benefit of multi-action policies over the best single action.
-
Physics-Informed Machine Learning Under Small-Data Constraints: Lessons from Abrasive Waterjet Milling
For small AWJM process data, treating statistical curation as competing hypotheses, using multi-fold evaluation, and residual physics with GPs yields more stable rankings and calibrated uncertainty than single-split pure ML.
-
Automated identification of Ichneumonoidea wasps via YOLO-based deep learning: Integrating HiresCam for Explainable AI
A YOLOv26 classifier reaches 96% accuracy on 11 Hymenoptera families, and HiResCAM maps link its decisions to traditional taxonomic traits.
-
Can human clinical rationales improve the performance and explainability of clinical text classification models?
Adding 96,679 human rationale highlights improves cancer-site classification less than adding the same number of full pathology reports, and the explainability gain is small.
-
Wirelessly transmitted subthalamic nucleus signals predict endogenous pain levels in Parkinson's disease patients
STN local field potentials, transmitted wirelessly, predicted binary pain ratings in four of six PD-related pain reports, with bilateral frequency-band contributions.
-
Machine Learning Methods for Small Data and Upstream Bioprocessing Applications: A Comprehensive Review
The paper's new contribution is a taxonomy that organizes small-data ML methods by ML lifecycle stage rather than by technique.
-
Interpretable Machine Learning for Air Pollution and Respiratory Health Prediction: A Socioeconomic Subgroup Analysis
On a synthetic country-level climate-health dataset, PM2.5 is the dominant predictor of respiratory disease rate and AQI-derived air quality, and removing it sharply lowers classification accuracy.
Discussion (0). Sign in to comment.