Domain-specific macOS features enable an ML detector to reach 98.5% accuracy on 41k samples and 99.5% on 9k fresh samples, beating prior methods by 16-50%.
Model Evaluation, Model Selection, and Algorithm Selection in Machine Learning
7 Pith papers cite this work, alongside 256 external citations. Polarity classification is still indexing.
abstract
The correct use of model evaluation, model selection, and algorithm selection techniques is vital in academic machine learning research as well as in many industrial settings. This article reviews different techniques that can be used for each of these three subtasks and discusses the main advantages and disadvantages of each technique with references to theoretical and empirical studies. Further, recommendations are given to encourage best yet feasible practices in research and applications of machine learning. Common methods such as the holdout method for model evaluation and selection are covered, which are not recommended when working with small datasets. Different flavors of the bootstrap technique are introduced for estimating the uncertainty of performance estimates, as an alternative to confidence intervals via normal approximation if bootstrapping is computationally feasible. Common cross-validation techniques such as leave-one-out cross-validation and k-fold cross-validation are reviewed, the bias-variance trade-off for choosing k is discussed, and practical tips for the optimal choice of k are given based on empirical evidence. Different statistical tests for algorithm comparisons are presented, and strategies for dealing with multiple comparisons such as omnibus tests and multiple-comparison corrections are discussed. Finally, alternative methods for algorithm selection, such as the combined F-test 5x2 cross-validation and nested cross-validation, are recommended for comparing machine learning algorithms when datasets are small.
years
2026 7representative citing papers
Releases the DAPWH dataset of 3556 wasp images including 1739 COCO-annotated examples to enable AI models for identifying Ichneumonoidea and associated families.
For small AWJM process data, treating statistical curation as competing hypotheses, using multi-fold evaluation, and residual physics with GPs yields more stable rankings and calibrated uncertainty than single-split pure ML.
ALL-FEM fine-tunes LLMs on a corpus of verified FEniCS scripts and uses multi-agent workflows to automate finite element code generation, achieving 71.79% success on 39 benchmarks across elasticity, flow, and coupled problems.
AI accuracy evaluation requires four normative choices on metrics, balancing, representative data, and thresholds that embed assumptions about risks and trade-offs, as analyzed through the EU AI Act.
The P3 selector achieves 0.9809 purity and 0.8869 completeness for QSO candidates in selected fields, outperforming Gaia's official probabilities.
Multi-modal ML framework on lncRNA expression, secondary structure, and sequence features from two cohorts identifies T2D associations with MEG3 dominant in both.
citing papers explorer
-
The Role of Domain-Specific Features in Malware Detection: A macOS Case Study
Domain-specific macOS features enable an ML detector to reach 98.5% accuracy on 41k samples and 99.5% on 9k fresh samples, beating prior methods by 16-50%.
-
Descriptor: Parasitoid Wasps and Associated Hymenoptera Dataset (DAPWH)
Releases the DAPWH dataset of 3556 wasp images including 1739 COCO-annotated examples to enable AI models for identifying Ichneumonoidea and associated families.
-
Physics-Informed Machine Learning Under Small-Data Constraints: Lessons from Abrasive Waterjet Milling
For small AWJM process data, treating statistical curation as competing hypotheses, using multi-fold evaluation, and residual physics with GPs yields more stable rankings and calibrated uncertainty than single-split pure ML.
-
ALL-FEM: Agentic Large Language models Fine-tuned for Finite Element Methods
ALL-FEM fine-tunes LLMs on a corpus of verified FEniCS scripts and uses multi-agent workflows to automate finite element code generation, achieving 71.79% success on 39 benchmarks across elasticity, flow, and coupled problems.
-
Is your AI Model Accurate Enough? The Difficult Choices Behind Rigorous AI Development and the EU AI Act
AI accuracy evaluation requires four normative choices on metrics, balancing, representative data, and thresholds that embed assumptions about risks and trade-offs, as analyzed through the EU AI Act.
-
A Gaia-linked High-purity QSO Candidate Catalog in Selected Fields with Extinction-binned Calibration and Spectrum-informed Training
The P3 selector achieves 0.9809 purity and 0.8869 completeness for QSO candidates in selected fields, outperforming Gaia's official probabilities.
-
Multi-Modal Machine Learning for Population- and Subject-Specific lncRNA-Type 2 Diabetes Association Analysis
Multi-modal ML framework on lncRNA expression, secondary structure, and sequence features from two cohorts identifies T2D associations with MEG3 dominant in both.