REVIEW 4 major objections 6 minor 45 references
Stabilizing Machine Learning for Reproducible and Explainable Results: A Novel Validation Approach to Subject-Specific Insights
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A repeated-trials validation approach with per-subject random seeds stabilizes feature importance in a single random forest model, giving subject- and group-level explanations without sacrificing accuracy.
desk verdict A genuinely new per-subject voting scheme for RF feature importance, but the correct-trial conditioning needs a null model before the stability claims can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is a randomized repeated-trial validation loop: for each subject, a random seed is drawn from the subject index and trial number, a random forest is trained on all other subjects and tested on the held-out subject, and this is repeated for up to 400 trials. Only trials with a correct prediction contribute their feature importance sets. Those sets are aggregated by voting, first per subject to yield the subject's top features, then across subjects to yield the group-level set. The fixed per-subject random seed, the trial-varying seed, and the voting aggregation together carry the argument.
What would settle it
Run the same 400-trial procedure on each dataset with class labels randomly permuted; if the top subject- and group-level feature sets still stabilize and rank as sharply as they do with true labels, the voting procedure is selecting seed-induced artifacts rather than subject-specific biology.
Extended reading notes
Core claim
The paper's central claim is that a single general random forest model can deliver both group-level and subject-specific feature importance if its leave-one-subject-out validation is repeated across up to 400 seeded trials, with feature importance collected only from trials that predicted the held-out subject's outcome correctly and then ranked by voting. Stability is the result: the repeated trials absorb the seed-induced variability that makes single-split or single-seed feature rankings unreliable. The paper shows that on the Alzheimer's disease dataset the stabilized rankings put four of the top five features among those previously reported as statistically significant, matching Spearman correlations of the same features with the outcome, at accuracy equal to conventional validation.
Load-bearing premise
Features recorded only in trials where the model predicted the subject's outcome correctly, ranked by votes across 400 seeded runs, correspond to real subject-specific biology rather than to which trials happened to be easiest.
Editorial extensions
If this is right
- Under the proposed validation scheme, a single general random forest can produce subject-level feature importance rankings, so clinicians can obtain individualized explanations without training a separate model per patient.
- Because the 400-trial voting procedure stabilizes feature rankings across random seeds, conclusions about which biomarkers matter become reproducible across runs that would otherwise disagree.
- On the Alzheimer's dataset, the stabilized rankings place four of the five top features among those previously reported statistically significant, suggesting the method can recover clinically meaningful signals from a small cohort.
- Execution time is lower than leave-one-subject-out validation while matching its accuracy, so the method is affordable enough for repeated use in medical experiments.
- Group-level feature importance derived from subject-level votes can prioritize cost-effective, low-risk biomarkers for clinical collection.
Reading between the lines
- Editorial: The correct-trial filtering implicitly weights trials the model finds easy, so the method likely mixes feature informativeness with task difficulty; a natural extension is to separate the two by comparing against a chance-level accuracy baseline.
- Editorial: Nothing in the procedure is random-forest-specific: the same per-subject-per-trial seeding and voting could be applied to gradient-boosted trees, support-vector machines, or neural networks, provided feature importance is computed per trial.
- Editorial: The voting step resembles ensemble feature selection; it could be combined with stability selection to attach confidence intervals to subject-level feature ranks, turning stable rankings into a statistical test rather than a heuristic.
- Editorial: Subject-level feature importance that matches group-level ranks in the Alzheimer's data may be an artifact of a small homogeneous cohort; testing on cohorts with known subgroups would show whether subject-level differentiation adds information beyond the group.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes a randomized repeated-trials validation protocol around a single random forest: for each subject, the model is trained on all other subjects (LOSO) for up to 400 trials with different random seeds, feature importance is recorded only when the model predicts that subject correctly, and the recorded importance sets are voted on to obtain subject-level top features and then a group-level importance set. The authors reproduce the seed-sensitivity of a prior schizophrenia study, compare validation schemes across seven public datasets plus an Alzheimer's disease dataset, and evaluate the AD feature rankings against earlier clinical findings and Spearman correlations. They report stable feature rankings, comparable predictive accuracy, and release their R code publicly.
Significance. If the central claim holds, the method would offer a practical way to obtain subject-specific and group-level feature importance from a single general model, with implications for reproducible and explainable clinical machine learning. The paper has genuine strengths: it reproduces a published study, uses multiple public datasets, anchors the AD feature rankings to an external clinical result and to Spearman correlations, and makes its code available. However, the subject-specific claims currently rest on an unvalidated filtering-and-voting mechanism, the trial-count hyperparameter is tuned on the same data, and the stated accuracy improvement is not supported by Table 2. These are load-bearing issues that need additional analysis rather than mere polishing.
major comments (4)
- [2.3, Figure 1] The method records feature importance only on trials where the LOSO model correctly predicts the current subject, but the paper never reports how many of the 400 trials are correct for each subject. For difficult or minority-class subjects this count could be small, making the 'subject-specific' ranking high-variance even if the group-level ranking appears stable. Correctness is also correlated with the random seed and with the subject's feature profile, so the filtered trials are a biased sample of model behavior. I request a label-permutation null model and a comparison against simpler aggregations, such as voting over all trials or averaging feature importance over all seeds. Without these, Figures 3, 4, and 7 do not establish that the observed stability reflects biological signal rather than an artifact of the filtering procedure.
- [Abstract, Section 4, Table 2] The abstract and the conclusion claim 'improved accuracy' and 'higher predictive accuracy,' but Table 2 shows that Random Trials achieves 99.50% on Breast Cancer versus 100.00% for the 80/20 split, and 100.00% on Alzheimer's Disease versus 100.00% for LOSO. Accuracy is at best comparable, not improved. The claim should be corrected, and the accuracy comparison should report variance or confidence intervals over subjects and trials rather than single point estimates.
- [2.3] The choice of 400 trials is justified only by the statement that 'experimentation using trial counts ranging from 50 to 1,000 showed an optimal maximum of 400.' The criterion for optimality, the datasets used for this tuning, and any held-out evaluation are not given. Because this hyperparameter is selected on the same data that are later used for the main evaluation, it functions as a tuned free parameter; the authors should report stability as a function of trial count and use a separate criterion or validation step to justify 400.
- [3.3, Figure 7] The external validation against Besga et al. is suggestive but is limited to one dataset and to group-level features. The claim that subject-level feature importance ranks corresponded with group-level ranks is asserted without quantitative evidence, and subject-level plots are shown only for one AD subject and one diabetes subject. I request a quantitative measure of subject-group agreement (e.g., rank correlation or overlap fraction per subject) across the datasets, especially for the Alzheimer's dataset where the subject-specific claim is central.
minor comments (6)
- [Table 1] The four repeated '8. Diamonds' rows make the table appear to list twelve datasets while the abstract says nine; consider a single entry with the four sample sizes clearly indicated.
- [3.2] The statement that the Breast Cancer result 'reaches feature importance stability within 256 trial iterations' is not reconciled with the 400-trial cap or with the 'optimal maximum of 400' in Section 2.3; please explain the relationship.
- [Table 2] The sample size for Breast Cancer is listed as 630 in Table 2 but 683 in Table 1; verify and unify the numbers.
- [References] Reference [18] appears malformed ('P. Futoma, Simons, The lancet digital health jf'); the complete author list and journal title should be restored.
- [2.3] The term 'subject' is used for medical datasets, but for non-medical datasets such as Cars, Glass, and Diamonds it presumably means a row or instance; please define this explicitly.
- [3.3] The agreement with Besga et al. is described as a 'strong correlation' in prose, but no rank-correlation or overlap statistic is reported for the top-k sets; a quantitative measure would strengthen the comparison.
Circularity Check
No significant circularity: the proposed repeated-trials validation is an algorithmic protocol whose outputs are externally checked against prior clinical findings and correlation benchmarks.
full rationale
The paper's central derivation chain is empirical rather than formal: it proposes a repeated random-seed LOSO protocol, votes over feature-importance sets from correct-prediction trials, and then compares the resulting rankings with an independent prior study (Besga et al., [37]) and Spearman correlations computed on the same data. No equation in the paper defines the claimed subject- or group-level feature importance in terms of the validation's inputs, and no fitted parameter is renamed as a prediction. The choice of 400 trials is a tuned hyperparameter, but the paper's main external check (Alzheimer's dataset 9 against prior clinical findings) does not depend on a held-out prediction being forced by that tuning, so the core stabilization claim does not reduce to the tuning exercise. Conditioning on correct-prediction trials could introduce selection bias, and per-subject correct-trial counts are not reported, but these are correctness and robustness limitations rather than circular reductions. The paper is self-contained against an external benchmark, so the circularity score is 0.
Assumptions & free parameters
free parameters (3)
- Trial count (400) =
400
- Random seeds for baseline comparisons =
42, 43
- Top-k feature selection (k unspecified)
assumptions (4)
- domain assumption Feature importance stability is a proxy for model explainability
- domain assumption Random Forest is an appropriate base model for all nine datasets
- ad hoc to paper Correct predictions form a valid basis for feature selection
- domain assumption Spearman correlation between features and class labels is a valid reference for feature significance
Cite this review
Pith. "Pith review of Stabilizing Machine Learning for Reproducible and Explainable Results: A Novel Validation Approach to Subject-Specific Insights." pith.science (2026). https://pith.science/paper/OJUMPTD3
@misc{pith2026241216199,
author = {Pith},
title = {Pith review of: Stabilizing Machine Learning for Reproducible and Explainable Results: A Novel Validation Approach to Subject-Specific Insights},
year = {2026},
howpublished = {\url{https://pith.science/paper/OJUMPTD3}},
note = {Machine review of arXiv:2412.16199}
}
read the original abstract
Machine Learning is transforming medical research by improving diagnostic accuracy and personalizing treatments. General ML models trained on large datasets identify broad patterns across populations, but their effectiveness is often limited by the diversity of human biology. This has led to interest in subject-specific models that use individual data for more precise predictions. However, these models are costly and challenging to develop. To address this, we propose a novel validation approach that uses a general ML model to ensure reproducible performance and robust feature importance analysis at both group and subject-specific levels. We tested a single Random Forest (RF) model on nine datasets varying in domain, sample size, and demographics. Different validation techniques were applied to evaluate accuracy and feature importance consistency. To introduce variability, we performed up to 400 trials per subject, randomly seeding the ML algorithm for each trial. This generated 400 feature sets per subject, from which we identified top subject-specific features. A group-specific feature importance set was then derived from all subject-specific results. We compared our approach to conventional validation methods in terms of performance and feature importance consistency. Our repeated trials approach, with random seed variation, consistently identified key features at the subject level and improved group-level feature importance analysis using a single general model. Subject-specific models address biological variability but are resource-intensive. Our novel validation technique provides consistent feature importance and improved accuracy within a general ML model, offering a practical and explainable alternative for clinical research.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
- [1]
-
[2]
A. Karampuri, S. Kundur, S. Perugu, Exploratory drug discovery in breast cancer patients: A multimodal deep learning approach to identify novel drug candidates targeting rtk signaling, Computers in Biology and Medicine 174 (2024) 108433. doi:10.1016/j.compbiomed.2024. 108433
-
[3]
S. Bhattacharjee, B. Saha, S. Saha, Symptom-based drug prediction of lifestyle-related chronic diseases using unsupervised machine learning techniques, Computers in Biology and Medicine 174 (2024) 108413.doi: 10.1016/j.compbiomed.2024.108413
arXiv 2024
-
[4]
T. Zhu, X. Liu, J. Wang, R. Kou, Y. Hu, M. Yuan, C. Yuan, L. Luo, W. Zhang, Explainable machine-learning algorithms to differentiate bipolar disorder from major depressive disorder using self-reported symptoms, vital signs, and blood-based markers, Computer Methods and Programs in Biomedicine 240 (2023) 107723.doi:10.1016/j.cmpb. 2023.107723
arXiv 2023
-
[5]
L. Liu, Y. Li, N. Liu, J. Luo, J. Deng, W. Peng, Y. Bai, G. Zhang, G. Zhao, N. Yang, C. Li, X. Long, Establishment of machine learning- based tool for early detection of pulmonary embolism, Computer Meth- ods and Programs in Biomedicine 244 (2024) 107977. doi:10.1016/j. cmpb.2023.107977. 20
arXiv 2024
-
[6]
M. Aslam, F. Rajbdad, S. Azmat, Z. Li, J. P. Boudreaux, R. Thi- agarajan, S. Yao, J. Xu, A novel method for detection of pancre- atic ductal adenocarcinoma using explainable machine learning, Com- puter Methods and Programs in Biomedicine 245 (2024) 108019. doi: 10.1016/j.cmpb.2024.108019
arXiv 2024
-
[7]
Y. Yuan, C. Shi, H. Zhao, Machine learning-enabled genome mining and bioactivity prediction of natural products, ACS Synthetic Biology 12 (9) (2023) 2650–2662. doi:10.1021/acssynbio.3c00234
-
[8]
G. Magazz, G. Zampieri, C. Angione, Clinical stratification improves the diagnostic accuracy of small omics datasets within machine learning and genome-scale metabolic modelling methods, Computers in Biology and Medicine 151 (2022) 106244. doi:10.1016/j.compbiomed.2022. 106244
Show all 45 references
-
[9]
M. Yang, J. Ma, Machine learning methods for exploring sequence de- terminants of 3d genome organization, Journal of Molecular Biology 434 (15) (2022) 167666. doi:10.1016/j.jmb.2022.167666
2022
-
[10]
Ball, Is ai leading to a reproducibility crisis in science?, Nature 624 (7990) (2023) 22–25
P. Ball, Is ai leading to a reproducibility crisis in science?, Nature 624 (7990) (2023) 22–25. doi:10.1038/d41586-023-03817-6
2023 doi
-
[11]
Kapoor, A
S. Kapoor, A. Narayanan, Leakage and the reproducibility crisis in machine-learning-based science, Patterns 4 (9) (2023) 100804. doi: 10.1016/j.patter.2023.100804
2023
-
[12]
Ameli, L
A. Ameli, L. Pea-Castillo, H. Usefi, Assessing the reproducibility of machine-learning-based biomarker discovery in parkinsons disease, Com- puters in Biology and Medicine 174 (2024) 108407. doi:10.1016/j. compbiomed.2024.108407
2024
-
[13]
Van Noorden, J
R. Van Noorden, J. M. Perkel, Ai and science: what 1,600 re- searchers think, Nature 621 (7980) (2023) 672–675. doi:10.1038/ d41586-023-02980-0
2023
-
[14]
Gunning, D
D. Gunning, D. W. Aha, Darpas explainable artificial intelligence pro- gram, AI Magazine 40 (2) (2019) 44–58. doi:10.1609/aimag.v40i2. 2850. 21
2019 doi
-
[15]
Rasheed, A
K. Rasheed, A. Qayyum, M. Ghaly, A. Al-Fuqaha, A. Razi, J. Qadir, Explainable, trustworthy, and ethical machine learning for healthcare: A survey, Computers in Biology and Medicine 149 (2022) 106043. doi: 10.1016/j.compbiomed.2022.106043
2022
-
[16]
Nagendran, P
M. Nagendran, P. Festor, M. Komorowski, A. C. Gordon, A. A. Faisal, Quantifying the impact of ai recommendations with explanations on prescription decision making, npj Digital Medicine 6 (1) (Nov. 2023). doi:10.1038/s41746-023-00955-z
2023 doi
- [17]
-
[18]
Futoma, Simons, The lancet digital health jf - the lancet digi- tal health (2020)
P. Futoma, Simons, The lancet digital health jf - the lancet digi- tal health (2020). doi:10.1016/S2589-7500(20)30186-2DO-10.1016/ S2589-7500(20)30186-2T2
2020 doi
-
[19]
A. M. Chekroud, M. Hawrilenko, H. Loho, J. Bondar, R. Gueorguieva, A. Hasan, J. Kambeitz, P. R. Corlett, N. Koutsouleris, H. M. Krumholz, J. H. Krystal, M. Paulus, Illusory generalizability of clinical prediction models, Science 383 (6679) (2024) 164–167. doi:10.1126/science. adg8538
2024 doi
-
[20]
Chekroud, M
A. Chekroud, M. Hawrilenko, H. Loho, J. Bondar, R. Gueorguieva, A. Hasan, J. Kambeitz, P. Corlett, N. Koutsouleris, H. Krumholz, J. Krystal, M. Paulus, Code to accompany illusory generalizability of clinical prediction models (2023). doi:10.5281/ZENODO.10086334
2023 doi
-
[21]
T. K. Ho, Random decision forests, in: Proceedings of 3rd international conference on document analysis and recognition, Vol. 1, IEEE, 1995, pp. 278–282
1995
-
[22]
K. P. Bennett, O. L. Mangasarian, Robust linear programming discrim- ination of two linearly inseparable sets, Optimization Methods and Soft- ware 1 (1) (1992) 23–34. doi:10.1080/10556789208805504
1992 doi
-
[23]
Smith, J
J. Smith, J. Everhart, W. Dickson, W. Knowler, R. Johannes, Using the adap learning algorithm to forecast the onset of diabetes mellitus, 22 in: Proceedings of the Annual Symposium on Computer Application in Medical Care, Orlando, 1988, pp. 261–265
1988
-
[24]
James, D
G. James, D. Witten, T. Hastie, R. Tibshirani, An introduction to sta- tistical learning (2013)
2013
-
[25]
Ezekiel, Methods of correlation analysis (1930)
M. Ezekiel, Methods of correlation analysis (1930)
1930
-
[26]
Hothorn, A
T. Hothorn, A. Zeileis, ipred: Improved Predictors (2023). URL https://CRAN.R-project.org/package=ipred
2023
-
[27]
German, Glass identification (1987)
B. German, Glass identification (1987). doi:10.24432/C5WW2P
1987 doi
-
[28]
Wickham, Dataset: Diamonds (2019)
H. Wickham, Dataset: Diamonds (2019). doi:10.5281/ZENODO. 3522106
2019 doi
-
[29]
Besga, M
A. Besga, M. Graa, D. Chyzhyk, Alzheimer’s disease versus bipolar dis- order versus health control mri data and processed results, Frontiers in Aging Neuroscience (2020). doi:10.5281/ZENODO.3935636
2020 doi
-
[30]
URL https://www.R-project.org/
R Core Team, R: A Language and Environment for Statistical Comput- ing, R Foundation for Statistical Computing, Vienna, Austria (2021). URL https://www.R-project.org/
2021
- [31]
-
[32]
M. T. Ribeiro, S. Singh, C. Guestrin, Why should i trust you?: Ex- plaining the predictions of any classifier (Aug. 2016). doi:10.1145/ 2939672.2939778
2016
-
[33]
Breiman, Random forests, Machine Learning 45 (1) (2001) 5–32
L. Breiman, Random forests, Machine Learning 45 (1) (2001) 5–32. doi: 10.1023/a:1010933404324
2001 doi
-
[34]
A. L. Beam, A. K. Manrai, M. Ghassemi, Challenges to the reproducibil- ity of machine learning models in health care, JAMA 323 (4) (2020) 305. doi:10.1001/jama.2019.20866
2020
-
[35]
Henderson, R
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, D. Meger, Deep reinforcement learning that matters, Proceedings of the AAAI Conference on Artificial Intelligence 32 (1) (Apr. 2018). doi:10.1609/ aaai.v32i1.11694. 23
2018
-
[36]
R. D. Peng, Reproducible research in computational science, Science 334 (6060) (2011) 1226–1227. doi:10.1126/science.1213847
2011 doi
-
[37]
Besga, I
A. Besga, I. Gonzalez, E. Echeburua, A. Savio, B. Ayerdi, D. Chyzhyk, J. L. M. Madrigal, J. C. Leza, M. Graa, A. M. Gonzalez-Pinto, Discrim- ination between alzheimers disease and late onset bipolar disorder using multivariate analysis, Frontiers in Aging Neuroscience 7 (Dec. ...
2015
-
[38]
J. H. Chen, S. M. Asch, Machine learning and prediction in medicine beyond the peak of inflated expectations, New England Journal of Medicine 376 (26) (2017) 2507–2509. doi:10.1056/nejmp1702071
2017 doi
-
[39]
Kopitar, P
L. Kopitar, P. Kocbek, L. Cilar, A. Sheikh, G. Stiglic, Early detec- tion of type 2 diabetes mellitus using machine learning-based pre- diction models, Scientific Reports 10 (1) (Jul. 2020). doi:10.1038/ s41598-020-68771-z
2020
-
[40]
B. A. Goldstein, A. M. Navar, M. J. Pencina, J. P. A. Ioannidis, Op- portunities and challenges in developing risk prediction models with electronic health records data: a systematic review, Journal of the American Medical Informatics Association 24 (1) (2016) 198–208. doi: 10...
2016 doi
-
[41]
Bzdok, M
D. Bzdok, M. Krzywinski, N. Altman, Machine learning: a primer, Na- ture Methods 14 (12) (2017) 1119–1120. doi:10.1038/nmeth.4526
2017 doi
-
[42]
Obermeyer, E
Z. Obermeyer, E. J. Emanuel, Predicting the future big data, ma- chine learning, and clinical medicine, New England Journal of Medicine 375 (13) (2016) 1216–1219. doi:10.1056/nejmp1606181
2016 doi
-
[43]
Bouthillier, C
X. Bouthillier, C. Laurent, P. Vincent, Unreproducible research is re- producible, in: K. Chaudhuri, R. Salakhutdinov (Eds.), Proceedings of the 36th International Conference on Machine Learning, Vol. 97 of Pro- ceedings of Machine Learning Research, PMLR, 2019, pp. 725–734. U...
2019
-
[44]
Ciobanu-Caraus, A
O. Ciobanu-Caraus, A. Aicher, J. M. Kernbach, L. Regli, C. Serra, V. E. Staartjes, A critical moment in machine learning in medicine: on repro- ducible and interpretable learning, Acta Neurochirurgica 166 (1) (Jan. 2024). doi:10.1007/s00701-024-05892-8 . 24
2024 doi
-
[45]
M. L. Wallace, L. Mentch, B. J. Wheeler, A. L. Tapia, M. Richards, S. Zhou, L. Yi, S. Redline, D. J. Buysse, Use and misuse of random forest variable importance metrics in medicine: demonstrations through incident stroke prediction, BMC Medical Research Methodology 23 (1) (Jun...
2023 doi
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.