{"id":"be6527c3-d517-4460-ad59-396981d34669","arxiv_id":"2412.16199","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Repeated randomized trials with per-subject leave-one-out validation stabilize random forest feature importance at subject and group levels.","lead":"This paper proposes a validation method that runs up to 400 randomized trials per subject to stabilize feature importance and accuracy of a single random forest model. It targets clinical researchers who want reproducible and explainable machine learning without building expensive subject-specific models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Voting over correct-prediction trials lacks a null model; observed feature stability may be an artifact of selection bias and unreported per-subject trial counts.","rationale":"The reader's weakest assumption correctly identifies that recording feature importance only from correct-prediction trials may introduce selection bias and lacks significance testing. My analysis agrees and sharpens this into a concrete, testable concern: without a permutation null or a comparison to all-trial voting, the claimed stability and clinical relevance of the top features are not established. The paper's own Table 2 also undercuts the 'improved accuracy' part of the central claim, but the feature-importance mechanism is the more load-bearing component. I do not see grounds to reject outright, because the method is plausible and the code is available; however, the current evidence is conditional on the proposed robustness checks being passed. Therefore I maintain the reader's conditional verdict, adding explicit requirements for a null model and per-subject convergence reporting.","tokens_in":11359,"tokens_out":4495,"duration_ms":40723,"concrete_test":"Run the proposed pipeline on the Alzheimer's dataset (48 subjects) with 1,000 label-permuted replicates to generate a null distribution for the frequency of each feature in the top-voted set. If the observed top features (FAST, M, TA3, A) do not exceed the 95th percentile of the null frequencies, the 'clinically meaningful' claim is unsupported. Additionally, for each subject, record the cumulative top-5 feature set after 50, 100, 200, and 400 trials; report the fraction of subjects whose top-5 set changes between 200 and 400 trials. Also compare subject-level rankings from correct trials vs. all trials; large divergence would confirm selection bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the assumption that voting over feature-importance sets from correct-prediction trials yields stable, meaningful subject- and group-level features. Section 2.3 states: 'If the model correctly predicts the outcome in each trial, the most important features are recorded.' This conditioning creates two unaddressed problems. First, the number of correct trials per subject is never reported; for difficult or minority-class subjects, the vote may be based on a small fraction of the 400 trials, so 'subject-specific' rankings can be high-variance even when group-level rankings look stable. Second, correctness is not independent of the random seed and the subject's feature profile, so the selected trials are a biased sample of model behavior. The paper provides no null model (e.g., label permutation) and no comparison to simpler aggregations (e.g., voting over all trials, or averaging feature importance over all seeds). Without these, the observed stability in Figures 3, 4, and 7 could be an artifact of the voting procedure rather than evidence of reproducible, biologically meaningful signal. The conclusion's 'improved accuracy' claim is also contradicted by Table 2 (99.5% vs 100% for 80/20 on breast cancer), though the accuracy claim is secondary to the feature-importance claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a randomized repeated-trials validation protocol around a single random forest: for each subject, the model is trained on all other subjects (LOSO) for up to 400 trials with different random seeds, feature importance is recorded only when the model predicts that subject correctly, and the recorded importance sets are voted on to obtain subject-level top features and then a group-level importance set. The authors reproduce the seed-sensitivity of a prior schizophrenia study, compare validation schemes across seven public datasets plus an Alzheimer's disease dataset, and evaluate the AD feature rankings against earlier clinical findings and Spearman correlations. They report stable feature rankings, comparable predictive accuracy, and release their R code publicly.","tokens_in":11622,"tokens_out":4652,"duration_ms":43108,"significance":"If the central claim holds, the method would offer a practical way to obtain subject-specific and group-level feature importance from a single general model, with implications for reproducible and explainable clinical machine learning. The paper has genuine strengths: it reproduces a published study, uses multiple public datasets, anchors the AD feature rankings to an external clinical result and to Spearman correlations, and makes its code available. However, the subject-specific claims currently rest on an unvalidated filtering-and-voting mechanism, the trial-count hyperparameter is tuned on the same data, and the stated accuracy improvement is not supported by Table 2. These are load-bearing issues that need additional analysis rather than mere polishing.","major_comments":[{"comment":"The method records feature importance only on trials where the LOSO model correctly predicts the current subject, but the paper never reports how many of the 400 trials are correct for each subject. For difficult or minority-class subjects this count could be small, making the 'subject-specific' ranking high-variance even if the group-level ranking appears stable. Correctness is also correlated with the random seed and with the subject's feature profile, so the filtered trials are a biased sample of model behavior. I request a label-permutation null model and a comparison against simpler aggregations, such as voting over all trials or averaging feature importance over all seeds. Without these, Figures 3, 4, and 7 do not establish that the observed stability reflects biological signal rather than an artifact of the filtering procedure.","section":"2.3, Figure 1"},{"comment":"The abstract and the conclusion claim 'improved accuracy' and 'higher predictive accuracy,' but Table 2 shows that Random Trials achieves 99.50% on Breast Cancer versus 100.00% for the 80/20 split, and 100.00% on Alzheimer's Disease versus 100.00% for LOSO. Accuracy is at best comparable, not improved. The claim should be corrected, and the accuracy comparison should report variance or confidence intervals over subjects and trials rather than single point estimates.","section":"Abstract, Section 4, Table 2"},{"comment":"The choice of 400 trials is justified only by the statement that 'experimentation using trial counts ranging from 50 to 1,000 showed an optimal maximum of 400.' The criterion for optimality, the datasets used for this tuning, and any held-out evaluation are not given. Because this hyperparameter is selected on the same data that are later used for the main evaluation, it functions as a tuned free parameter; the authors should report stability as a function of trial count and use a separate criterion or validation step to justify 400.","section":"2.3"},{"comment":"The external validation against Besga et al. is suggestive but is limited to one dataset and to group-level features. The claim that subject-level feature importance ranks corresponded with group-level ranks is asserted without quantitative evidence, and subject-level plots are shown only for one AD subject and one diabetes subject. I request a quantitative measure of subject-group agreement (e.g., rank correlation or overlap fraction per subject) across the datasets, especially for the Alzheimer's dataset where the subject-specific claim is central.","section":"3.3, Figure 7"}],"minor_comments":[{"comment":"The four repeated '8. Diamonds' rows make the table appear to list twelve datasets while the abstract says nine; consider a single entry with the four sample sizes clearly indicated.","section":"Table 1"},{"comment":"The statement that the Breast Cancer result 'reaches feature importance stability within 256 trial iterations' is not reconciled with the 400-trial cap or with the 'optimal maximum of 400' in Section 2.3; please explain the relationship.","section":"3.2"},{"comment":"The sample size for Breast Cancer is listed as 630 in Table 2 but 683 in Table 1; verify and unify the numbers.","section":"Table 2"},{"comment":"Reference [18] appears malformed ('P. Futoma, Simons, The lancet digital health jf'); the complete author list and journal title should be restored.","section":"References"},{"comment":"The term 'subject' is used for medical datasets, but for non-medical datasets such as Cars, Glass, and Diamonds it presumably means a row or instance; please define this explicitly.","section":"2.3"},{"comment":"The agreement with Besga et al. is described as a 'strong correlation' in prose, but no rank-correlation or overlap statistic is reported for the top-k sets; a quantitative measure would strengthen the comparison.","section":"3.3"}],"recommendation":"major_revision","confidential_remarks":"I do not see a scope problem; the manuscript fits the journal's interests. The main report lists technical issues that require re-analysis rather than rejection, so I recommend major revision rather than reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: this is a plausible method paper with one genuinely new idea and one serious gap. The new idea is to run leave-one-subject-out up to 400 times per subject with different random seeds, record feature importance only from trials where the model predicts the subject correctly, then vote across trials for subject-level top features and across subjects for group-level features. That combination is new as far as I can tell—Henderson et al. suggested multiple seeds for performance, not this per-subject voting procedure.\n\nWhat the paper does well: it shows concrete seed sensitivity in the Chekroud et al. reproduction, it ships code and data on GitHub, and it validates the Alzheimer's result against Spearman correlations and the prior Besga et al. findings. That external check is real evidence and the best part of the paper.\n\nThe soft spots are real. First, conditioning on correct-prediction trials creates a biased sample of model behavior, and the paper never says how many of the 400 trials were correct per subject. For hard or minority-class subjects that count could be small, making the subject-level vote high-variance even if group-level looks stable. There is no null model (label permutation) and no comparison to simpler baselines like voting over all trials or averaging importance over seeds. The observed stability in Figures 3, 4, and 7 could be an artifact of the voting procedure rather than reproducible biological signal. This is the load-bearing gap.\n\nSecond, the abstract and conclusion claim improved accuracy, but Table 2 shows Random Trials underperforming the 80/20 split on Breast Cancer (99.5% vs 100%). That is a direct contradiction.\n\nThird, the 400-trial cutoff was chosen by tuning on the same datasets, so the optimal is suspect. Minor issues: no error bars on the accuracy comparisons, and the top-k selection is unspecified.\n\nOverall, the central claim is not disproven, but it is under-supported. The method deserves a serious referee, and the fix is clear: add a permutation null model, report per-subject trial counts, and compare to simpler aggregations. I would like to see the paper in a reading group as a case study in validation design.\n\nRecommendation: send to peer review, but expect major revisions.","headline":"A genuinely new per-subject voting scheme for RF feature importance, but the correct-trial conditioning needs a null model before the stability claims can be trusted.","tokens_in":12132,"tokens_out":2708,"would_cite":true,"duration_ms":23956,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T01","92-08"],"pacs":["07.05.Mh","87.19.La"],"model":"deepseek-v4-flash","headline":"A repeated-trials validation approach with per-subject random seeds stabilizes feature importance in a single random forest model, giving subject- and group-level explanations without sacrificing accuracy.","keywords":["random seed variation","repeated trials validation","feature importance stability","reproducibility","random forest","subject-specific feature importance","explainable AI","leave-one-subject-out"],"falsifier":"Run the same 400-trial procedure on each dataset with class labels randomly permuted; if the top subject- and group-level feature sets still stabilize and rank as sharply as they do with true labels, the voting procedure is selecting seed-induced artifacts rather than subject-specific biology.","tokens_in":11175,"feed_emoji":"🎲","tokens_out":4982,"duration_ms":46326,"temperature":0.7,"pith_summary":"This paper tries to establish that the instability of machine-learning results in medical research, where different random seeds and validation schemes produce different accuracy and feature rankings, can be removed by a repeated-trials validation scheme instead of by building subject-specific models. The scheme trains one general random forest, leaves each subject out in turn, and repeats up to 400 times with a fresh random seed per subject per trial. If the model predicts the subject's outcome correctly, the trial's feature importance set is recorded; voting across trials produces a per-subject ranking and then a group-level ranking. Tested on nine datasets, the paper argues this yields reproducible accuracy and stable, clinically plausible feature importance, offering a practical and explainable alternative to costly per-patient models.","feed_headline":"Repeat trials, not new models, stabilize medical ML results","feed_subtitle":"Single random forest, 400 seeded runs per patient: stable feature rankings that agree with Alzheimer's biomarkers.","key_machinery":"The mechanism is a randomized repeated-trial validation loop: for each subject, a random seed is drawn from the subject index and trial number, a random forest is trained on all other subjects and tested on the held-out subject, and this is repeated for up to 400 trials. Only trials with a correct prediction contribute their feature importance sets. Those sets are aggregated by voting, first per subject to yield the subject's top features, then across subjects to yield the group-level set. The fixed per-subject random seed, the trial-varying seed, and the voting aggregation together carry the argument.","core_discovery":"The paper's central claim is that a single general random forest model can deliver both group-level and subject-specific feature importance if its leave-one-subject-out validation is repeated across up to 400 seeded trials, with feature importance collected only from trials that predicted the held-out subject's outcome correctly and then ranked by voting. Stability is the result: the repeated trials absorb the seed-induced variability that makes single-split or single-seed feature rankings unreliable. The paper shows that on the Alzheimer's disease dataset the stabilized rankings put four of the top five features among those previously reported as statistically significant, matching Spearman correlations of the same features with the outcome, at accuracy equal to conventional validation.","pith_inferences":["Editorial: The correct-trial filtering implicitly weights trials the model finds easy, so the method likely mixes feature informativeness with task difficulty; a natural extension is to separate the two by comparing against a chance-level accuracy baseline.","Editorial: Nothing in the procedure is random-forest-specific: the same per-subject-per-trial seeding and voting could be applied to gradient-boosted trees, support-vector machines, or neural networks, provided feature importance is computed per trial.","Editorial: The voting step resembles ensemble feature selection; it could be combined with stability selection to attach confidence intervals to subject-level feature ranks, turning stable rankings into a statistical test rather than a heuristic.","Editorial: Subject-level feature importance that matches group-level ranks in the Alzheimer's data may be an artifact of a small homogeneous cohort; testing on cohorts with known subgroups would show whether subject-level differentiation adds information beyond the group."],"forward_implications":["Under the proposed validation scheme, a single general random forest can produce subject-level feature importance rankings, so clinicians can obtain individualized explanations without training a separate model per patient.","Because the 400-trial voting procedure stabilizes feature rankings across random seeds, conclusions about which biomarkers matter become reproducible across runs that would otherwise disagree.","On the Alzheimer's dataset, the stabilized rankings place four of the five top features among those previously reported statistically significant, suggesting the method can recover clinically meaningful signals from a small cohort.","Execution time is lower than leave-one-subject-out validation while matching its accuracy, so the method is affordable enough for repeated use in medical experiments.","Group-level feature importance derived from subject-level votes can prioritize cost-effective, low-risk biomarkers for clinical collection."],"supporting_citations":[{"why":"The Science study whose reproducibility is tested; supplies the initial schizophrenia datasets and the claim of illusory generalizability that motivates the new approach.","marker":"[19]"},{"why":"Source code and data from the original study, used to run the seed-changing reproduction experiment.","marker":"[20]"},{"why":"The random decision forests paper that introduces the RF algorithm used throughout the experiments.","marker":"[21]"},{"why":"The Alzheimer's disease dataset used to validate the proposed method in a real-world clinical context.","marker":"[29]"},{"why":"The random forests paper that defines out-of-bag variable importance and the bootstrap seeding process, which is the source of the instability the method targets.","marker":"[33]"},{"why":"A study showing that random seed choice can inflate performance estimates by up to two-fold, motivating the repeated-seeded-trials design.","marker":"[35]"},{"why":"A prior multivariate analysis of the same Alzheimer's dataset that defined statistically significant features, used to corroborate the stabilized feature rankings.","marker":"[37]"}],"fun_headline_variants":["Repeat seeded trials, not models, for stable ML rankings","One random forest, 400 seeds: reproducible feature importance","Seed variation stabilizes subject-specific ML insights","Repeated validation yields explainable AI without new models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Features recorded only in trials where the model predicted the subject's outcome correctly, ranked by votes across 400 seeded runs, correspond to real subject-specific biology rather than to which trials happened to be easiest.","fun_headline_variants_meta":{"raw":{"variants":["Repeat seeded trials, not models, for stable ML rankings","One random forest, 400 seeds: reproducible feature importance","Seed variation stabilizes subject-specific ML insights","Repeated validation yields explainable AI without new models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0006,"raw_usage":{"total_tokens":2800,"prompt_tokens":938,"completion_tokens":1862,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":1808}},"tokens_in":554,"tokens_out":1862,"duration_ms":13611,"temperature":1.0,"reasoning_tokens":1808,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:06:28.933299+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 400-trial procedure on each dataset with class labels randomly permuted; if the top subject- and group-level feature sets still stabilize and rank as sharply as they do with true labels, the voting procedure is selecting seed-induced artifacts rather than subject-specific biology.","supporting_citations":[{"cited_title":"Chekroud, M","cited_arxiv_id":null,"evidence_quote":"Source code and data from the original study, used to run the seed-changing reproduction experiment."},{"cited_title":"Besga, M","cited_arxiv_id":null,"evidence_quote":"The Alzheimer's disease dataset used to validate the proposed method in a real-world clinical context."},{"cited_title":"Henderson, R","cited_arxiv_id":null,"evidence_quote":"A study showing that random seed choice can inflate performance estimates by up to two-fold, motivating the repeated-seeded-trials design."},{"cited_title":"Besga, I","cited_arxiv_id":null,"evidence_quote":"A prior multivariate analysis of the same Alzheimer's dataset that defined statistically significant features, used to corroborate the stabilized feature rankings."}],"review_version":1}