{"id":"c0b4fde7-4b15-4251-8e4a-76ac9422182b","arxiv_id":"2605.24179","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"7T qMRI with DL segmentation and feature-selected ML reached 82% accuracy for PD vs HC, 100% for PIGD vs TD, and 73% multiclass on 45 subjects.","lead":"This paper applies 7 Tesla quantitative MRI maps and U-Net brain segmentation to train machine learning models that distinguish healthy controls from Parkinson's patients and separate motor subtypes. Feature selection on the imaging data raised reported accuracies substantially in a cohort of 45 subjects.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Feature selection on n=24 without confirmed nested CV produces optimistically biased accuracies (esp. 1.00), undermining the claim of reliable low-dimensional signatures.","rationale":"Reader's weakest assumption (non-generalizable sample + overfit from feature selection on n=24) is exactly the load-bearing risk for the classification claims. No other internal inconsistency (e.g., DSC or qMRI map issues) rises to the same level of threat to the central claim.","tokens_in":1858,"tokens_out":367,"duration_ms":21834,"concrete_test":"Re-derive the Approach B results using strictly nested 5-fold CV (feature selection performed only on each training fold, never on held-out data) and compare to the published numbers; if task-2 accuracy drops below ~0.80 or the gap versus Approach A shrinks substantially, the headline improvement is an artifact of leakage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that DL segmentation + qMRI feature selection (Approach B) improves classification over all features (Approach A) and supports interpretable signatures for PD diagnosis and subtype stratification. This rests on 5-fold CV results showing gains to 0.82/1.00/0.73 accuracy. With only 24 PD patients total (split across PIGD/TD for task 2), any non-nested feature selection—i.e., choosing the 'optimal subset' on the full cohort before CV—leaks test information and selects features that separate the current sample by chance. Even nested selection risks overfitting on such small n. The reported perfect separation for PIGD vs TD is the clearest symptom; the improvement claim cannot be trusted until the selection procedure is shown to be properly isolated from evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that deep-learning U-Net segmentation of 7T quantitative MRI maps, followed by machine-learning classification with feature selection (Approach B), improves performance over using all features (Approach A) across three tasks—HC vs PwP, PIGD vs TD, and multiclass—on a cohort of 21 HC and 24 PwP, yielding accuracies up to 0.82/1.00/0.73 and supporting low-dimensional interpretable imaging signatures for PD diagnosis and motor-subtype stratification.","tokens_in":2079,"tokens_out":601,"duration_ms":28078,"significance":"If validated, the work would demonstrate a pipeline linking automated qMRI segmentation to phenotype stratification in a small PD cohort, potentially aiding biological subtyping; however, the reported gains rest on unverified feature-selection procedures whose stability on n=24 is unproven.","major_comments":[{"comment":"Abstract and Methods (Approach B description): Feature selection is described as finding 'the optimal subset of features for the classification tasks' prior to reporting 5-fold CV results, with no statement that selection occurs inside each training fold. On a total of 24 PD patients this procedure risks selecting sample-specific features, directly undermining the claim that Approach B improves classification (e.g., Task 2 accuracy rising from 0.69 to 1.00).","section":"Abstract / Methods (Approach B)"},{"comment":"Results (Task 2, PIGD vs TD): The reported accuracy of 1.00 and AUC of 1.00 after feature selection on only 24 patients is statistically implausible without external validation or nested cross-validation; this single result is load-bearing for the central claim of 'interpretable, low-dimensional imaging signatures.'","section":"Results (classification performance)"},{"comment":"Methods (cross-validation details): The manuscript provides no description of nested cross-validation, held-out test set, or baseline clinical classifiers, leaving the reported gains (Approach B vs A) without a control for selection bias on small N.","section":"Methods (classification approaches)"}],"minor_comments":[{"comment":"Abstract: The sentence 'The U-Net achieved mean DSC of 0.86 for all ROIs during training' should clarify whether this is training or validation DSC and list the ROIs.","section":"Abstract"},{"comment":"The total sample size (N=45) and the split between PIGD and TD subgroups should be stated explicitly in the Methods when describing the classification tasks.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's reliance on a single small cohort without external validation or proper nested selection makes the headline performance numbers difficult to interpret; this is a methodological rather than interpretive issue."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments highlighting important methodological concerns with our small-cohort study. We address each major comment below and commit to revisions that improve transparency and rigor without overstating the current results.","responses":[{"response":"We agree that the manuscript description is ambiguous and does not confirm feature selection occurred inside the CV folds. This is a valid concern for selection bias on n=24. In revision we will rewrite the Methods to specify nested cross-validation, with feature selection performed independently on each training fold only, and will recompute and report the revised performance metrics.","revision_made":"yes","referee_comment":"[Abstract / Methods (Approach B)] Abstract and Methods (Approach B description): Feature selection is described as finding 'the optimal subset of features for the classification tasks' prior to reporting 5-fold CV results, with no statement that selection occurs inside each training fold. On a total of 24 PD patients this procedure risks selecting sample-specific features, directly undermining the claim that Approach B improves classification (e.g., Task 2 accuracy rising from 0.69 to 1.00)."},{"response":"We acknowledge that perfect separation on this small subsample is likely inflated by the non-nested feature selection and should not be presented as robust evidence. In the revision we will replace the current Task 2 numbers with results from properly nested CV, add explicit caveats in Results and Discussion about small-sample instability, and soften the language around the signatures to reflect that they are exploratory.","revision_made":"yes","referee_comment":"[Results (classification performance)] Results (Task 2, PIGD vs TD): The reported accuracy of 1.00 and AUC of 1.00 after feature selection on only 24 patients is statistically implausible without external validation or nested cross-validation; this single result is load-bearing for the central claim of 'interpretable, low-dimensional imaging signatures.'"},{"response":"We will expand the Methods to detail the nested CV procedure and state that a held-out test set was not used because of the limited total sample (n=45). We did not compare against clinical classifiers because the study focus was imaging-derived features; a brief note on this scope limitation will be added, but a full clinical baseline comparison is outside the current scope.","revision_made":"partial","referee_comment":"[Methods (classification approaches)] Methods (cross-validation details): The manuscript provides no description of nested cross-validation, held-out test set, or baseline clinical classifiers, leaving the reported gains (Approach B vs A) without a control for selection bias on small N."}],"tokens_in":1586,"tokens_out":572,"duration_ms":44285,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that the 1.00 accuracy and AUC for PIGD versus TD after feature selection on only 24 patients is almost certainly overfit. The abstract does not describe nested cross-validation, so the optimal subset was probably chosen on the full cohort before the 5-fold evaluation, which leaks test information and inflates the numbers.\n\nThe paper applies a standard U-Net to segment 7T quantitative maps and then runs off-the-shelf classifiers on the extracted features. They compare using every feature versus a selected subset across HC versus PD, PIGD versus TD, and the three-class task. The segmentation reaches a mean DSC of 0.86, which is solid for the regions involved, and the selected-feature approach does improve the cross-validation scores over the all-features baseline. That is the concrete positive result.\n\nThe soft spots are the sample size and validation design. With 24 PD patients total, even modest feature selection on the same data tends to pick up noise that separates the current groups by chance. No external validation set, no held-out test data, and no comparison against simple clinical scores are reported. These gaps make the claim of interpretable low-dimensional signatures hard to trust at face value.\n\nNothing methodologically new is presented; the methods are established tools applied to a new but small 7T cohort. The work is aimed at researchers already studying quantitative imaging in movement disorders who might want to see 7T data tried on motor subtyping. A reader seeking reliable biomarkers will find the evidence too preliminary. It deserves peer review if the authors can show the selection step was isolated from evaluation and add some form of external check; otherwise the overfitting issue is central and the paper is not ready.","headline":"Feature selection on 24 patients without nested CV produces suspiciously perfect scores that likely reflect overfitting rather than stable signatures.","tokens_in":2613,"tokens_out":420,"would_cite":false,"duration_ms":30811,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Selected quantitative MRI features from 7T scans let machine learning separate Parkinson's motor subtypes with high accuracy in cross-validation.","keywords":["Parkinson's disease","quantitative MRI","machine learning","subtype stratification","7 Tesla","deep learning segmentation","motor phenotypes","feature selection"],"falsifier":"Running the same pipeline on an independent cohort of comparable size and finding that the selected-feature classifiers drop below 0.80 accuracy on the binary tasks would falsify the reported utility of the signatures.","tokens_in":2798,"feed_emoji":"🧠","tokens_out":654,"duration_ms":30605,"temperature":0.7,"pith_summary":"The paper examines whether automatic segmentation of 7 Tesla quantitative MRI maps can yield brain-region features that, after selection, improve machine-learning classification of healthy controls versus Parkinson's patients and of the two main motor subtypes within patients. Three tasks were tested with 5-fold cross-validation on 21 controls and 24 patients: binary diagnosis, binary subtype separation, and three-class stratification. Using all features gave moderate performance, while selecting the best subset raised accuracies to 0.82, 1.00, and 0.73 respectively, indicating that compact imaging signatures may support objective subtype identification.","feed_headline":"Selected qMRI features classify Parkinson's subtypes at 100% accuracy","feed_subtitle":"Feature selection after U-Net segmentation on 7T scans lifts cross-validated performance on 45 subjects above use of all features.","key_machinery":"Optimal subset selection performed on quantitative MRI values extracted from U-Net-segmented brain regions; the step reduces feature count and raises cross-validated accuracy on the three classification tasks.","core_discovery":"Deep-learning U-Net segmentation of quantitative 7T MRI maps followed by feature selection produced classifiers whose performance exceeded that of models using every extracted feature, reaching perfect accuracy and AUC on the postural instability/gait difficulty versus tremor-dominant task and supporting the feasibility of low-dimensional, interpretable signatures for diagnosis support and phenotype stratification.","pith_inferences":["If the signatures prove stable, clinical rating scales could be supplemented or partially replaced by objective imaging metrics.","The same workflow might be tested on other movement disorders that also exhibit motor heterogeneity.","Scanner harmonization studies would be needed before multi-site deployment of the selected features.","Longitudinal scans could check whether the signatures track disease progression within each subtype."],"forward_implications":["Low-dimensional imaging signatures become feasible for supporting objective motor-subtype assignment.","Feature selection after deep-learning segmentation demonstrably improves classification over use of all features.","Quantitative 7T maps combined with automatic segmentation can highlight differences between controls and the two motor phenotypes.","The approach opens a route toward imaging-supported study design and personalized treatment planning in heterogeneous Parkinson's disease."],"fun_headline_variants":["Feature selection on 7T qMRI hits 100% accuracy for Parkinson's subtypes","7T qMRI feature selection classifies PIGD and TD at 100% accuracy","U-Net segmented qMRI maps stratify Parkinson's subtypes with selected features","Selected 7T features yield perfect accuracy in PD motor subtype classification"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 24 Parkinson's patients form a representative sample whose feature distributions will generalize, and feature selection on this cohort will not produce classifiers whose reported accuracies, especially the perfect score, will replicate on new data.","fun_headline_variants_meta":{"raw":{"variants":["Feature selection on 7T qMRI hits 100% accuracy for Parkinson's subtypes","7T qMRI feature selection classifies PIGD and TD at 100% accuracy","U-Net segmented qMRI maps stratify Parkinson's subtypes with selected features","Selected 7T features yield perfect accuracy in PD motor subtype classification"]},"model":"grok-4.3","cost_usd":0.006745,"raw_usage":{"total_tokens":3123,"prompt_tokens":796,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":67453000,"prompt_tokens_details":{"text_tokens":796,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2247,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":796,"tokens_out":80,"duration_ms":16313,"temperature":1.0,"reasoning_tokens":2247,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T14:19:56.434891+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the same pipeline on an independent cohort of comparable size and finding that the selected-feature classifiers drop below 0.80 accuracy on the binary tasks would falsify the reported utility of the signatures.","supporting_citations":[],"review_version":1}