{"id":"a74558c6-6e39-4f82-becb-ea592f48404c","arxiv_id":"2505.23916","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A 3D CNN trained on synthetic motion artifacts predicts head motion in real structural MRI scans, reproducing known motion-related cortical thickness and age effects in 12 of 15 external datasets.","lead":"A deep-learning model trained on synthetically corrupted brain scans estimates head motion from ordinary structural MRI data, and the estimate tracks known motion-related biases in cortical thickness across 15 real datasets. It could give researchers a cheap, retrospective way to detect motion artifacts without special hardware.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Direct validation against continuous ground-truth motion (e.g., vNavs) is missing; MR-ART's 3-level label and indirect thickness correlations cannot distinguish motion from general image-quality bias.","rationale":"The reader's weakest assumption is the fidelity of synthetic motion artifacts to real motion across scanners, protocols, and populations (Sec. 2.3.2). I agree that this is the load-bearing premise for the central claim. My stress-test sharpens it: the only direct real-motion validation is a single dataset (MR-ART) with a three-level ordinal label, which cannot separate motion from other image-quality degradations. The thickness correlations in 12/15 datasets are consistent with a general quality index, since blur and ghosting also reduce FreeSurfer thickness and are age-related. The paper is otherwise strong: it provides code, weights, and a reproducible pipeline; the synthetic test performance (R^2 = 0.94) is high; and the MR-ART ranking is encouraging. The missing piece is a validation against continuous, recorded motion. Fortunately, the authors have access to HBN vNavs data at the CBIC and CUNY sites, where navigator-based motion estimates are available. Using a held-out subset of those participants would directly test whether the predicted score tracks real motion magnitude. If that validation succeeds, the central claim is much better supported; if it fails, the model may be a quality index rather than a motion estimator. This does not change the reader's conditional verdict, but it makes the condition concrete and testable.","tokens_in":13922,"tokens_out":6971,"duration_ms":74704,"concrete_test":"Run the trained model on held-out HBN participants scanned with vNavs at the CBIC and CUNY sites, and compute the Spearman correlation between the predicted motion score and the vNavs-recorded RMS displacement (or rotation/translation RMS) for each T1w volume. If the correlation is substantially lower than the MR-ART value (e.g., <0.4) or not significant, the synthetic-to-real transfer claim is unsupported; if the correlation is moderate-to-high (e.g., 0.6–0.8), the concern is resolved. This directly tests whether the score tracks continuous real motion rather than a coarse three-level quality grade.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, that the regressor estimates head motion from synthetic artifacts, rests on the assumption that TorchIO's k-space motion corruption (Sec. 2.3.2) transfers to real in-scanner motion. The only direct evidence is a Spearman correlation of 0.71 against MR-ART's three-level clinical-usability grades (Sec. 3.1.2), which conflates motion with overall image quality. The thickness and age correlations (Secs. 3.2.2 and 3.2.3) are indirect: any image-quality index that tracks blur or ghosting would reproduce them, because these artifacts also bias FreeSurfer thickness estimates and are more prevalent in children. The model's training labels are RMS deviations of simulated transforms, and the Discussion (Sec. 4) lists 'more accurate artifact simulators' as future work, conceding that the simulator is imperfect. Critically, the authors already possess data with continuous ground-truth motion: HBN CBIC/CUNY vNavs acquisitions (Sec. 2.1) record navigator-based motion, yet no validation against these measurements is reported. Without a direct correlation between the predicted score and recorded motion on held-out subjects, the claim that the score is motion-specific, rather than a generic quality score, is underdetermined.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains a 3D convolutional neural network (SFCN2) to regress a scalar head-motion score from T1-weighted structural MRI. Training labels are RMS deviation values computed from synthetic rigid-body motion applied to clean HBN volumes using TorchIO. The model is evaluated on a synthetic test set, on MR-ART manual three-level motion grades, and on 15 real datasets by correlating predicted motion with FreeSurfer cortical thickness and age. The authors report a synthetic test R² of 0.94, a Spearman correlation of 0.71 with MR-ART labels, significant thickness–motion associations in 12/15 datasets after FDR correction, and age correlations consistent with prior literature.","tokens_in":14151,"tokens_out":4273,"duration_ms":41874,"significance":"If the central claim holds—that a model trained on synthetic motion artifacts yields a motion-specific score transferable across scanners and protocols—the work would provide a scalable, retrospective motion-assessment tool for structural MRI, with direct relevance to population studies of neurodevelopment and psychiatric disorders. Strengths include the unusually broad external evaluation (14 independent datasets plus one held-out site from the training cohort), explicit FDR correction across 525 fitted models, and publicly released code, model weights, manual labels, and an open-source CLI tool, all of which support reproducibility. The paper also demonstrates that simulated motion reproduces the known negative thickness–motion relationship in synthetic data, a useful internal consistency check.","major_comments":[{"comment":"The abstract states that the method achieves \"a representative R² = 0.65 versus manual labels,\" but the body reports only a Spearman rank correlation of 0.71 against MR-ART's three-level grades; no R² value is reported in Section 3.1.2 or anywhere else in the results. The authors must either report the actual R² and its computation (e.g., against which continuous variable, since MR-ART labels are ordinal) or correct the abstract, because the abstract's headline performance claim is currently unsupported by the presented evidence.","section":"Abstract; §3.1.2"},{"comment":"No validation is reported against continuous ground-truth motion, despite the availability of vNavs data. Section 2.1 states that CBIC and CUNY include T1w volumes acquired with vNavs-based prospective correction, and these same sites are used to generate synthetic training data. Yet the paper never compares predicted motion scores against the recorded vNavs motion estimates on those real volumes. Such a comparison would directly test whether the score tracks real in-scanner motion rather than generic image quality, and it is feasible with data already in the authors' possession. Without it, the claim that the model estimates head motion (Conclusion, Section 5) rather than a general artifact index is underdetermined.","section":"§2.1; §3.1.2"},{"comment":"The RMS deviation label depends on an estimated brain radius Rc, but the paper never specifies how Rc is obtained. Because Rc scales the rotation contribution to the label, different choices change the label distribution and thus the learned score. Please state whether Rc is a fixed constant, derived from the affine-registered images, or taken from a template, and report its value; without this, the label definition is not fully reproducible.","section":"§2.3.2, Eq. (1)"},{"comment":"The real-data validation strategy is indirect: MR-ART's three-level clinical-usability grades conflate motion with overall image quality, and the thickness and age correlations could plausibly be reproduced by any quality index that tracks blur or ghosting. The MR-ART dataset includes paired STAND/HM1/HM2 scans with known instructed motion levels, but the manuscript only reports the Spearman correlation against the three-level grade and does not report whether the predicted score discriminates the instructed motion condition in a paired analysis (e.g., STAND vs. HM1 vs. HM2). Adding such an analysis, or the vNavs comparison above, is necessary to support the conclusion that the model learns \"purely synthetic motion artifacts\" rather than a general quality score.","section":"§3.1.2; §3.2.2"}],"minor_comments":[{"comment":"The word \"generalisabiliy\" is misspelled and should be \"generalizability.\"","section":"§2.2"},{"comment":"The sentence \"These findings agree are in line with the literature\" contains a duplicated verb; it should read \"These findings are in line with the literature.\"","section":"§3.2.3"},{"comment":"The caption states \"bottom left and top bottom\"; the last phrase should be \"bottom right\" or similar.","section":"Figure 6 caption"},{"comment":"Section 2.4 says the network is trained by minimizing KL divergence, while Section 2.5 says the validation selection criterion is Jensen-Shannon divergence; please clarify which objective is used for training and which for model selection.","section":"§2.4; §2.5"},{"comment":"The variable name \"emotion\" is easily misread as \"emotion\" and should be renamed to something like \"e_motion\" or \"target_motion\" for clarity.","section":"§2.3.2"},{"comment":"The claim \"This is the first attempt at correcting motion-related biases...\" is stronger than what the paper demonstrates and could be softened to \"a method for estimating motion that can be used to correct for motion-related biases.\"","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"The abstract/body discrepancy on R² is a straightforward fix but should be caught before publication. The more substantive issue is the absence of any direct validation against continuous motion ground truth; given that the authors already have vNavs data in their training sites, adding this validation is likely feasible within the manuscript's scope and would substantially strengthen the motion-specificity claim. The paper is otherwise a solid empirical study with unusually broad external evaluation and good reproducibility practices."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper: it's a serious methods submission on estimating head motion from structural MRI using a CNN trained on simulated motion. The concrete contributions are a slightly deeper SFCN variant (two convs per block), a specific augmentation pipeline, and validation across 15 datasets—one held-out training site, MR-ART with manual ratings, and 11 OpenNeuro datasets. That scale is genuinely unusual for this kind of work, and the authors release code, model weights, data splits, and manual labels. The validation is mostly careful: they apply FDR correction over 525 thickness models, report per-dataset coefficients, and show age correlations that match prior findings.\n\nThe soft spots are real but not fatal. The abstract says R²=0.65 versus manual labels; the body reports a Spearman rank correlation of 0.71 on MR-ART and never shows that R². That is an inconsistency the authors need to fix. More substantively, the claim that the score measures motion rather than generic image quality is underdetermined. The only direct validation is against MR-ART's 3-level clinical usability grade, which is a coarse, quality-based label. The authors have HBN vNavs data—which provides continuous, recorded motion—yet never report a validation against it. That is the single most useful addition a revision could make. The thickness and age correlations are indirect: any blur/ghosting-sensitive index would likely reproduce them, because those artifacts also bias FreeSurfer and are more common in children.\n\nI'd also like to see a direct comparison to Pollak et al. 2023, whose real-motion depth-camera labels give a proper regression target. Without that, the advantage of this synthetic pipeline is asserted, not demonstrated.\n\nThe paper is honest about its own limits: the Discussion lists \"more accurate artifact simulators\" as future work, which is the right concession. The synthetic-to-real transfer is the load-bearing premise, and the paper doesn't fully prove it.\n\nBottom line: this deserves peer review. A good referee can push on the vNavs validation, the abstract/body metric discrepancy, and the Pollak baseline. It's not a desk reject, and it's worth discussing in reading group if you're working on image QC.","headline":"A well-validated synthetic-to-real motion score, but the claim that it measures motion rather than generic image quality is only partially supported because the authors skip the continuous vNavs ground-truth they already hold.","tokens_in":14736,"tokens_out":2378,"would_cite":true,"duration_ms":23704,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A motion score learned on synthetic artifacts transfers to real retrospective MRI scans.","keywords":["head motion estimation","structural MRI","cortical thickness","synthetic motion artifacts","deep learning","quality control","MPRAGE","convolutional neural network"],"falsifier":"Scan a cohort with simultaneous optical motion tracking during MPRAGE acquisition, run the model on each volume, and compute the Spearman correlation between predicted score and tracked motion magnitude; if it approaches zero or reverses sign, the synthetic-to-real transfer claim is refuted. A cheaper partial test: on MR-ART's paired clean and deliberately moved scans from the same subject, check that the model consistently assigns higher motion to the moved scan in every individual, not just on average.","tokens_in":13689,"feed_emoji":"🧠","tokens_out":8670,"duration_ms":72837,"temperature":0.7,"pith_summary":"Head motion during a five-minute structural MRI scan is unavoidable and systematically biases automated measurements like cortical thickness, but retrospective scans usually lack an objective motion label. This paper claims that a 3D convolutional network trained only on synthetically motion-corrupted brain volumes can assign a continuous motion score to real retrospective scans, without prospective motion correction or special hardware. On an independent dataset with manual ratings the score reaches a representative R² of 0.65, and the predicted motion is significantly associated with FreeSurfer mean cortical thickness in 12 of 15 datasets, with the same negative direction previously reported for real motion. Predicted motion also tracks the known age pattern: more motion in children and in older adults. If correct, the method gives researchers a cheap, scalable way to enter motion as a covariate or quality metric in studies of brain anatomy.","feed_headline":"Synthetic motion trains a score that works on real brain scans","feed_subtitle":"Trained only on corrupted clean scans, the model ranks real head motion and flags thickness bias in 15 datasets.","key_machinery":"The mechanism is the synthetic motion-augmentation plus RMS labeling loop. Clean volumes are corrupted by sampling a sequence of rigid-body transforms, applying them in image domain, and blending the resulting k-space lines; the RMS deviation of the composed transform (Jenkinson's metric) provides a continuous, objective label for every corrupted volume. The network learns to regress this label from the corrupted image alone, and heavy dropout plus noise and contrast augmentation is what makes the regressor survive the shift to real scanners. The architecture details matter less than this labeling loop: a two-convolution SFCN with 50-bin soft classification and KL-divergence training.","core_discovery":"The authors build SFCN2, a lightweight 3D fully convolutional network that outputs a probability distribution over 50 motion bins and is trained by minimizing KL divergence to a Gaussian centered on the true motion label. The labels come from an entirely synthetic pipeline: 449 manually screened 'clean' T1 volumes from the Healthy Brain Network are repeatedly corrupted by TorchIO's k-space rigid-body motion augmentation, and each corrupted volume is labeled with the RMS deviation of the net transformation, an objective geometric quantity. Without any fine-tuning on real artifacts, the model ranks the MR-ART manual motion grades with Spearman 0.71, and its score shows a significant negative association with mean left-hemisphere cortical thickness in 12 of 15 held-out datasets (FDR-corrected), matching the known thickness-reducing effect of motion. The score also reproduces the established age pattern, with higher predicted motion in younger and older participants. The paper's central claim is that this synthetic-to-real transfer is genuine motion estimation, not just image-quality scoring.","pith_inferences":["A natural next step the paper does not take is to threshold the continuous score into a formal inclusion or exclusion criterion; the current validation uses the score as a continuous covariate, not as a QC gate.","The same synthetic-labeling loop could be extended to other artifact families (e.g., susceptibility, flow, ghosting) or other morphometric outputs (surface area, volume) by training separate regressors, since the label generation is modular.","The transfer claim would be strengthened by a direct comparison against measured motion from vNavs or an optical tracker on the same subjects; the paper only validates against manual ratings and indirect thickness and age associations.","Because the score is derived from a single volume, its per-subject reproducibility across repeated scans of the same person is untested; a test–retest study would clarify how much of the score is motion versus acquisition-specific noise."],"forward_implications":["Researchers can add the predicted motion score as a covariate in statistical models of cortical thickness, absorbing a bias that otherwise distorts group differences in pediatric and clinical populations.","Quality control becomes an automated, continuous rating on already-acquired MPRAGE-like scans, replacing or complementing noisy manual review.","The method transfers across scanner manufacturers and protocol variants without retraining, as shown by the significant associations on GE and Siemens datasets.","Individual FreeSurfer estimates can be confidence-weighted or flagged when predicted motion is high, before downstream analyses are run.","Studies whose subjects move more (children, ADHD, schizophrenia) can now quantify that motion post hoc, addressing a confound previously requiring prospective hardware."],"supporting_citations":[{"why":"Supplies the k-space motion artifact augmentation method used to corrupt clean training volumes.","marker":"Shaw et al., 2019"},{"why":"Defines the RMS deviation metric that labels each synthetic corruption with a continuous motion score.","marker":"Jenkinson, n.d."},{"why":"Provides the MR-ART dataset of paired clean and motion-corrupted scans with manual ratings used to validate real-motion prediction.","marker":"Narai et al., 2022"},{"why":"Establishes the known negative effect of head motion on FreeSurfer cortical thickness estimates that the paper reproduces.","marker":"Reuter et al., 2015"},{"why":"Demonstrates systematic thickness and curvature biases from subtle motion, the pattern the predicted score is checked against.","marker":"Alexander-Bloch et al., 2016"},{"why":"Introduces the SFCN backbone architecture that the SFCN2 network extends.","marker":"Peng et al., 2021"},{"why":"Prior use of SFCN for head-motion estimation with camera-derived labels; the paper builds on this architecture choice.","marker":"Pollak et al., 2023"},{"why":"The Healthy Brain Network dataset supplies the clean training volumes for synthetic corruption.","marker":"Alexander et al., 2017"},{"why":"FreeSurfer is the tool that computes the cortical thickness measurements used in validation.","marker":"Fischl, 2012"}],"fun_headline_variants":["Synthetic motion trains MRI model to score real head motion","AI trained on fake motion flags real cortical thickness bias","Synthetic-to-real motion scoring for brain MRI scanners","Neural net learns from synthetic motion, predicts real artifacts","Motion scoring from synthetic training generalizes across sites"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The approach stands or falls on the assumption that TorchIO's synthetic rigid-body k-space corruptions look enough like real head motion in MPRAGE scans that a network trained on them will score real scans correctly across scanner brands, protocols, and age groups.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic motion trains MRI model to score real head motion","AI trained on fake motion flags real cortical thickness bias","Synthetic-to-real motion scoring for brain MRI scanners","Neural net learns from synthetic motion, predicts real artifacts","Motion scoring from synthetic training generalizes across sites"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000824,"raw_usage":{"total_tokens":3616,"prompt_tokens":970,"completion_tokens":2646,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":586,"completion_tokens_details":{"reasoning_tokens":2569}},"tokens_in":586,"tokens_out":2646,"duration_ms":17224,"temperature":1.0,"reasoning_tokens":2569,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:38:06.525750+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Scan a cohort with simultaneous optical motion tracking during MPRAGE acquisition, run the model on each volume, and compute the Spearman correlation between predicted score and tracked motion magnitude; if it approaches zero or reverses sign, the synthetic-to-real transfer claim is refuted. A cheaper partial test: on MR-ART's paired clean and deliberately moved scans from the same subject, check that the model consistently assigns higher motion to the moved scan in every individual, not just on average.","supporting_citations":[],"review_version":1}