{"id":"7bf39734-3707-4421-af6f-cb0ac90dc79e","arxiv_id":"2608.07092","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Stochastic cortical self-reconstruction, trained on UK Biobank, transfers to a Chinese lifespan cohort, with a fine-tuned Spherical UNet achieving 0.848 average pairwise AUC for CN versus MCI versus AD.","lead":"A brain-atrophy mapping method trained only on UK Biobank data was tested on a separate Chinese cohort spanning ages 4 to 85. It detected Alzheimer's and mild cognitive impairment across populations, and fine-tuning on Chinese data gave the best diagnostic accuracy, with average AUC of 0.848.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 0.848-vs-0.815 SUNet advantage and the strong-transfer claim rest on point estimates from 60 subjects per group with no confidence intervals or group-balance reporting; sampling noise or demographic confounds could reverse the ordering.","rationale":"The reader's conditional verdict correctly identifies the central weakness: the headline AUC ordering and the cross-population transfer claim depend on small, possibly unrepresentative subsamples summarized by point estimates. My reading of the full text finds the same load-bearing gap. The methods are clearly described, the architecture comparison is internally consistent, and the reconstruction-error results provide some independent support for transferability, but the downstream clinical claim—the reason the paper is important—is not statistically secured. The fine-tuning advantage over direct application for SUNet is small (0.848 vs 0.815) and concentrated in one pairwise comparison; without intervals, this could easily be sampling noise. The lack of reported demographic and scanner balance across diagnostic groups adds a second route by which the AUC could be inflated independently of atrophy. Since the source code and Chinese dataset are not available for independent recomputation, the appropriate verdict remains CONDITIONAL, not ACCEPT or REJECT. I therefore leave the reader's verdict unchanged and propose a single decisive statistical re-analysis as the concrete test.","tokens_in":9550,"tokens_out":5755,"duration_ms":59235,"concrete_test":"Recompute Table 2 with nonparametric bootstrap (10,000 resamples over the 180 test subjects) to obtain 95% confidence intervals for each pairwise AUC and for the SUNet direct-vs-fine-tuned difference, plus DeLong tests; also report age and sex distributions per diagnostic group. If the direct-vs-fine-tuned difference interval includes zero, the 'fine-tuning is most effective' conclusion should be softened; if demographic imbalance is found, rerun the AUC analysis adjusted for age and sex.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claims—'robust detection of cortical atrophy' and 'highest discriminative performance by fine-tuned SUNet (0.848), followed closely by UKB-trained SUNet (0.815)'—rest on Table 2, which reports pairwise AUCs computed on 60 CN, 60 MCI, and 60 AD subjects with no confidence intervals, significance tests, or multiple-comparison correction across eight configurations. The fine-tuning advantage over direct application is driven mainly by the CN|MCI pair (0.787 vs 0.718; +0.069), whereas CN|AD and MCI|AD change by only +0.013 and +0.015. With n=60 per group, a 0.069 AUC shift is within plausible sampling variability; DeLong or bootstrap intervals would likely overlap substantially. Section 3.4 states only that CN controls are 'age-matched (46–85 years)', but does not report age, sex, education, scanner, or disease-severity distributions per group. If CN differ from MCI/AD in age or sex within that range, or if site/scanner composition differs across diagnostic groups, the AUCs—especially CN|MCI—could reflect demography rather than atrophy. Because fine-tuning is trained on Chinese healthy data, it may also absorb site/scanner biases, inflating apparent discrimination without improving true disease-related transfer. The claim that 'every configuration remains clearly above chance' is likewise unsupported without intervals.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates the transfer of Stochastic Cortical Self-Reconstruction (SCSR), a neural-network-based normative model of cortical thickness, from a UK Biobank-trained model to an independent Chinese cohort spanning ages 4–85. Eight configurations are compared (MLP vs. Spherical UNet; direct application, fine-tuning, scratch training, and joint training). The authors report reconstruction MAE across the lifespan and pairwise AUCs for CN/MCI/AD discrimination from SCSR Z-scores in an AD ROI. They claim robust atrophy detection across all configurations, with the fine-tuned SUNet achieving the best average AUC (0.848) followed by the UKB-trained SUNet (0.815), and strong cross-population transferability based on low reconstruction errors.","tokens_in":9804,"tokens_out":3901,"duration_ms":36719,"significance":"If the empirical claims hold, the paper provides valuable evidence that a self-reconstruction-based normative model can transfer across populations and detect AD/MCI-related atrophy without target-domain training, with practical implications for deploying such models on new international cohorts. The experimental design is clean: four adaptation strategies are compared systematically under two architectures, and the public code repository supports reproducibility. The main weakness is that the central quantitative claims rest on point estimates from 60 subjects per diagnostic group without confidence intervals, significance tests, or adjustment for potential confounds; these issues are fixable and do not invalidate the underlying approach.","major_comments":[{"comment":"The central claims that the fine-tuned SUNet achieves the highest average AUC (0.848) and that all eight configurations are 'clearly above chance' rest on point estimates from 60 subjects per diagnostic group with no confidence intervals, significance tests, or multiple-comparison correction. The fine-tuning advantage over direct application is driven mainly by the CN|MCI pair (0.787 vs. 0.718, +0.069); with n=60 per group, this difference is within plausible sampling variability (DeLong or bootstrap intervals would likely overlap substantially). Please report confidence intervals for all AUCs, test differences between configurations (e.g., DeLong or bootstrap with correction across the eight configurations), and assess the power of the current sample size.","section":"Section 4.2, Table 2"},{"comment":"The manuscript states only that CN controls are 'age-matched (46–85 years)' and does not report per-group distributions of age, sex, education, scanner/site, or disease severity for the 60 CN, 60 MCI, and 60 AD subjects. If these variables differ across diagnostic groups, the reported AUCs—particularly CN|MCI (0.787 for the fine-tuned SUNet)—could reflect demographic or scanner confounds rather than atrophy. Please report these distributions for each diagnostic group and, if possible, provide analyses adjusted for age and sex, or at least demonstrate balance across groups.","section":"Section 3.1 and Section 3.4"},{"comment":"The evaluation assumes that FreeSurfer thickness values are comparable across the UK Biobank and Chinese datasets without cross-site harmonization, even though scanner hardware and acquisition protocols differ between the populations. This assumption affects both the cross-population transferability claim and the comparison of training strategies, since fine-tuning on Chinese data may absorb site-specific biases. Please discuss or test this directly, for example by reporting scanner distributions per group, adding a site/scanner covariate in the analysis, or performing a sensitivity analysis on a subset matched for scanner characteristics.","section":"Sections 3.1 and 4.2"},{"comment":"The abstract and conclusion claim 'strong cross-population transferability' partly based on reconstruction errors that are described as 'comparable' across populations, yet Table 1 shows a 40% relative increase in MAE for the direct MLP on Chinese data (0.359 mm) versus UKB validation (0.256 mm), and no statistical comparisons across cohorts or age brackets are provided. Please provide statistical comparisons (e.g., confidence intervals or tests for age-bracket differences) or temper the claim to reflect the observed magnitudes.","section":"Section 4.1 and Table 1"}],"minor_comments":[{"comment":"The validation set used to estimate sigma is not specified separately for each of the eight configurations; please clarify whether a distinct validation set was used for each training strategy and whether sigma differs across configurations.","section":"Section 3.2, Eq. (3)"},{"comment":"The Z-score color scale is not shown in either figure; adding a colorbar and specifying the displayed hemisphere and anatomical views would improve interpretability of the atrophy maps.","section":"Figures 3 and 4"},{"comment":"The text states '640 healthy scans for training/finetuning and 160 healthy scans for validation' but does not clarify whether the 139 lifespan test subjects and the 60 CN controls come from the same source pool or whether they overlap with the validation set; please make the data splits explicit.","section":"Section 3.1"},{"comment":"The related work on transfer learning would benefit from a comparison with recent cross-cohort normative modeling benchmarks in the neuroimaging literature, which would help position the contribution beyond the cited segmentation and shape classification studies.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The paper's abstract and conclusion make stronger claims than the statistical evidence supports. The primary fix is straightforward: add confidence intervals and significance tests for the AUC results, report per-group demographics and scanner distributions, and temper the transferability claim accordingly. The architecture effect (SUNet outperforming MLP) is consistently observed and is a useful contribution. I do not see grounds for rejection, but the current manuscript is not yet sufficient for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take on arXiv:2608.07092. This is the first independent-cohort evaluation of the SCSR normative model, originally trained on UK Biobank, applied to a Chinese lifespan dataset. That alone makes it a useful paper for the normative modeling subfield. The experimental matrix is clean: four adaptation strategies (direct, fine-tuned, scratch, joint) crossed with two backbones (MLP, SUNet), evaluated both on reconstruction error and on downstream atrophy AUC in AD/MCI. The public code and unchanged hyperparameters from the original SCSR paper make the experiments reproducible in principle.\n\nWhat holds up: the reconstruction error results across age brackets are the strongest part. Direct application of the UKB-trained model to Chinese data, including ages 4–20 that are entirely outside the UKB training range, gives errors comparable to models that saw Chinese data. That is a genuine and practically useful finding. The MLP-overfits-on-small-data story (scratch MLP much worse than SUNet scratch) is also coherent and supported by the parameter counts. The discussion is measured; the authors explicitly note that fine-tuning helps SUNet more than MLP and that joint training might benefit from re-weighting.\n\nThe soft spot is exactly where the stress-test note lands: the headline AUC numbers are point estimates from 60 subjects per group with no confidence intervals or significance tests across eight configurations. The fine-tuned SUNet advantage over direct SUNet (0.848 vs 0.815) is driven mostly by the CN|MCI pair (0.787 vs 0.718); with n=60 that shift is within plausible sampling variability, and the paper does not report group-wise age, sex, scanner, education, or severity distributions beyond \"age-matched 46–85.\" The sentence claiming every configuration is \"clearly above chance\" is not backed by intervals. This matters because the abstract's \"robust detection\" phrasing rests on those point estimates. I would not call the paper wrong—the direction of effects is plausible and the direct-transfer baseline at 0.815 is strong enough that the broad conclusion likely survives—but the specific fine-tuning claim should not be taken at face value until intervals or DeLong tests are added.\n\nCitation pattern is fine: self-citation of [20] and [27] is appropriate here since the method and the Chinese dataset come from those papers.\n\nWho this is for: researchers doing normative modeling, cortical surface transfer learning, or practical atrophy mapping across populations. It deserves a serious referee and likely a conditional accept after the authors add error bars or significance testing, describe the group composition, and soften the \"clearly above chance\" wording. I'd send it to review.","headline":"A useful first external transfer study of SCSR to a Chinese cohort; the direct-transfer baseline is the solid result, while the fine-tuning AUC advantage is a point estimate that needs error bars before it carries weight.","tokens_in":10368,"tokens_out":2416,"would_cite":true,"duration_ms":23338,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SCSR, a cortical-thickness reference model trained only on UK Biobank adults, detects Alzheimer's and MCI atrophy in an independent Chinese cohort, with fine-tuning on local healthy scans raising the best average AUC from 0.815 to 0.848.","keywords":["normative modeling","cortical surfaces","transfer learning","Alzheimer's disease","cortical atrophy","mild cognitive impairment","stochastic cortical self-reconstruction","spherical UNet"],"falsifier":"Compute the average pairwise AUC for all eight configurations on a larger, scanner- and demographically matched Chinese test set with bootstrapped confidence intervals; the paper's transfer claim holds up only if the directly applied UK-trained SUNet's 0.815 remains well above chance and statistically comparable to the fine-tuned 0.848, and collapses if the gap disappears or the direct model falls to chance under re-sampling.","tokens_in":9332,"feed_emoji":"🧠","tokens_out":9191,"duration_ms":77556,"temperature":0.7,"pith_summary":"SCSR is a reference model that learns the typical shape of the healthy cortex from a population, then gives each person a personalized healthy baseline by repeatedly reconstructing their own thickness map. The paper asks whether such a model, built only on healthy UK adults aged 45–82, transfers to an independent Chinese cohort aged 4–85 and to Chinese Alzheimer's disease (AD) and mild cognitive impairment (MCI) patients. It claims yes: without any adaptation, the SUNet version separates cognitively normal controls from MCI and AD with average pairwise AUC (area under the ROC curve) of 0.815, and fine-tuning on 640 Chinese healthy scans raises that to 0.848, the best of eight configurations. Reconstruction error stays low across the lifespan even in age groups absent from training, which the paper reads as strong cross-population transferability.","feed_headline":"Model trained on UK data finds Alzheimer's atrophy in Chinese scans","feed_subtitle":"Fine-tuning lifts the detection score to 0.848; using the model as-is still hits 0.815.","key_machinery":"The central object is stochastic cortical self-reconstruction (SCSR), a normative reference model that builds a personalized healthy baseline from a subject's own cortex rather than from demographic covariates. During training a neural network learns to predict masked vertices of a cortical thickness map from a randomly sampled 20% of its vertices; at test time the sampling is repeated $m=100$ times and the per-vertex 95th centile across repetitions becomes the healthy reference $R$, yielding Z-scores $Z=(Y-R)/\\sigma$ with $\\sigma$ estimated from validation residuals. The paper runs this machinery with two backbones: a 20M-parameter multilayer perceptron with no spatial structure and a 1.7M-parameter Spherical UNet that convolves directly on an icosahedral mesh, and compares direct application, fine-tuning, training from scratch, and joint training on UK and Chinese data.","core_discovery":"The central empirical claim is that SCSR's healthy cortical reference transfers across international populations and across architectures. On the Chinese cohort, the directly applied UK-trained SUNet achieves an average pairwise AUC of 0.815 for CN vs MCI, CN vs AD, and MCI vs AD using mean Z-scores in the AD ROI, and fine-tuning the same model on the 640 Chinese healthy training scans improves this to 0.848. All eight configurations—two backbones times four adaptation strategies—stay above chance, and the SUNet backbone outperforms the much larger MLP for this downstream detection task while the MLP shows lower raw reconstruction error. The paper interprets this as evidence that population-specific adaptation is a modest, optional gain rather than a prerequisite, and that the choice of adaptation strategy should depend on the goal: fine-tuning for diagnostic discrimination, joint training for reconstruction fidelity across populations.","pith_inferences":["A testable extension the paper leaves open is deploying the directly applied UK-trained model in a site with no local healthy training data, to see whether the 0.815 AUC generalizes to other scanners and populations.","The 60-subject-per-group test set means the 0.033 AUC gap between direct and fine-tuned SUNet could be sampling noise; re-estimating with confidence intervals is an editorial caution, not a paper claim.","Because the paper does not harmonize between acquisition sites, a version of this experiment with site-harmonized thickness maps would show whether the transfer signal is biological or partly scanner-specific.","Re-weighting the joint training so the small Chinese cohort counts more than the large UK cohort could close the MLP's joint-training deficit, a possibility the paper mentions but does not test."],"forward_implications":["A UK-trained SCSR model can be applied to a new international cohort without retraining and still yield usable atrophy detection, with reconstruction error comparable to locally trained models.","For maximizing diagnostic separation between CN, MCI, and AD, fine-tuning the SUNet backbone on local healthy data is the best of the four strategies tested, ahead of joint training and training from scratch.","For maintaining reconstruction fidelity across both source and target populations, joint training on UK and Chinese data is the best configuration.","The SUNet backbone transfers better than a twelve-times-larger MLP: it stays accurate when trained from scratch on 640 subjects, while the scratch MLP overfits.","Age groups entirely missing from UK training data, such as children aged 4–20, show only mildly higher reconstruction error, so the reference generalizes beyond the training age range."],"supporting_citations":[{"why":"Defines SCSR and the original UK-trained models and configuration this paper transfers.","marker":"[20]"},{"why":"Supplies the UK Biobank imaging cohort used for pretraining and validation.","marker":"[11]"},{"why":"Supplies the Chinese healthy lifespan and AD/MCI dataset that is the target population.","marker":"[27]"},{"why":"Defines the Spherical UNet architecture used as the SUNet backbone.","marker":"[26]"},{"why":"Defines the AD region of interest whose mean Z-scores produce the reported AUC values.","marker":"[17]"},{"why":"Provides the FreeSurfer pipeline that produces the vertex-wise cortical thickness maps.","marker":"[8]"}],"fun_headline_variants":["UK-trained AI spots Alzheimer's atrophy in Chinese brains","Cross-border brain map: UK model detects Alzheimer's in China","Alzheimer's atrophy detection crosses continents via transfer learning","Fine-tuned AI boosts Alzheimer's detection across populations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 60 CN, 60 MCI, and 60 AD scans used for the AUC comparison are representative and internally balanced, so the reported fine-tuning advantage is not sampling noise.","fun_headline_variants_meta":{"raw":{"variants":["UK-trained AI spots Alzheimer's atrophy in Chinese brains","Cross-border brain map: UK model detects Alzheimer's in China","Alzheimer's atrophy detection crosses continents via transfer learning","Fine-tuned AI boosts Alzheimer's detection across populations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000778,"raw_usage":{"total_tokens":3474,"prompt_tokens":1014,"completion_tokens":2460,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":2395}},"tokens_in":630,"tokens_out":2460,"duration_ms":14882,"temperature":1.0,"reasoning_tokens":2395,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:04:10.558993+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the average pairwise AUC for all eight configurations on a larger, scanner- and demographically matched Chinese test set with bootstrapped confidence intervals; the paper's transfer claim holds up only if the directly applied UK-trained SUNet's 0.815 remains well above chance and statistically comparable to the fine-tuned 0.848, and collapses if the gap disappears or the direct model falls to chance under re-sampling.","supporting_citations":[{"cited_title":"Medical Image Analysis107, 103788 (Jan 2026)","cited_arxiv_id":null,"evidence_quote":"Defines SCSR and the original UK-trained models and configuration this paper transfers."},{"cited_title":"Nature communications11(1), 2624 (2020)","cited_arxiv_id":null,"evidence_quote":"Supplies the UK Biobank imaging cohort used for pretraining and validation."},{"cited_title":"Nature Neuroscience29(2), 420–434 (Dec 2025)","cited_arxiv_id":null,"evidence_quote":"Supplies the Chinese healthy lifespan and AD/MCI dataset that is the target population."},{"cited_title":"In: Information Processing in Medical Imaging","cited_arxiv_id":null,"evidence_quote":"Defines the Spherical UNet architecture used as the SUNet backbone."},{"cited_title":"NeuroImage: Clinical11, 802–812 (2016)","cited_arxiv_id":null,"evidence_quote":"Defines the AD region of interest whose mean Z-scores produce the reported AUC values."}],"review_version":1}