{"id":"9de53194-18e6-459b-94a7-82994633fe9c","arxiv_id":"2505.03715","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"DISARM++ harmonizes multi-scanner T1-weighted brain MRI into a scanner-free or reference-scanner domain and reports improved downstream prediction accuracy over STGAN and IGUANe.","lead":"This paper presents DISARM++, a deep learning model that transforms brain MRI scans from any scanner into a scanner-free version or into the style of a chosen reference scanner. The authors report that this harmonization improves age prediction, Alzheimer's disease classification, and diagnosis prediction compared with existing methods, without requiring skull-stripping.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Scanner-free output may retain scanner contrast because the scanner-free loss (Eq. 7) only enforces that the re-encoded scanner effect of two same-noise outputs match, not that the output lies in a genuinely scanner-free distribution.","rationale":"The reader's weakest_assumption is the load-bearing disentanglement assumption, and I agree that it is the central concern. The strongest evidence in the paper is the reduction of voxel-intensity divergence between scanners (JSD 0.17 to 0.009), but this evidence only demonstrates that the output distributions are close to each other, not that they are close to a scanner-free distribution in any absolute sense. The scanner-free loss L_sf (Eq. 7) is a consistency loss between two same-noise outputs of two different source images in the re-encoded scanner-effect space; it does not anchor the scanner-free output to the Gaussian prior or to any canonical reference distribution. Thus the model could in principle collapse all scanners to a new shared but scanner-affected appearance (e.g., a learned mean scan), while still satisfying the cycle, identity, and classification losses, provided the brain encoder retains enough information to reconstruct. The downstream experiments do not resolve this, since they show a difference between DISARM++ and alternatives but do not establish that the residual is biologically complete rather than a loss of scanner-related anatomical signal. The paper does not test scanner information leakage in z_b or in the scanner-free images directly, and the code is not actually linked, which limits the ability to check whether the reported scanner-free inference (using N(0,1) and c0) is implemented as described. I recommend keeping the verdict CONDITIONAL, because the direct harmonization and traveling-subject results are plausible and likely reproducible, but the core scanner-free identity is not sufficiently validated. My concrete test would settle whether the scanner-free output is genuinely scanner-free by measuring residual scanner information, and would also test the alternative explanation that the reported improvement is essentially global intensity normalization.","tokens_in":36608,"tokens_out":3086,"duration_ms":23935,"concrete_test":"Train a simple 3D linear/logistic scanner classifier on the original test images from the 10 test scanners; then evaluate the same classifier on (a) raw images, (b) scanner-free DISARM++ outputs, (c) IGUANe/STGAN outputs using identical preprocessing. If scanner classification accuracy on DISARM++ scanner-free outputs remains far above chance (or above the level achievable on images that have been contrast-normalized to a common template), the scanner-free space is not actually scanner-free and the harmonization claim fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The scanner-free space F is the load-bearing artifact: all 'scanner-free' downstream claims assume G(z_b, epsilon, c0) produces images whose contrast is independent of the acquisition scanner while preserving anatomy. The training signal for this space is the scanner-free loss L_sf (Eq. 7) plus a KL prior on the scanner encoder (Eq. 1/8). L_sf only compares the re-encoded scanner-effect latents Es(G(z_b_xk, eps, c0)) and Es(G(z_b_xh, eps, c0)) for two same-noise outputs, pulling them together but not fixing them to the N(0,1) prior or to any canonical appearance. Nothing forces Es(G(z_b, eps, c0)) to equal eps, or forces the scanner-free images from different source scanners to match in voxel/intensity space. The reported JSD=0.009 is computed on pairwise scanner-mean intensity distributions and an AD-test null acceptance, but this only shows within-distribution spread is small after harmonization, not that the harmonized distribution is scanner-free rather than a new scanner-like domain (e.g., a learned blend). The diagnostic failure mode: if scanner and anatomy are entangled, the brain encoder can partially hide scanner info, and the cycle/adversarial losses can be satisfied while the 'scanner-free' output still contains source-scanner intensity texture or, worse, discards anatomy that correlates with scanner. The paper provides no direct measurement of scanner information remaining in the scanner-free output (e.g., scanner classification accuracy on the scanner-free images or in the latent z_b), and no direct check that the encoder's representation is anatomy-complete. The AD/diagnosis sections also carry known confounds, but the scanner-free identity is the more fundamental issue: it supports the entire 'beyond feature standardization' claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes DISARM++, an unsupervised image-to-image translation model for harmonizing 3D T1-weighted MRI. The model assumes each image is generated from a brain-structure latent B and a scanner-effect latent S; after training with cycle-consistency, adversarial, classification, KL, and a new scanner-free loss, inference maps a new image either to one of the training scanners or to a 'scanner-free' space by substituting random Gaussian noise for the scanner effect. The authors train on 701 healthy-control images from five scanners and evaluate on ten scanners (including scanners unseen in training), on traveling subjects, and on four downstream tasks: age prediction, inter-scanner volume variability, AD-versus-healthy classification, and diagnosis prediction. They report strong harmonization improvements over STGAN and IGUANe (e.g., pairwise JSD dropping from about 0.17 to 0.009), better traveling-subject SSIM, and better downstream results (age R2≈0.60, AD accuracy≈0.86, diagnosis AUC≈0.95).","tokens_in":36875,"tokens_out":8085,"duration_ms":79854,"significance":"The practical goal is important: direct voxel-level harmonization that does not require skull-stripping and generalizes to unseen scanners would be useful for multi-center neuroimaging. The paper's strengths include a broad evaluation across multiple public and private datasets, a traveling-subject test, comparison with two strong baselines, and several downstream tasks; this is a substantial amount of evidence for the core harmonization effect. The traveling-subject SSIM improvements and the reduction in inter-scanner volume variability (Table 9) are the most convincing parts of the empirical package. However, the 'scanner-free' claim is not independently established by the current metrics, and the headline diagnosis AUC appears to be an in-sample estimate. The paper is therefore promising but needs additional validation before the central claims can be accepted.","major_comments":[{"comment":"The 'scanner-free' property is the load-bearing claim of the paper, but the scanner-free loss in Eq. (7) only enforces Es(hat{x}_{k->f}) ≈ Es(hat{x}_{h->f}) for two outputs generated from the same Gaussian noise epsilon; it does not anchor these latents to the prior N(0,1) or to any canonical reference appearance. Nothing in Eq. (8) as described applies L_KL or L_lat to the scanner-free encodings, so the generator could map all inputs to a new, arbitrary common domain F that is not 'scanner-free' in any biologically meaningful sense. The JSD/HD/WD reductions in Table 6 and the AD-test acceptance only show that the ten scanner-mean intensity distributions are closer to each other after harmonization; they do not distinguish the claimed scanner-free space from a learned blend. Please add a direct probe of scanner information remaining in scanner-free outputs (e.g., accuracy of a scanner classifier trained on original images and tested on scanner-free images), clarify exactly which losses are applied to Es(hat{x}_{->f}), and show that scanner-free outputs from the same brain but different source scanners agree at the voxel level beyond the traveling-subject SSIM already reported.","section":"Section 3.4, Eq. (7)"},{"comment":"The diagnosis-prediction result is the clearest over-claim. With 41 subjects and a PCA plus logistic-regression pipeline, the single AUC of 0.9459 reported in Section 6.3.4 is almost certainly an in-sample estimate: Section 5.4.4 describes PCA and logistic regression but no cross-validation or held-out test set. An in-sample AUC on 41 subjects can be severely optimistic and cannot be compared with the other methods' AUCs, which appear to be computed in the same way. The authors should re-estimate the AUC with repeated stratified cross-validation, with PCA fitted inside each training fold, and report the distribution over folds plus a permutation null.","section":"Section 5.4.4 / Section 6.3.4"},{"comment":"The AD-versus-healthy classification is confounded: healthy images come from RIN, IXI, and PPMI, while AD images come from NeuroArtP3 and ADNI3 (Section 5.4.3). Scanner and dataset are therefore partially predictive of the label even before any biology is considered. The improved accuracy after DISARM++ harmonization (0.858 ± 0.03) could reflect removal of some scanner effects, but it does not by itself show preservation of disease-related signal. Please provide scanner-balanced cross-validation folds or report a scanner-label classification accuracy on the same harmonized images as a sanity check.","section":"Section 5.4.3 / Table 10"}],"minor_comments":[{"comment":"The text says 'yij and xij represent the age and volume, respectively', while the model y = β0 + β1 x + u + ε predicts volume from age; the wording appears to be reversed and should be corrected.","section":"Section 5.4.2"},{"comment":"The sentence 'DISARM++ shows significantly better precision, F1 score, and recall, although no significant difference in precision is observed' is self-contradictory; the first occurrence of 'precision' should presumably be 'accuracy' or the sentence should be rephrased.","section":"Section 5.4.3 / Figure 12"},{"comment":"The configuration without Lsf achieves a lower post-harmonization JSD (0.004) than the full model (0.008); the statement that all components are 'crucial' for harmonization is therefore only supported jointly with the structural metrics, and should be phrased accordingly.","section":"Table 5"},{"comment":"There is a notation inconsistency in the method description: Section 3.2 defines the brain discriminator as Db : X → C, processing images, but Eq. (4) evaluates Db on the brain-structure latent zb. If the discriminator actually operates on the latent space B, the architecture description and Figure 1 should be corrected; if it operates on images, Eq. (4) is wrong.","section":"Section 3.2 / Eq. (4)"},{"comment":"The vector c0 used for scanner-free generation is not defined in Section 3.1, where C is defined as one-hot vectors with ∥c∥1 = 1; since c0 appears to be neither one-hot nor a member of C, its meaning should be stated explicitly.","section":"Figures 3 and 5 / Section 3.1"},{"comment":"The age prediction comparison reports R2 and RMSE from 10-fold cross-validation but no paired significance test; given the overlap between the DISARM++ and raw-image standard deviations, a paired test would strengthen the claim of superiority.","section":"Section 5.4.1"},{"comment":"Several typos remain, e.g., 'distict' (Section 5.3.2), 'Hellringer' (Appendix D), and 'F ormula' in Table 2; also, the code link is given only as 'this link' and should be a permanent URL.","section":"General"},{"comment":"Inference for scanner-free harmonization uses random Gaussian noise epsilon; the paper should specify whether the same epsilon is used when harmonizing two scans of the same subject from different scanners, otherwise part of the traveling-subject SSIM comparison may be affected by stochasticity.","section":"Section 3.5"}],"recommendation":"major_revision","confidential_remarks":"The paper is an extended version of a conference paper, and the incremental novelty over DISARM is modest (new loss, attention layers, larger training set). The reported diagnosis AUC of 0.9459 is likely to attract attention, so it is important that the revision addresses the in-sample nature of that estimate. The scanner-free claim also needs a more direct test than the current distributional metrics. There is no indication of misconduct, but the current evidence is not yet sufficient for the central claims to be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Direct and to the point: the image harmonization results are the real contribution. DISARM++ produces visibly uniform intensities across ten scanners (JSD down to ~0.009) and improves traveling-subject SSIM substantially, beating STGAN and IGUANe on those metrics. That alone is a useful engineering result, and the decision to keep skull and full-head volume is a sensible practical choice. The additions over the DISARM conference paper—partial volumes, attention layers, the consistency loss, and the larger training set—are incremental but coherent, and the generalization to scanners not seen in training is demonstrated.\n\nThe soft spots are downstream. The diagnosis AUC of 0.9459 is computed on 41 subjects with no described cross-validation; as reported, it looks like an in-sample logistic-regression fit, which is circular. The AD-vs-healthy classification mixes disease status with data source and scanner, and the age-prediction improvement over raw images (R2 0.60 vs 0.52) is within one standard deviation, so without a proper significance test it does not support an 'outperforms' claim. The code link is absent; the abstract promises a link but no URL appears in the text, which hurts reproducibility.\n\nThe stress-test concern about the scanner-free space is worth taking seriously but not fatal. The scanner-free loss only pulls the re-encoded scanner effects of two same-noise outputs together, which does not by itself guarantee a scanner-invariant distribution. However, the pairwise JSD across ten scanners dropping to roughly 0.009 is direct evidence that, at the intensity-distribution level, the outputs no longer separate by source scanner. What is missing is a check for residual scanner-specific texture: a scanner classifier on the harmonized images, or a probe of the brain latent for scanner information. That is a cheap experiment and should be added.\n\nOverall, this is a paper for neuroimaging method developers and for anyone pooling T1-weighted MRI across sites. It deserves peer review, but with major revisions: fix the diagnosis evaluation, add a proper validation split, report statistical tests for age prediction, release the code, and test scanner-invariance of the scanner-free output directly.","headline":"Harmonization evidence is strong; downstream validation and the scanner-free identity need more work before the SOTA claims hold.","tokens_in":37600,"tokens_out":3543,"would_cite":false,"duration_ms":35353,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DISARM++ removes scanner bias from 3D MRI without skull-stripping and beats two benchmark harmonizers on every test.","keywords":["image harmonization","image-to-image translation","magnetic resonance imaging","scanner-free imaging","disentangled representations","brain age prediction","Alzheimer's disease classification","multi-site MRI"],"falsifier":"On a traveling-subjects set where the same person is scanned on several machines, compute the pairwise similarity of scanner-free outputs across scanners and compare each output against the person's known anatomy, such as a real lesion. The claim fails if scanner-free outputs from the same subject remain scanner-dependent, or if two different subjects' scanner-free outputs become more alike than their raw scans, or if a visible lesion is erased in the scanner-free version.","tokens_in":36308,"feed_emoji":"🧠","tokens_out":6563,"duration_ms":63615,"temperature":0.7,"pith_summary":"The paper proposes that scanner effects in T1-weighted MR images can be removed at the image level rather than at the feature level, by teaching a network to separate what a brain looks like from which scanner acquired it. Its model, DISARM++, can send any scan into a scanner-free reference space, or restyle it to match a chosen training scanner, working on full-head volumes with no skull-stripping and no retraining for new scanner types. The authors test it on healthy controls, traveling subjects scanned on multiple machines, and Alzheimer's disease patients, and report that it outperforms two image-based state-of-the-art methods across harmonization quality, age prediction, inter-scanner volume variability, Alzheimer's classification, and diagnosis prediction. If these results hold, the contribution is a drop-in harmonization step that makes multi-site neuroimaging data comparable while preserving the whole head for downstream analysis.","feed_headline":"MRI harmonizer erases scanner bias, even on unseen scanners","feed_subtitle":"DISARM++ maps scans into a scanner-free space and beats two state-of-the-art harmonizers on downstream tasks.","key_machinery":"The argument is carried by a disentangled latent-space design. A brain encoder $E_b$ maps an image to an anatomical-content vector $z^b$, a variational scanner encoder $E_s$ maps the same image and its scanner label to a scanner-effect distribution whose sampled vector $z^s_i = \\sigma_i \\epsilon + \\mu_i$ encodes the scanner style, and a generator $G$ recomposes the image from both. Training swaps the anatomical vectors between two scanner domains to enforce cycle-consistency, while a new scanner-free loss $L_{\\mathrm{sf}}$ demands that images generated from the same Gaussian noise through different anatomies map to the same scanner encoding, which is what makes the scanner-free output act like a denoised, scanner-independent image. Attention layers in the encoders and generator, together with a 26-slice moving window, are the architectural changes that let the model retain anatomical detail while processing thinner volumes.","core_discovery":"On its own terms, the paper claims that a single generator can produce a scanner-free version of any T1-weighted brain scan by discarding the scanner-specific component of the image and regenerating the anatomy from latent content plus Gaussian noise. The same machinery also allows the image to be transplanted into any training-scanner domain. Across healthy controls, traveling subjects, and Alzheimer's patients, the harmonized images come out visually consistent, their voxel-intensity distributions converge, and downstream models perform better; age prediction reaches $R^2 \\approx 0.60$, Alzheimer's versus healthy classification accuracy reaches $0.86$, and diagnosis of mild cognitive impairment versus Alzheimer's disease reaches an AUC of $0.95$, in each case above the two benchmark harmonizers. The paper further claims this works for scanners never seen during training and without skull-stripping, which it treats as an advantage for full-head applications.","pith_inferences":["If the disentanglement holds, the same scanner-free space could serve as a common substrate for cross-site federated learning, where models share harmonized images or features without sharing raw data; the paper does not discuss this.","A testable extension is to apply DISARM++ to a paired traveling-subject dataset with a visible anatomical abnormality; the scanner-free outputs should agree across scanners while still showing the abnormality, directly testing whether pathology is preserved.","The scanner-free encoding could also be used as a normalization step for image retrieval or for training generative models on pooled multi-site data, though the paper only evaluates the four downstream tasks.","Replacing the scanner code with Gaussian noise may also suppress scanner-specific image fingerprints, which raises a possible anonymization use that the paper does not claim."],"forward_implications":["Multi-site MRI studies can be pooled without feature-level harmonization or skull-stripping, because the harmonization happens on the image itself.","Scans from new scanners can be harmonized by inference only, with no retraining or fine-tuning, which lowers the barrier to adding new sites to a study.","Two harmonization modes are available: scanner-free output for general pooling, and reference-scanner style transfer when a downstream model was trained on one specific scanner's look.","Because the whole head is preserved, the method can support analyses outside brain tissue, such as head trauma and cranial deformation.","Downstream predictive tasks such as age estimation, Alzheimer's screening, and diagnosis inherit the harmonization benefit directly, rather than requiring a separate feature-level correction."],"supporting_citations":[{"why":"Supplies the disentangled image-to-image translation architecture that DISARM++ builds on.","marker":"[18]"},{"why":"The conference-paper baseline DISARM that this extended version modifies with attention, a new loss, and larger training data.","marker":"[4]"},{"why":"STGAN is one of the two state-of-the-art image-based baselines the paper must beat; its pre-trained model is used for comparison.","marker":"[7]"},{"why":"IGUANe is the other state-of-the-art baseline, a 3D CycleGAN harmonizer to a reference dataset, also used pre-trained.","marker":"[32]"},{"why":"Provides the traveling-subjects data used to test whether the same person's scans from different scanners become more similar after harmonization.","marker":"[40]"},{"why":"Defines SSIM and Struct-SSIM, the metrics used to quantify anatomical-structure preservation and traveling-subject similarity.","marker":"[46]"},{"why":"Provides the whole-brain segmentation method used to extract volumetric biomarkers for the downstream analyses.","marker":"[8]"}],"fun_headline_variants":["MRI harmonizer wipes scanner bias without skull-stripping","Scanner-free harmonization generalizes to unseen scanners","DISARM++ harmonizes brains for any scanner without retraining","MRI harmonizer: no skull-stripping, no retraining needed","Brain scan harmonizer works on new scanners instantly"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model assumes an MR image splits cleanly into anatomy and scanner effect, so that replacing the scanner part with Gaussian noise leaves all biologically relevant signal untouched.","fun_headline_variants_meta":{"raw":{"variants":["MRI harmonizer wipes scanner bias without skull-stripping","Scanner-free harmonization generalizes to unseen scanners","DISARM++ harmonizes brains for any scanner without retraining","MRI harmonizer: no skull-stripping, no retraining needed","Brain scan harmonizer works on new scanners instantly"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000699,"raw_usage":{"total_tokens":3201,"prompt_tokens":1036,"completion_tokens":2165,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":652,"completion_tokens_details":{"reasoning_tokens":2083}},"tokens_in":652,"tokens_out":2165,"duration_ms":16176,"temperature":1.0,"reasoning_tokens":2083,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:44:25.522763+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a traveling-subjects set where the same person is scanned on several machines, compute the pairwise similarity of scanner-free outputs across scanners and compare each output against the person's known anatomy, such as a real lesion. The claim fails if scanner-free outputs from the same subject remain scanner-dependent, or if two different subjects' scanner-free outputs become more alike than their raw scans, or if a visible lesion is erased in the scanner-free version.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the disentangled image-to-image translation architecture that DISARM++ builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The conference-paper baseline DISARM that this extended version modifies with attention, a new loss, and larger training data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"STGAN is one of the two state-of-the-art image-based baselines the paper must beat; its pre-trained model is used for comparison."},{"cited_title":"IGUANe: a 3D generalizable CycleGAN for multicenter harmonization of brain MR images","cited_arxiv_id":"2402.03227","evidence_quote":"IGUANe is the other state-of-the-art baseline, a 3D CycleGAN harmonizer to a reference dataset, also used pre-trained."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the traveling-subjects data used to test whether the same person's scans from different scanners become more similar after harmonization."},{"cited_title":"C., Sheikh, H","cited_arxiv_id":null,"evidence_quote":"Defines SSIM and Struct-SSIM, the metrics used to quantify anatomical-structure preservation and traveling-subject similarity."},{"cited_title":"H., Busa, E., Albert, M., Dieterich, M., Haselgrove, C., van der Kouwe, A., Killiany, R., Kennedy, D., Klaveness, S., Montillo, A., Makris, N., Rosen, B., and Dale, A","cited_arxiv_id":null,"evidence_quote":"Provides the whole-brain segmentation method used to extract volumetric biomarkers for the downstream analyses."}],"review_version":1}