{"id":"a420b777-db8b-45f9-b412-3d7fbdd36285","arxiv_id":"2505.14722","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Under controlled simulations and two real datasets, ComBAT harmonization fails when sites differ in age slope, have under 16 to 32 subjects, narrow age ranges, or mix pathological cases into parameter estimation.","lead":"This paper reviews the assumptions behind ComBAT, a standard method for removing site-to-site biases in brain MRI, and tests what happens when those assumptions break. It proposes a pairwise variant that aligns every site to one fixed reference dataset and issues practical rules for sample size, age range, and patient mix.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Quantitative thresholds in the recommendations are calibrated only on synthetic moving sites made from CamCAN via Eq. (16); their portability to real, nonlinear, or confounded site effects is not established.","rationale":"The reader's weakest assumption is the same concern I identify: the experimental program treats Eq. (16) as a faithful generative model of real inter-site differences, and the paper's quantitative recommendations inherit that assumption. The analytical demonstration of the slope-mismatch limitation is correct and independently checkable, so the paper's broad message (do not use ComBAT as a black box; check covariate slopes and demographics) is defensible. The problem is that the specific cutoffs (16-32 subjects; 40-year age span) are presented as general best practices, while the supporting experiments are synthetic transformations of a single reference population. This does not force rejection, because the recommendations are plausible and partly consistent with prior literature, such as Orlhac et al. recommending N at least 30 and Karayumak et al. recommending 16-18 well-matched controls. It does strengthen the conditional nature of the reader's verdict: the guidelines should be labeled as bounded to the tested linear, single-reference setting until reproduced on independent real multi-site data. Hence the verdict remains CONDITIONAL rather than ACCEPT or REJECT.","tokens_in":23475,"tokens_out":9972,"duration_ms":99584,"concrete_test":"Re-run Experiments 2 and 3 using a traveling-subject multi-site diffusion dataset, where each participant is scanned at two or more sites, and compare harmonized values to the same participant's reference-site measurement at N=8, 16, and 32 and age spans of 10, 20, and 40 years. If the 16-32-subject and 40-year thresholds do not reproduce the reported MAD/BD performance, the guidelines are not portable beyond the Eq. (16) simulation setting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The analytical point about ComBAT's shared-slope assumption (Eq. 14) is sound: when the generative model of Eq. (14) holds, a site-specific slope multiplier S leaves a residual x^T beta (S-1) that no additive or multiplicative batch correction can remove. The load-bearing weakness is in the quantitative leap from that point to the paper's recommendations. Experiments 2 and 3, which ground the \"at least 16-32 subjects\" and \"age range at least 40 years\" guidelines, use only Modified-CamCAN: the moving site is created by applying the global linear transform in Eq. (16) to the same CamCAN data used as the reference, and the test set is generated by the same transform. In this self-consistent world ComBAT's linearity assumptions hold by construction, so the measured MAD curves in Figures 5 and 7 do not test whether the method works under the nonlinear age effects, voxel-dependent site effects, or acquisition-specific artifacts that occur in real multi-site data. The paper's own Eq. (14) also shows that a slope change S implies a coupled intercept change (gamma_iv = alpha_v(S-1)+A_i), whereas the simulation in Eq. (16) treats A and S as independent; to the extent that real scanner effects couple slope and intercept, the failure modes shown in Figure 3(b) may not transfer directly. The recommendations are also presented without error bars or statistical tests on the 30 repetitions, and the code is not yet public. The central analytical claim survives, but the externally valid scope of the guidelines does not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript re-derives the linear model underlying ComBAT harmonization and pinpoints the assumption that the covariate slope β_v is identical across sites. It introduces Pairwise-ComBAT, which harmonizes each site to a fixed reference dataset, and proposes a Bhattacharyya-distance goodness-of-fit measure. Using a synthetic \"modified-CamCAN\" dataset created by applying additive, slope, and multiplicative transforms to CamCAN, plus two real datasets (ADNI and NIMH), the authors run experiments varying bias magnitude, sample size, age range, sex covariates, and pathological contamination. They conclude with recommendations: ComBAT should not be used as a black box; covariate slopes should be matched; at least 16–32 subjects per site are recommended; the age range should span at least 40 years; sex distributions should be balanced; and harmonization parameters should be estimated on healthy controls only.","tokens_in":23787,"tokens_out":8420,"duration_ms":84101,"significance":"If the central claims hold, the paper is a useful service to the ComBAT user community. The formal observation around Eq. (14)—that a site-specific slope multiplier S_i leaves a residual x^T β(S_i−1) that no additive or multiplicative correction can remove—is correct and is illustrated in the synthetic and NIMH examples. The train/test decomposition of harmonization error is a good methodological instinct, and the recommendation to estimate batch parameters on healthy controls is clinically sensible. The ADNI pathology illustration in Fig. 10 convincingly shows the risk of compressing patient distributions into the normative range. However, the quantitative thresholds that are the paper's headline contribution rest on simulations that satisfy ComBAT's linear assumptions by construction, and no uncertainty is reported on the 30-repeat averages. The significance is therefore real, but the strength of the quantitative recommendations currently exceeds what the evidence supports.","major_comments":[{"comment":"The generative model used to create the synthetic moving site is not the model analyzed in the theory. Eq. (14) shows that a slope multiplier S_i necessarily changes the intercept through the term γ_iv = α_v(S_i−1)+A_i, so slope and additive site effects are coupled. Eq. (16), by contrast, applies the additive factor A and slope factor S independently to the CamCAN baseline. The simulated failure modes in Figs. 3–7 therefore do not directly instantiate the theoretical failure mode of Eq. (14), and the measured effects may not transfer to situations where slope and intercept changes are coupled. Please justify the independent parameterization or rerun the central experiments under the coupled model derived in Eq. (14).","section":"Mathematical limitations and Dataset studied (Eqs. 14 and 16)"},{"comment":"The recommendations \"at least 16 to 32 subjects\" and \"age range at least 40 years\" are read off MAD curves that are averaged over 30 repetitions but presented without confidence intervals, error bars, or any inferential test. In Fig. 5(a) and Fig. 7, the training and testing curves approach the reference floor smoothly, and it is not possible to tell whether the apparent plateau at N≈32 or age span≈40 is statistically distinguishable from the floor at larger N or wider age spans. Because these thresholds are a central product of the paper, please report per-condition variability (e.g., bootstrap CIs or box plots over the 30 repeats) and, where possible, a test comparing the error at the recommended operating point with the reference error.","section":"Results (Figs. 5 and 7)"},{"comment":"The sample-size and age-range experiments are conducted exclusively on Modified-CamCAN, in which the moving site and the test set are generated by applying the same global linear transform to the same CamCAN distribution used as the reference. This means the data satisfy ComBAT's model assumptions by construction, with no nonlinear age trajectories, voxel-dependent scanner effects, or acquisition-specific artifacts. The experiments therefore test interpolation under a correct model, not robustness to model misspecification. The ADNI and NIMH analyses are qualitative illustrations rather than quantitative validation of the thresholds. The recommendations should be labeled as conditional on the synthetic model, or supported by additional real-data analyses with varying sample size and age range.","section":"Experiments 2 and 3; Dataset studied (Modified-CamCAN)"},{"comment":"Eq. (15) adds back the covariate effect as x^T_Rj β̂_v, i.e., using the covariates of a reference-site subject rather than those of the moving-site subject being harmonized. If this is not a typographical error, the method replaces each moving subject's age/sex effect with the reference subject's covariates and thereby destroys biological variability. If it is a typo, it should read x^T_Mj β̂_v. Either way, the current text cannot be implemented faithfully, and the accompanying code is not yet public, so a reader cannot resolve the ambiguity from the repository.","section":"Pairwise-ComBAT (Eq. 15)"},{"comment":"The formula for the Bhattacharyya distance is inconsistent and, as printed, incorrect. The notation switches between subscripts R and T (z_Rjv is defined while the expression uses µ_Tv and σ_Tv), and the logarithmic term should have denominator 2 σ_Tv σ_Mv, not 2 σ_Tv + σ_Mv. Since all BD values in Figs. 3, 4, 9, and 10 are quantitative evidence for the paper's conclusions, the printed formula and notation must be corrected.","section":"Goodness of fit"}],"minor_comments":[{"comment":"The abstract states \"five essential recommendations,\" but the Recommendations section lists six bullet points and the Conclusion also says \"six recommendations.\" Please align the count.","section":"Abstract and Recommendations"},{"comment":"The availability statement says the code and data \"will be rendered public upon acceptance.\" For a paper whose stated goals are reproducibility and open science, please make the Pairwise-ComBAT code and the modified-CamCAN generation scripts available during review, at least as supplementary material.","section":"Code and Data Availability"},{"comment":"\"Mean Absolution Difference\" should be \"Mean Absolute Difference.\" In addition, the MAD metric is used without a formal definition; please define it in the Method section or figure captions.","section":"Captions of Figs. 8 and 9"},{"comment":"The label \"Error Ref = 0.465e10-5\" appears to mean 0.465×10⁻⁵ but the typesetting is confusing. Please write it unambiguously.","section":"Figure 5(a)"}],"recommendation":"major_revision","confidential_remarks":"The core analytical point about ComBAT's shared-slope assumption is sound and worth publishing. The main risk is overclaiming the quantitative thresholds: they rest on a self-consistent synthetic model and have no uncertainty quantification. I believe the paper can be made acceptable by softening the external-validity claims, adding CIs or bootstrapped intervals, and fixing the equation-level issues in Eqs. (15) and the Bhattacharyya formula. There is no sign of circularity or fabrication in the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my take on 2505.14722. The paper is worth a look if you use ComBAT, but it is not a theoretical advance. The main point—ComBAT assumes a common slope across sites and fails when that breaks—is already in Kim et al. and Bayer et al., and the authors say Pairwise-ComBAT is only a reconfiguration of the original. What they add is a systematic, controlled look at how sample size, age range, sex balance, and patient mix affect harmonization, with clear recommendations at the end. That is genuinely useful for applied multi-site dMRI work.\n\nThe theory section is correctly written and the slope-mismatch derivation (Eq. 14) is sound. I also credit them for treating harmonization as a train/test problem rather than just looking at training fit. The ADNI and NIMH examples are real data and support the qualitative conclusions: slope differences produce bad harmonization, and mixing patients into the parameter estimation compresses pathology.\n\nThe soft spots are in the quantitative recommendations. The '16-32 subjects' and 'at least 40 years of age span' thresholds come from the modified-CamCAN simulations, where the moving site is created by applying a linear transform to the same CamCAN data used as the reference. That means the simulation is self-consistent with ComBAT's own generative model; it does not test nonlinear age effects, voxel-dependent site effects, or the messy interactions you see across real scanners. The MAD curves in Figures 5 and 7 are averages over 30 repeats with no confidence intervals and no statistical tests, so the specific cutoffs are not well-grounded. The central qualitative claims hold up; the precise numbers should be treated as provisional.\n\nA couple of other things: the code and data are promised only upon acceptance, which limits reproducibility now, and there are typos in key formulas (Eq. 10 and the Bhattacharyya distance expression). These are fixable, not fatal.\n\nWho benefits? Practitioners who use ComBAT without knowing its failure modes, and normative modeling groups. I would cite it if I were writing a methods paper on multi-site harmonization, but not as the authoritative source for the thresholds.\n\nMy recommendation: send it to peer review. It is a solid, honest evaluation of a widely used tool, and the recommendations matter to a large community. The referee should insist on error bars, a public code release, and a discussion of how the simulated site effects might differ from real ones.","headline":"A useful, honest practical evaluation of ComBAT's known limitations, but the new quantitative thresholds rest on simulations that mirror ComBAT's own linear assumptions and are not statistically grounded.","tokens_in":24340,"tokens_out":3264,"would_cite":true,"duration_ms":30517,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P10","62F15","62J05"],"pacs":[],"model":"deepseek-v4-flash","headline":"ComBAT, the standard method for pooling MRI measurements across sites, works only when age's effect is identical at every site; violating that hidden assumption misaligns populations and corrupts variance estimates.","keywords":["ComBAT harmonization","diffusion MRI","multi-site imaging","batch effects","covariate slope","normative modeling","sample size","empirical Bayes"],"falsifier":"Take two real sites whose age-versus-metric slopes are known to differ markedly, run standard ComBAT with more than 32 subjects and a full age span at each site, and measure the Bhattacharyya distance between the harmonized and reference distributions on held-out subjects: the paper predicts a large distance whenever the slope factor departs from 1, so near-zero distance on such a pair would contradict the central claim. A second check targets the variance claim: on two sites with matched slopes, compare ComBAT's estimated site variance with the empirical within-site variance; if they converge well before 100 subjects, the paper's claim that the prior keeps disproportionately high weight near 100 subjects would be refuted.","tokens_in":23286,"feed_emoji":"🧠","tokens_out":14496,"duration_ms":127467,"temperature":0.7,"pith_summary":"ComBAT is the most widely used statistical tool for combining MRI-derived measurements collected at different hospitals and scanners, removing site-specific additive and multiplicative biases while trying to preserve genuine biological variability. This paper argues that ComBAT hides a strong assumption: the slope linking covariates such as age to the measured metric must be the same at every site, and when that fails the harmonization is misleading. Using a large healthy-aging cohort whose values are artificially rescaled with known factors, the authors show exactly when harmonization breaks and when it works, and they repackage ComBAT as a pairwise procedure that aligns each site to one fixed reference dataset. The payoff is a set of concrete conditions for safe use — at least 16 to 32 subjects per site, an age span of roughly 40 years, balanced sex composition, and parameters estimated from healthy controls only — which matters for any multi-site study, normative modeling effort, or clinical deployment that pools brain measurements.","feed_headline":"ComBAT breaks when age effects differ across MRI sites","feed_subtitle":"Safe-use rules for the popular harmonizer: 16+ subjects per site, 40-year age span, balanced sex, healthy controls only.","key_machinery":"The load-bearing object is the linear generative model $y_{ijv} = \\alpha_v + x_{ij}^T\\beta_v + \\gamma_{iv} + \\delta_{iv}\\varepsilon_{ijv}$, where $x_{ij}$ collects covariates such as age and sex, $\\beta_v$ is the common population slope, $\\gamma_{iv}$ the additive site effect, and $\\delta_{iv}$ the multiplicative site effect. Because ComBAT standardizes data by removing the population trend $x_{ij}^T\\hat{\\beta}_v$, any site-specific multiplication of that slope — written as the factor $S$ in $y_{ijv} = \\alpha_v + \\gamma_{iv}A + x_{ij}^T\\beta_v S + \\delta_{iv}\\varepsilon_{ijv}M$ — bypasses the correction and survives as a misalignment. The experimental vehicle is Pairwise-ComBAT, a reconfiguration that harmonizes each moving site independently onto a fixed reference population rather than onto a pooled average, and a synthetic-protocol design in which a real healthy-aging cohort is rescaled with known additive ($A$), slope ($S$), and variance ($M$) factors so that harmonization error can be measured exactly. A closed-form Bhattacharyya distance between rectified Gaussian populations serves as the quality metric.","core_discovery":"ComBAT models each measured value as a linear function of covariates plus an additive site effect and a multiplicative site effect, and assumes the covariate regression vector $\\beta_v$ — the population slope relating age to the measured metric — is identical at every site. The central demonstration is that this slope assumption is often wrong: when a moving site has been rescaled by a slope factor $S \\neq 1$ relative to the reference, ComBAT compensates the additive bias but leaves the multiplicative bias uncorrected, so the harmonized populations stay misaligned and the estimated site variance $\\hat{\\delta}^{2*}_{iv}$ carries a growing quadratic error. In controlled experiments with known injected bias, slope, and variance factors, harmonization quality deteriorates sharply for slope factors of 0 and 1.8 even though pure additive and pure multiplicative differences are handled well. The paper also shows that the variance estimator keeps a disproportionately large weight on its prior even for populations above roughly 100 subjects, and introduces a Bhattacharyya-distance measure on covariate-rectified data as a goodness-of-fit metric for harmonization.","pith_inferences":["A cheap pre-flight check the authors do not propose: fit the covariate model inside each site separately, compare slope estimates and their confidence intervals, and restrict ComBAT harmonization to site pairs whose slopes overlap, turning the paper's negative result into a positive screening test.","Because the variance estimator keeps strong prior weight even near 100 subjects, a natural extension is a sample-size-aware prior or reporting both the shrunk and the empirical variance; the paper documents the symptom but does not propose the remedy.","The equal-slope requirement is stated for linear slopes, yet many diffusion metrics change nonlinearly with age; a testable prediction is that the 40-year age-range rule is necessary but not sufficient on cohorts with curvilinear age trajectories, where a global linear fit can mask locally mismatched slopes.","The reference-anchored framing implies a frozen-parameter protocol for live clinical data: estimate the harmonization function once on a normative reference, then harmonize each incoming subject with those fixed parameters, eliminating the moving-target problem by construction rather than by re-estimation."],"forward_implications":["Multi-site studies should verify that covariate slopes, such as the age-versus-metric relationship, are approximately equal across sites before running ComBAT, because large slope discrepancies predict failed harmonization even when additive and multiplicative biases are modest.","Reliable harmonization has concrete lower bounds on the moving site: at least 16 to 32 training subjects, an age span of roughly 40 years or more, and balanced male/female composition; below these, test-set generalization degrades even when training fits look good.","For normative modeling and clinical applications, harmonization parameters should be estimated on healthy control subjects only and then applied to pathological subjects; including pathology in the estimation compresses diseased populations into the normal range and erases the disease signal.","Harmonizing each site to a single fixed reference dataset, rather than to the pooled average of all sites, allows new sites to be added without re-estimating previously established parameters, a direct benefit for open-data sharing, longitudinal studies, and clinical data streams.","Because a slope mismatch corrupts the estimated site variance, any downstream statistic built from post-harmonization variance — such as normative deviation scores — inherits the bias unless the slope condition is checked first."],"supporting_citations":[{"why":"Supplies the original empirical Bayes ComBAT formulation whose standardization equations and priors the paper dissects.","marker":"[18]"},{"why":"Defines Bayesian ComBAT for diffusion tensor imaging, the algorithm under study and the source of the parameter-estimation steps.","marker":"[19]"},{"why":"Provides the healthy-aging reference cohort from which the modified moving-site populations are generated.","marker":"[34]"},{"why":"Documents the reference cohort's demographics and acquisition protocol that make it the anchor site for Pairwise-ComBAT.","marker":"[35]"},{"why":"Empirical assessment of ComBAT's assumptions on diffusion tensor data, cited as evidence that the equal-slope assumption is not always valid.","marker":"[43]"},{"why":"Practical guide to ComBAT harmonization that supplies the sample-size and clinical-status precedents the recommendations align with.","marker":"[32]"},{"why":"Lifespan harmonization method cited for the requirement of wide, overlapping age distributions across sites.","marker":"[29]"},{"why":"Modified ComBAT for radiomic features that Pairwise-ComBAT conceptually reconfigures toward a fixed reference site.","marker":"[28]"},{"why":"Overview of site-effect correction, cited for the claim that ComBAT's assumptions are often violated in practice.","marker":"[11]"}],"fun_headline_variants":["ComBAT fails when age effects differ across sites","If age effects vary by site, ComBAT harmonization is flawed","ComBAT's slope assumption: a hidden pitfall in MRI harmonization","ComBAT assumes identical age effects across sites—often false","ComBAT corrects additive bias, but not multiplicative when slopes differ"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The experiments assume real inter-site differences are exactly the kind the authors inject — a global additive offset, a slope change, and a variance change applied to one healthy population — so if genuine site effects are messier, such as nonlinear age trajectories or acquisition-specific artifacts, the diagnosed failure modes and the recommended sample-size and age-range thresholds may not transfer to real data.","fun_headline_variants_meta":{"raw":{"variants":["ComBAT fails when age effects differ across sites","If age effects vary by site, ComBAT harmonization is flawed","ComBAT's slope assumption: a hidden pitfall in MRI harmonization","ComBAT assumes identical age effects across sites—often false","ComBAT corrects additive bias, but not multiplicative when slopes differ"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001005,"raw_usage":{"total_tokens":4245,"prompt_tokens":933,"completion_tokens":3312,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":3221}},"tokens_in":549,"tokens_out":3312,"duration_ms":26835,"temperature":1.0,"reasoning_tokens":3221,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:15:41.927576+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take two real sites whose age-versus-metric slopes are known to differ markedly, run standard ComBAT with more than 32 subjects and a full age span at each site, and measure the Bhattacharyya distance between the harmonized and reference distributions on held-out subjects: the paper predicts a large distance whenever the slope factor departs from 1, so near-zero distance on such a pair would contradict the central claim. A second check targets the variance claim: on two sites with matched slopes, compare ComBAT's estimated site variance with the empirical within-site variance; if they converge well before 100 subjects, the paper's claim that the prior keeps disproportionately high weight near 100 subjects would be refuted.","supporting_citations":[{"cited_title":"Adjusting batch effects in microarray expression data using empirical Bayes methods","cited_arxiv_id":null,"evidence_quote":"Supplies the original empirical Bayes ComBAT formulation whose standardization equations and priors the paper dissects."},{"cited_title":"Harmonization of multi-site diffusion tensor imaging data","cited_arxiv_id":null,"evidence_quote":"Defines Bayesian ComBAT for diffusion tensor imaging, the algorithm under study and the source of the parameter-estimation steps."},{"cited_title":"The Cambridge Centre for Age- ing and Neuroscience (Cam-CAN) study protocol: a cross- sectional, lifespan, multidisciplinary examination of healthy cognitive ageing","cited_arxiv_id":null,"evidence_quote":"Provides the healthy-aging reference cohort from which the modified moving-site populations are generated."},{"cited_title":"The Cambridge Centre for Ageing and Neuroscience (Cam-CAN) data repository: Structural and functional MRI, MEG, and cognitive data from a cross- sectional adult lifespan sample","cited_arxiv_id":null,"evidence_quote":"Documents the reference cohort's demographics and acquisition protocol that make it the anchor site for Pairwise-ComBAT."},{"cited_title":"Empirical assessment of the assump- tions of ComBat with diffusion tensor imaging","cited_arxiv_id":null,"evidence_quote":"Empirical assessment of ComBAT's assumptions on diffusion tensor data, cited as evidence that the equal-slope assumption is not always valid."},{"cited_title":"A guide to ComBat harmonization of imaging biomarkers in multicenter studies","cited_arxiv_id":null,"evidence_quote":"Practical guide to ComBAT harmonization that supplies the sample-size and clinical-status precedents the recommendations align with."},{"cited_title":"Harmonization of large MRI datasets for the analysis of brain imaging patterns through- out the lifespan","cited_arxiv_id":null,"evidence_quote":"Lifespan harmonization method cited for the requirement of wide, overlapping age distributions across sites."},{"cited_title":"Performance comparison of modified Com- Bat for harmonization of radiomic features for multicenter studies","cited_arxiv_id":null,"evidence_quote":"Modified ComBAT for radiomic features that Pairwise-ComBAT conceptually reconfigures toward a fixed reference site."},{"cited_title":"Site effects how-to and when: An overview of retrospective techniques to accommodate site effects in multi-site neuroimaging analyses","cited_arxiv_id":null,"evidence_quote":"Overview of site-effect correction, cited for the claim that ComBAT's assumptions are often violated in practice."}],"review_version":1}