REVIEW 2 major objections 4 minor 1 cited by
Brain foundation model embeddings encode scanner and site differences more strongly than diagnosis.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 20:17 UTC pith:X6TD6OMV
load-bearing objection Solid empirical workshop paper showing frozen BrainLM/SwiFT embeddings carry site structure that often dominates diagnosis; the finding is well-supported and worth citing. the 2 major comments →
Batch Effects In Brain Foundation Model Embeddings
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Subject-level embeddings extracted from two frozen neuroimaging foundation models (BrainLM and SwiFT) encode substantial acquisition-related batch effects that often dominate diagnosis-related information across multi-site fMRI datasets; harmonization can attenuate the batch signal without necessarily amplifying clinical signal, and the two models preferentially encode different biological summaries consistent with their architectures.
What carries the argument
A multi-pronged evaluation of frozen subject-level embeddings (CLS token for BrainLM; average of temporal-window embeddings for SwiFT) that combines PCA/LDA visualization, PERMANOVA, site-versus-diagnosis classifiers, deliberately confounded train/test splits, post-hoc ComBat, and decoding of ALFF versus FNC.
Load-bearing premise
That taking the CLS token or averaging window embeddings from already-trained, frozen models, without any task-specific fine-tuning or site-aware pre-training, is a fair and sufficient way to measure what the models have learned about biology versus acquisition.
What would settle it
If, after the same embedding extraction on the same multi-site cohorts, site classification accuracy fell to chance while diagnosis accuracy rose above that of classical functional connectivity, or if confounded site-by-diagnosis splits no longer produced near-perfect accuracy, the central claim would fail.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper evaluates subject-level embeddings from two frozen neuroimaging foundation models (BrainLM CLS tokens and SwiFT window-averaged embeddings) on multi-site resting-state fMRI (FBIRN, ADHD-200, ABIDE I). Using PCA/LDA visualizations, PERMANOVA, site vs. diagnosis classifiers, controlled within-site / multi-site / confounded designs, ComBat harmonization, and decoding of ALFF vs. FNC, the authors show that batch/site structure often dominates diagnosis-related information in the embedding space. Harmonization reduces site predictability without reliably improving diagnosis prediction, and the two models preferentially encode regional activity (BrainLM) versus inter-regional interactions (SwiFT), consistent with their architectures.
Significance. If the result holds under the stated frozen-embedding protocol, it is a timely and practically important caution for the growing use of brain foundation models in multi-site clinical neuroimaging. The work supplies multiple convergent empirical lines of evidence (visualization, multivariate tests, predictive accuracy, confounded-site controls, harmonization, and biological decoding) rather than a single metric, and it cleanly separates acquisition-related from diagnosis-related structure. Strengths include transparent experimental design, demographic regression controls, comparison against handcrafted FNC, and architecture-consistent interpretability findings. For a non-archival workshop paper the contribution is solid and actionable: it motivates batch-aware evaluation and future disentanglement methods without overclaiming that the models cannot learn biology under different training or fine-tuning regimes.
major comments (2)
- The central claim is well supported under the frozen CLS / window-average protocol, but the manuscript should state more explicitly (Methods §3 and Discussion) that this protocol is a deliberate probe of released checkpoints rather than a claim about all possible uses of these models. Without that boundary, readers may over-generalize the dominance finding to fine-tuned or site-aware pre-trained settings that the paper does not test.
- Tables 2 and 13 (confounded and multi-site settings) are load-bearing for the shortcut interpretation. Sample sizes for some site pairs are modest; the paper should report confidence intervals or bootstrap variability for the near-100% confounded accuracies and for the multi-site degradation cases so that the strength of the shortcut claim is quantified rather than left as point estimates.
minor comments (4)
- Figure numbering and cross-references are occasionally inconsistent (e.g., main-text references to Figure 2/3/4 vs. appendix figures); a single pass to align captions and in-text citations would help.
- Several appendix tables (PERMANOVA Table 11, full prediction Tables 12–16) are essential to the argument; consider moving a compact summary of site vs. diagnosis effect sizes into the main text for readers who do not open the appendix.
- Clarify the exact PCA dimensionality used before LDA/classifiers in each experiment (main text mentions 20 components in places; free parameter should be fixed and stated once).
- Minor typographical issues (e.g., spacing around PERMANOVA, occasional missing spaces after periods) and a few incomplete sentences in the harmonization paragraph of §4 should be cleaned.
Circularity Check
No significant circularity: purely empirical evaluation with external labels and frozen extractors
full rationale
The paper is an empirical study of frozen subject-level embeddings (BrainLM CLS token; SwiFT window-averaged) from publicly released checkpoints. Site and diagnosis labels are external ground truth from multi-site fMRI cohorts (FBIRN, ADHD-200, ABIDE I). Analyses (PCA/LDA visualizations, PERMANOVA, site vs. diagnosis classifiers, within-site vs. multi-site vs. confounded controlled settings, ComBat harmonization, ALFF/FNC decoding) report observed statistics; no quantity is defined in terms of a fitted parameter that is later presented as a prediction, and no first-principles derivation is claimed. Self-citations (e.g., COINSTAC, NeuroMark, FNC/ICA background) supply standard methodological context and are not load-bearing for the dominance-of-batch claim. The evaluation protocol is transparent and conventional for foundation-model embedding probes; the results stand or fall on the reported experiments rather than on circular construction. Score 0 is therefore appropriate.
Axiom & Free-Parameter Ledger
free parameters (2)
- PCA dimensionality for visualization and pre-classifier reduction =
20
- AAL-424 atlas for BrainLM ROI timeseries
axioms (3)
- domain assumption Site/scanner differences introduce non-biological variability in multi-site fMRI that can be quantified by PERMANOVA and classification accuracy.
- ad hoc to paper Frozen CLS-token or averaged-window embeddings from the released checkpoints are valid subject-level representations of what the foundation models have learned.
- domain assumption ComBat is a reasonable first-line harmonization method for high-dimensional embeddings.
read the original abstract
Foundation models show strong potential for large-scale, high-dimensional biomedical applications, yet their ability to capture relevant neurobiological characteristics remains underexplored. We systematically evaluate embeddings from two neuroimaging foundation models, BrainLM and SwiFT, across multi-site fMRI datasets using a comprehensive evaluation framework. Our results show that foundation model embeddings encode substantial batch-related variability, often dominating diagnosis-related information across heterogeneous datasets. We further investigate how harmonization, applied to reduce batch effects, influences these embeddings. In addition, we find that BrainLM prefers to capture fine-grained regional activity, whereas SwiFT tends to represent interactions between regions, consistent with their respective model architectures. Our study highlights the importance of accounting for batch effects in foundation models and motivates future work on disentangling biologically meaningful signals from acquisition-related variability.
Figures
Forward citations
Cited by 1 Pith paper
-
Foundation Models for EEG Are Blind to Long-Range Temporal Correlations: A Spectral-Temporal Dissociation Behind Their Cross-Population Fragility
EEG foundation models fail to encode the alpha-envelope DFA exponent, a disease-relevant temporal-scaling feature, while spectral-input models still encode the static 1/f slope.
Reference graph
Works this paper leans on
-
[1]
O., Fonseca, A
Caro, J. O., Fonseca, A. H. d. O., Averill, C., Rizvi, S. A., Rosati, M., Cross, J. L., Mittal, P., Zappala, E., Levine, D., Dhodapkar, R. M., et al. BrainLM: A foundation model for brain activity recordings.�������, pp. 2023–09,
2023
-
[2]
C., James, G
Craddock, R. C., James, G. A., Holtzheimer III, P. E., Hu, X. P., and Mayberg, H. S. A whole brain fMRI atlas generated via spatially constrained spectral clustering. ����� ����� �������, 33(8):1914–1928,
1914
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.