Pith. sign in

REVIEW 2 major objections 4 minor 1 cited by

Brain foundation model embeddings encode scanner and site differences more strongly than diagnosis.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 20:17 UTC pith:X6TD6OMV

load-bearing objection Solid empirical workshop paper showing frozen BrainLM/SwiFT embeddings carry site structure that often dominates diagnosis; the finding is well-supported and worth citing. the 2 major comments →

arxiv 2604.14441 v2 pith:X6TD6OMV submitted 2026-04-15 eess.SP

Batch Effects In Brain Foundation Model Embeddings

classification eess.SP
keywords foundation modelsfMRIbatch effectssite effectsembeddingsharmonizationBrainLMSwiFT
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks what subject-level embeddings from two frozen neuroimaging foundation models actually contain when applied to multi-site resting-state fMRI. Across three public cohorts the embeddings show clear batch structure: site identity can be recovered at high accuracy, PERMANOVA attributes more multivariate variance to site than to diagnosis, and deliberately confounded train/test splits reach near-perfect accuracy by exploiting site rather than disease. Standard ComBat harmonization applied after embedding extraction reduces site predictability but does not reliably strengthen diagnosis signals. The two models also differ in the biological features they preserve: one aligns more with regional activity amplitude, the other with inter-region connectivity, matching their architectures. The practical message is that high downstream accuracy on multi-site fMRI can be an artifact of acquisition shortcuts, so batch-aware evaluation and representation design are required before these embeddings can be trusted as clinical biomarkers.

Core claim

Subject-level embeddings extracted from two frozen neuroimaging foundation models (BrainLM and SwiFT) encode substantial acquisition-related batch effects that often dominate diagnosis-related information across multi-site fMRI datasets; harmonization can attenuate the batch signal without necessarily amplifying clinical signal, and the two models preferentially encode different biological summaries consistent with their architectures.

What carries the argument

A multi-pronged evaluation of frozen subject-level embeddings (CLS token for BrainLM; average of temporal-window embeddings for SwiFT) that combines PCA/LDA visualization, PERMANOVA, site-versus-diagnosis classifiers, deliberately confounded train/test splits, post-hoc ComBat, and decoding of ALFF versus FNC.

Load-bearing premise

That taking the CLS token or averaging window embeddings from already-trained, frozen models, without any task-specific fine-tuning or site-aware pre-training, is a fair and sufficient way to measure what the models have learned about biology versus acquisition.

What would settle it

If, after the same embedding extraction on the same multi-site cohorts, site classification accuracy fell to chance while diagnosis accuracy rose above that of classical functional connectivity, or if confounded site-by-diagnosis splits no longer produced near-perfect accuracy, the central claim would fail.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper evaluates subject-level embeddings from two frozen neuroimaging foundation models (BrainLM CLS tokens and SwiFT window-averaged embeddings) on multi-site resting-state fMRI (FBIRN, ADHD-200, ABIDE I). Using PCA/LDA visualizations, PERMANOVA, site vs. diagnosis classifiers, controlled within-site / multi-site / confounded designs, ComBat harmonization, and decoding of ALFF vs. FNC, the authors show that batch/site structure often dominates diagnosis-related information in the embedding space. Harmonization reduces site predictability without reliably improving diagnosis prediction, and the two models preferentially encode regional activity (BrainLM) versus inter-regional interactions (SwiFT), consistent with their architectures.

Significance. If the result holds under the stated frozen-embedding protocol, it is a timely and practically important caution for the growing use of brain foundation models in multi-site clinical neuroimaging. The work supplies multiple convergent empirical lines of evidence (visualization, multivariate tests, predictive accuracy, confounded-site controls, harmonization, and biological decoding) rather than a single metric, and it cleanly separates acquisition-related from diagnosis-related structure. Strengths include transparent experimental design, demographic regression controls, comparison against handcrafted FNC, and architecture-consistent interpretability findings. For a non-archival workshop paper the contribution is solid and actionable: it motivates batch-aware evaluation and future disentanglement methods without overclaiming that the models cannot learn biology under different training or fine-tuning regimes.

major comments (2)
  1. The central claim is well supported under the frozen CLS / window-average protocol, but the manuscript should state more explicitly (Methods §3 and Discussion) that this protocol is a deliberate probe of released checkpoints rather than a claim about all possible uses of these models. Without that boundary, readers may over-generalize the dominance finding to fine-tuned or site-aware pre-trained settings that the paper does not test.
  2. Tables 2 and 13 (confounded and multi-site settings) are load-bearing for the shortcut interpretation. Sample sizes for some site pairs are modest; the paper should report confidence intervals or bootstrap variability for the near-100% confounded accuracies and for the multi-site degradation cases so that the strength of the shortcut claim is quantified rather than left as point estimates.
minor comments (4)
  1. Figure numbering and cross-references are occasionally inconsistent (e.g., main-text references to Figure 2/3/4 vs. appendix figures); a single pass to align captions and in-text citations would help.
  2. Several appendix tables (PERMANOVA Table 11, full prediction Tables 12–16) are essential to the argument; consider moving a compact summary of site vs. diagnosis effect sizes into the main text for readers who do not open the appendix.
  3. Clarify the exact PCA dimensionality used before LDA/classifiers in each experiment (main text mentions 20 components in places; free parameter should be fixed and stated once).
  4. Minor typographical issues (e.g., spacing around PERMANOVA, occasional missing spaces after periods) and a few incomplete sentences in the harmonization paragraph of §4 should be cleaned.

Circularity Check

0 steps flagged

No significant circularity: purely empirical evaluation with external labels and frozen extractors

full rationale

The paper is an empirical study of frozen subject-level embeddings (BrainLM CLS token; SwiFT window-averaged) from publicly released checkpoints. Site and diagnosis labels are external ground truth from multi-site fMRI cohorts (FBIRN, ADHD-200, ABIDE I). Analyses (PCA/LDA visualizations, PERMANOVA, site vs. diagnosis classifiers, within-site vs. multi-site vs. confounded controlled settings, ComBat harmonization, ALFF/FNC decoding) report observed statistics; no quantity is defined in terms of a fitted parameter that is later presented as a prediction, and no first-principles derivation is claimed. Self-citations (e.g., COINSTAC, NeuroMark, FNC/ICA background) supply standard methodological context and are not load-bearing for the dominance-of-batch claim. The evaluation protocol is transparent and conventional for foundation-model embedding probes; the results stand or fall on the reported experiments rather than on circular construction. Score 0 is therefore appropriate.

Axiom & Free-Parameter Ledger

2 free parameters · 3 axioms · 0 invented entities

The paper is an empirical evaluation study. It inherits standard neuroimaging assumptions (BOLD signal, ROI parcellation, existence of site effects) and uses off-the-shelf models and harmonization. No new physical entities or free parameters are fitted to produce the central claim; the few numerical choices (PCA dimensionality, classifier families) are conventional and do not define the result.

free parameters (2)
  • PCA dimensionality for visualization and pre-classifier reduction = 20
    Typically reduced to 20 components before LDA or classification; conventional choice that does not alter the qualitative dominance of batch over diagnosis.
  • AAL-424 atlas for BrainLM ROI timeseries
    Fixed parcellation chosen by the original BrainLM authors; treated as given rather than optimized here.
axioms (3)
  • domain assumption Site/scanner differences introduce non-biological variability in multi-site fMRI that can be quantified by PERMANOVA and classification accuracy.
    Standard premise of the batch-effect literature (Johnson 2007, Fortin 2017) invoked throughout §1–4.
  • ad hoc to paper Frozen CLS-token or averaged-window embeddings from the released checkpoints are valid subject-level representations of what the foundation models have learned.
    Explicit methodological choice in §3; if a different pooling or fine-tuning regime were used the batch dominance might change.
  • domain assumption ComBat is a reasonable first-line harmonization method for high-dimensional embeddings.
    Applied in §4 following its established use on FNC; no claim that it is optimal.

pith-pipeline@v1.1.0-grok45 · 18469 in / 2418 out tokens · 33649 ms · 2026-07-12T20:17:18.187481+00:00 · methodology

0 comments
read the original abstract

Foundation models show strong potential for large-scale, high-dimensional biomedical applications, yet their ability to capture relevant neurobiological characteristics remains underexplored. We systematically evaluate embeddings from two neuroimaging foundation models, BrainLM and SwiFT, across multi-site fMRI datasets using a comprehensive evaluation framework. Our results show that foundation model embeddings encode substantial batch-related variability, often dominating diagnosis-related information across heterogeneous datasets. We further investigate how harmonization, applied to reduce batch effects, influences these embeddings. In addition, we find that BrainLM prefers to capture fine-grained regional activity, whereas SwiFT tends to represent interactions between regions, consistent with their respective model architectures. Our study highlights the importance of accounting for batch effects in foundation models and motivates future work on disentangling biologically meaningful signals from acquisition-related variability.

Figures

Figures reproduced from arXiv: 2604.14441 by Anand D. Sarwate, Bradley T. Baker, Sandeep Panta, Sergey Plis, Vince D. Calhoun, Ye Tao, Yu Wu.

Figure 1
Figure 1. Figure 1: Overview of the fMRI representation pipeline using foundation models. fMRI scans are encoded into low-dimensional embeddings by foundation models. These embeddings are used for dimensionality reduction (e.g., PCA, LDA) and predictive modeling. interpretability of these embeddings to study which biologi￾cal signals are emphasized by different foundation models, revealing systematic differences consistent wi… view at source ↗
Figure 2
Figure 2. Figure 2: (or [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: and [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Visualization of subject-level embeddings extracted from the pre-trained SwiFT model. In the last row, embeddings are first projected using PCA with 20 components, followed by further dimensionality reduction using LDA. All points are colored according to site identity. For the ABIDE I dataset, which includes 17 imaging sites, only the 10 sites with the largest sample sizes are shown for clarity. C.3. Addi… view at source ↗
Figure 5
Figure 5. Figure 5: Visualization of subject-level FNC features. In the last row, features are first projected using PCA with 20 components, followed by further dimensionality reduction using LDA. All points are colored according to site identity. For the ABIDE I dataset, which includes 17 imaging sites, only the 10 sites with the largest sample sizes are shown for clarity. higher site classification accuracy but lower diagno… view at source ↗
Figure 6
Figure 6. Figure 6: PCA visualization of subject-level representations colored by diagnostic labels. Top, middle, and bottom rows correspond to FBIRN, ADHD-200, and ABIDE I, respectively. Columns represent different feature types: FNC, BrainLM embeddings, and SwiFT embeddings. This comparison illustrates how the various representations separate diagnostic groups across datasets. 15 [PITH_FULL_IMAGE:figures/full_fig_p015_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Spatial maps of ALFF decoding performance (R 2 ) after ComBat harmonization. Only the top 30% predictive regions are shown for visualization. Visual Somatomotor Dorsal Attention Frontoparietal Default Ventral Attention Limbic Subcortical Cerebellar Functional Network 0.002 0.000 0.002 0.004 0.006 0.008 0.010 0.012 M e a n R 2 Foundation Model BrainLM SwiFT (a) FBIRN Visual Somatomotor Dorsal Attention Fron… view at source ↗
Figure 8
Figure 8. Figure 8: Network-level mean R 2 of FNC decoding after ComBat harmonization. Bars show the average predictive performance within each functional network for BrainLM and SwiFT embeddings. 16 [PITH_FULL_IMAGE:figures/full_fig_p016_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Foundation Models for EEG Are Blind to Long-Range Temporal Correlations: A Spectral-Temporal Dissociation Behind Their Cross-Population Fragility

    q-bio.NC 2026-07 conditional novelty 6.0

    EEG foundation models fail to encode the alpha-envelope DFA exponent, a disease-relevant temporal-scaling feature, while spectral-input models still encode the static 1/f slope.

Reference graph

Works this paper leans on

2 extracted references · cited by 1 Pith paper

  1. [1]

    O., Fonseca, A

    Caro, J. O., Fonseca, A. H. d. O., Averill, C., Rizvi, S. A., Rosati, M., Cross, J. L., Mittal, P., Zappala, E., Levine, D., Dhodapkar, R. M., et al. BrainLM: A foundation model for brain activity recordings.�������, pp. 2023–09,

  2. [2]

    C., James, G

    Craddock, R. C., James, G. A., Holtzheimer III, P. E., Hu, X. P., and Mayberg, H. S. A whole brain fMRI atlas generated via spatially constrained spectral clustering. ����� ����� �������, 33(8):1914–1928,