{"id":"e80a1ad3-6e38-4019-a594-12fe13ae28a9","arxiv_id":"2604.14441","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"Pre-trained BrainLM and SwiFT fMRI embeddings encode batch effects that dominate diagnosis information across multi-site datasets; ComBat harmonization reduces site predictability but does not reliably strengthen clinical signals.","lead":"Embeddings from two brain foundation models (BrainLM and SwiFT) encode scanner/site batch effects that often dominate diagnosis signals in multi-site fMRI. This matters for anyone using or sharing these models, because high downstream accuracy can be a shortcut to acquisition artifacts rather than biology, and site identity can leak from the embeddings.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader correctly identifies the strongest claim and the most natural methodological caveat (frozen CLS/average extraction). That caveat is real but not load-bearing against the paper’s actual claim, which is carefully scoped to pre-trained, frozen embeddings used as feature extractors (Methods §3). The multi-pronged empirical package (visualization, PERMANOVA, predictive accuracy, controlled confounded designs, FNC baselines, ComBat, and architecture-consistent ALFF/FNC decoding) makes the batch-dominance observation hard to dismiss as an artifact of a single analysis choice. For a non-archival workshop contribution the evidence is sufficient; the suggested concrete test would only strengthen generality, not rescue a failing argument. Therefore the ACCEPT verdict and low correctness risk stand.","tokens_in":14279,"tokens_out":521,"duration_ms":5822,"concrete_test":"Re-extract subject embeddings after a light site-balanced fine-tuning pass (or with an alternative pooling strategy such as mean-pooling all ROI tokens for BrainLM) on one held-out multi-site cohort; recompute the site-vs-diagnosis accuracy gap and confounded-setting accuracy. If the gap collapses and confounded accuracy falls well below 100%, the frozen-extraction protocol would be shown to overstate batch dominance; if the gap persists, the claim is robust to that design choice.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that frozen subject-level embeddings from BrainLM (CLS token) and SwiFT (window-averaged) encode batch variability that often dominates diagnosis—is supported by multiple orthogonal lines of evidence that cohere: PCA/LDA site clustering (Figs. 2–4), PERMANOVA site vs. diagnosis effects (Table 11), high site vs. lower diagnosis classifier accuracies (Tables 1, 12), and confounded-site experiments that reach ~100% accuracy while multi-site pooling can degrade within-site diagnosis performance (Tables 2, 13). The reader’s weakest assumption (that the frozen CLS/average extraction protocol is a faithful probe of what the models learned) is transparent, conventional for foundation-model embedding studies, and does not undermine the empirical observation under the stated protocol. Harmonization and ALFF/FNC decoding results further corroborate rather than contradict the claim. No internal inconsistency or untested premise that would reverse the dominance finding is apparent for a non-archival workshop paper of this scope.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper evaluates subject-level embeddings from two frozen neuroimaging foundation models (BrainLM CLS tokens and SwiFT window-averaged embeddings) on multi-site resting-state fMRI (FBIRN, ADHD-200, ABIDE I). Using PCA/LDA visualizations, PERMANOVA, site vs. diagnosis classifiers, controlled within-site / multi-site / confounded designs, ComBat harmonization, and decoding of ALFF vs. FNC, the authors show that batch/site structure often dominates diagnosis-related information in the embedding space. Harmonization reduces site predictability without reliably improving diagnosis prediction, and the two models preferentially encode regional activity (BrainLM) versus inter-regional interactions (SwiFT), consistent with their architectures.","tokens_in":14566,"tokens_out":811,"duration_ms":7111,"significance":"If the result holds under the stated frozen-embedding protocol, it is a timely and practically important caution for the growing use of brain foundation models in multi-site clinical neuroimaging. The work supplies multiple convergent empirical lines of evidence (visualization, multivariate tests, predictive accuracy, confounded-site controls, harmonization, and biological decoding) rather than a single metric, and it cleanly separates acquisition-related from diagnosis-related structure. Strengths include transparent experimental design, demographic regression controls, comparison against handcrafted FNC, and architecture-consistent interpretability findings. For a non-archival workshop paper the contribution is solid and actionable: it motivates batch-aware evaluation and future disentanglement methods without overclaiming that the models cannot learn biology under different training or fine-tuning regimes.","major_comments":[{"comment":"The central claim is well supported under the frozen CLS / window-average protocol, but the manuscript should state more explicitly (Methods §3 and Discussion) that this protocol is a deliberate probe of released checkpoints rather than a claim about all possible uses of these models. Without that boundary, readers may over-generalize the dominance finding to fine-tuned or site-aware pre-trained settings that the paper does not test.","section":null},{"comment":"Tables 2 and 13 (confounded and multi-site settings) are load-bearing for the shortcut interpretation. Sample sizes for some site pairs are modest; the paper should report confidence intervals or bootstrap variability for the near-100% confounded accuracies and for the multi-site degradation cases so that the strength of the shortcut claim is quantified rather than left as point estimates.","section":null}],"minor_comments":[{"comment":"Figure numbering and cross-references are occasionally inconsistent (e.g., main-text references to Figure 2/3/4 vs. appendix figures); a single pass to align captions and in-text citations would help.","section":null},{"comment":"Several appendix tables (PERMANOVA Table 11, full prediction Tables 12–16) are essential to the argument; consider moving a compact summary of site vs. diagnosis effect sizes into the main text for readers who do not open the appendix.","section":null},{"comment":"Clarify the exact PCA dimensionality used before LDA/classifiers in each experiment (main text mentions 20 components in places; free parameter should be fixed and stated once).","section":null},{"comment":"Minor typographical issues (e.g., spacing around PERMANOVA, occasional missing spaces after periods) and a few incomplete sentences in the harmonization paragraph of §4 should be cleaned.","section":null}],"recommendation":"minor_revision","confidential_remarks":"Fits a non-archival workshop well. The empirical package is stronger than many similar embedding audits; I would not block on the frozen-protocol limitation if the authors add the explicit scope sentence and uncertainty on the confounded tables. No novelty or citation concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean, useful empirical note. The new result is straightforward: frozen subject-level embeddings from BrainLM (CLS) and SwiFT (window average) on three standard multi-site rs-fMRI cohorts (FBIRN, ADHD-200, ABIDE I) encode batch/site structure that frequently swamps diagnosis structure. That is not a rehash of the old FNC/ComBat literature; it is the first systematic check of the same problem inside these two transformer latents, plus a short architecture-linked decoding result (BrainLM closer to ALFF-style regional activity, SwiFT closer to FNC-style interactions).\n\nWhat they do well is the multi-pronged design. PCA/LDA site clustering, PERMANOVA site-vs-diagnosis, linear and kernel classifiers, within-site vs multi-site vs deliberately confounded splits (the last hitting ~100% accuracy while pooling can hurt diagnosis), ComBat before/after, and ALFF/FNC decoding all point the same way. Demographic regression is reported and does not erase the site structure. Comparison to handcrafted FNC is fair and shows the expected contrast: FNC keeps more diagnosis signal; the foundation embeddings keep more site signal. Citations are standard and light on self-citation. Circularity is essentially zero.\n\nSoft spots are real but proportionate. The probe is frozen public checkpoints with conventional CLS/average extraction—no fine-tuning, no site-aware pre-training, only two models, only ComBat. That is transparent and conventional for an embedding audit, not a hidden flaw, but it does limit how far you can generalize to “what foundation models can learn.” Sample sizes in some within-site cells are small, and there is no code release in the text. None of that overturns the dominance finding under the stated protocol.\n\nThis is for people building or evaluating neuroimaging foundation models, multi-site pipelines, or privacy-aware sharing. It is not a methods breakthrough and not a clinical tool. For a non-archival workshop paper it is exactly the right scope: honest documentation of a practical failure mode that changes how you should read zero-shot or embedding-based claims. I would send it to peer review, bring it to reading group, and cite the batch-dominance result when discussing these models.","headline":"Solid empirical workshop paper showing frozen BrainLM/SwiFT embeddings carry site structure that often dominates diagnosis; the finding is well-supported and worth citing.","tokens_in":15128,"tokens_out":567,"would_cite":true,"duration_ms":5709,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Brain foundation model embeddings encode scanner and site differences more strongly than diagnosis.","keywords":["foundation models","fMRI","batch effects","site effects","embeddings","harmonization","BrainLM","SwiFT"],"falsifier":"If, after the same embedding extraction on the same multi-site cohorts, site classification accuracy fell to chance while diagnosis accuracy rose above that of classical functional connectivity, or if confounded site-by-diagnosis splits no longer produced near-perfect accuracy, the central claim would fail.","tokens_in":15226,"feed_emoji":"🧠","tokens_out":608,"duration_ms":5585,"temperature":0.7,"pith_summary":"This paper asks what subject-level embeddings from two frozen neuroimaging foundation models actually contain when applied to multi-site resting-state fMRI. Across three public cohorts the embeddings show clear batch structure: site identity can be recovered at high accuracy, PERMANOVA attributes more multivariate variance to site than to diagnosis, and deliberately confounded train/test splits reach near-perfect accuracy by exploiting site rather than disease. Standard ComBat harmonization applied after embedding extraction reduces site predictability but does not reliably strengthen diagnosis signals. The two models also differ in the biological features they preserve: one aligns more with regional activity amplitude, the other with inter-region connectivity, matching their architectures. The practical message is that high downstream accuracy on multi-site fMRI can be an artifact of acquisition shortcuts, so batch-aware evaluation and representation design are required before these embeddings can be trusted as clinical biomarkers.","feed_headline":"Brain AI embeddings track scanners more than diagnosis","feed_subtitle":"Multi-site fMRI tests show site shortcuts dominate clinical signal; harmonization helps only partly","key_machinery":"A multi-pronged evaluation of frozen subject-level embeddings (CLS token for BrainLM; average of temporal-window embeddings for SwiFT) that combines PCA/LDA visualization, PERMANOVA, site-versus-diagnosis classifiers, deliberately confounded train/test splits, post-hoc ComBat, and decoding of ALFF versus FNC.","core_discovery":"Subject-level embeddings extracted from two frozen neuroimaging foundation models (BrainLM and SwiFT) encode substantial acquisition-related batch effects that often dominate diagnosis-related information across multi-site fMRI datasets; harmonization can attenuate the batch signal without necessarily amplifying clinical signal, and the two models preferentially encode different biological summaries consistent with their architectures.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Brain foundation embeddings capture scanners more than diagnoses","Batch effects dominate diagnosis signals in neuroimaging AI embeddings","Harmonization only partly tames site bias in BrainLM and SwiFT","Frozen brain models encode acquisition noise over clinical info","Site shortcuts outrank diagnosis in multi-site fMRI foundation embeddings"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That taking the CLS token or averaging window embeddings from already-trained, frozen models, without any task-specific fine-tuning or site-aware pre-training, is a fair and sufficient way to measure what the models have learned about biology versus acquisition.","fun_headline_variants_meta":{"raw":{"variants":["Brain foundation embeddings capture scanners more than diagnoses","Batch effects dominate diagnosis signals in neuroimaging AI embeddings","Harmonization only partly tames site bias in BrainLM and SwiFT","Frozen brain models encode acquisition noise over clinical info","Site shortcuts outrank diagnosis in multi-site fMRI foundation embeddings"]},"model":"grok-4.5","effort":"low","cost_usd":0.002784,"raw_usage":{"total_tokens":976,"prompt_tokens":672,"num_sources_used":0,"completion_tokens":83,"cost_in_usd_ticks":27840000,"prompt_tokens_details":{"text_tokens":672,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":221,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":672,"tokens_out":83,"duration_ms":3139,"temperature":1.0,"reasoning_tokens":221,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T20:17:18.187481+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If, after the same embedding extraction on the same multi-site cohorts, site classification accuracy fell to chance while diagnosis accuracy rose above that of classical functional connectivity, or if confounded site-by-diagnosis splits no longer produced near-perfect accuracy, the central claim would fail.","supporting_citations":[],"review_version":2}