REVIEW 2 major objections 6 minor 45 references
EEG foundation-model clinical gains shrink or reverse once evaluation unit, population shift, strong classical baselines, and negative controls are taken seriously.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 12:33 UTC pith:IGP2XJR4
load-bearing objection A careful evaluation paper with real controls and one load-bearing soft spot: the CAUEEG failure is not cleanly isolated from wrapper mismatch. the 2 major comments →
Stress-Testing EEG Foundation Models for Clinical Decoding: Dataset Identity and Targeted Negative Controls
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Conclusions about EEG foundation-model gains on clinical tasks depend strongly on evaluation unit, dataset shift, comparator strength, and targeted controls. In particular, frozen REVE fails a Korean dementia stress test relative to classical features, carries readily decodable dataset identity, and can be beaten by random initialization, while its clearest controlled positive is narrow cross-subject ictal detection rather than broad clinical transfer.
What carries the argument
A harmonized frozen linear-probe benchmark plus a targeted negative-control suite—random initialization, random features, label permutation, scrambled-label adaptation, and projection sensitivity—used to separate pretrained representation quality from architecture, probe capacity, label dependence, and dataset identity.
Load-bearing premise
The Korean failure claim rests on recording-level splits plus one patient-disjoint held-out bound, even though public patient IDs are missing and dataset-identity probes cannot separate site from hardware, preprocessing, population, and clinical mix.
What would settle it
Re-run the CAUEEG dementia comparison with true patient-grouped splits, matched montage and preprocessing across cohorts, and the same classical and random-init controls: if pretrained REVE then reliably beats both classical features and random initialization on diagnosis while dataset identity collapses, the cross-population failure claim fails.
If this is right
- A benchmark gain on frozen EEG embeddings should not be read as transferable clinical signal until subject leakage, pretraining exposure, and population shift are audited.
- Classical spectral baselines and random-initialization controls should be required comparators before claiming foundation-model advantage.
- Dataset-identity probes can expose when embeddings mainly separate recording conditions rather than disease labels.
- Subject-level aggregation can reverse epoch-level model rankings on Alzheimer’s-style tasks.
- The practical product is a checklist: match evaluation unit to clinical decision, strengthen comparators, and choose controls for the specific alternative explanation.
Where Pith is reading between the lines
- If dataset identity dominates leading embedding directions across many public EEG corpora, multi-site foundation pretraining may need explicit site-invariance objectives rather than scale alone.
- Regulatory or trial use of EEG foundation models would likely need locked subject-level protocols and external population stress tests, not only within-corpus frozen probes.
- The random-init win on Korean dementia suggests some reported pretraining benefits may be corpus-specific feature shaping that can actively hurt under shift.
- Ictal detection may remain the most forgiving clinical probe for current EEG foundation models because same-session negatives reduce session confounds that plague trait diagnosis.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript benchmarks six pretrained EEG foundation models (LaBraM, EEGMamba, CBraMod, REVE, BENDR, BIOT) on five clinical tasks across four public datasets under frozen linear probing, with targeted negative controls applied mainly to REVE. Central claims: (1) on the Korean CAUEEG dementia cohort, frozen REVE trails classical spectral features by ~20–30 pp AUROC (0.568 vs 0.769 3-way, recording-level CV; 0.565 vs 0.768 on the authors' patient-disjoint held-out split); (2) dataset identity is near-perfectly decodable from frozen REVE embeddings (AUROC 1.000 at PCA-50, 0.9998 band-limited) while the same pipeline decodes Korean diagnosis at 0.528; (3) a randomly initialised REVE encoder outperforms pretrained REVE on CAUEEG (0.659 vs 0.570); (4) on ds004504 Alzheimer's, the nominal epoch-level REVE advantage disappears under subject-level aggregation and stronger comparators; (5) the clearest controlled positive is CHB-MIT cross-subject ictal detection (REVE 0.793, +9.2 pp over random-init, permutation at chance). The paper frames these as evidence that EEG-FM conclusions depend on evaluation unit, dataset shift, comparator strength, and controls.
Significance. If the results hold, this is a useful corrective to a fast-moving literature. The paper's strengths are concrete: (i) a genuine patient-disjoint leakage bound on CAUEEG derived from the authors' no-overlap split (gap narrows from 0.337 to 0.286 binary and 0.252 to 0.203 3-way, Table 9), which largely settles the recording-level leakage concern; (ii) a real negative-control suite (random initialisation with three seeds, label permutation, scrambled-label fine-tuning, GRP-vs-PCA projection sensitivity) rather than a single comparator; (iii) a controlled positive on CHB-MIT with same-session negatives, guard-band and separate-recording sensitivities, and honest disclosure that the retired 0.920 headline was replaced by 0.793 under a preregistered correction; (iv) consistent subject-level aggregation discipline on ds004504 that reverses the epoch-level ordering; and (v) unusually candid statistical self-reporting (pseudoreplication in the 15-pair Wilcoxon is flagged by the authors themselves). The dataset-identity probes (AUROC 1.000 at PCA-50 vs 0.528 for diagnosis under the identical pipeline) are a crisp, falsifiable demonstration that dataset shift dominates the leading variance.
major comments (2)
- [§4.6, Table 9] §3.3, §4.6 (and contrast §3.1 BENDR caveat, Limitation 5): the headline CAUEEG result (frozen REVE 0.568 vs classical 0.769, persisting at 0.565 vs 0.768 patient-disjoint) is obtained entirely through the harmonized pipeline (200 Hz, 0.5–70 Hz FIR, CAR, 4 s epochs, per-epoch z-score, ±8 SD clip), which is not documented to match REVE's pretraining preprocessing. The manuscript applies precisely this reasoning to quarantine BENDR ('confounded with a bandwidth and channel-count mismatch... attribution to its contrastive objective is suggestive', §3.1) and extends it to LaBraM (Limitation 5), but no equivalent wrapper-sensitivity analysis is reported for REVE on the one cohort that carries the central claim. The random-init control (0.659 vs 0.570) rules out a purely architectural failure but is compatible with the alternative that pretrained weights, calibrated to REVE's native filtering/z
- [§4.6, Architectural-bias control] §4.6, 'Architectural-bias control': the claim that random initialisation exceeds pretrained REVE rests on 3 seeds × 5 folds sharing the same fold partitions. The manuscript itself notes the Wilcoxon p=6.1e-5 over all 15 pairs is anticonservative through pseudoreplication, and that the fold-level two-sided test floors at p=0.0625 (n=5). Given that 'a randomly-initialised encoder outperforms pretrained REVE (0.659 vs 0.570)' appears in the abstract with no uncertainty statement, the effective statistical unit (5 folds; unanimous direction, mean gap +8.9 pp) should be stated alongside the point estimates wherever the claim is made, and ideally strengthened with a recording-level bootstrap within folds. Relatedly, the interpretive sentence 'REVE's Western pretraining does not add, and numerically removes, Korean-dementia discriminative signal' should be made conditional on the evaluated
minor comments (6)
- [§4.1, Table 3] Table 3 and §4.1: LaBraM (79.2±10.6) is nominally tied with REVE (79.3±18.3) on CHB-MIT, but the random-init/permutation/random-feature controls were applied only to REVE. The conclusion 'the clearest controlled positive is cross-subject ictal detection' is accurate as phrased, but the Discussion should state explicitly that no equivalent controlled positive (or negative) can yet be asserted for LaBraM.
- [§3.2, CHB-MIT methods notes] §3.2 CHB-MIT methods note (ii): per-epoch z-scoring removes absolute amplitude from both the encoder inputs and the classical features, handicapping the classical comparator on ictal detection. A cheap amplitude-only baseline (e.g., logistic regression on per-epoch RMS or peak-to-peak amplitude) would bound how much of the 70.0% classical AUROC is recoverable from amplitude alone and would sharpen the interpretation of the +9.3 pp margin; the random-init control is unaffected by this concern.
- [Data and Code Availability] Data and Code Availability: for a benchmark whose stated contribution is a reusable control suite, 'code... can be supplied privately... and will be released publicly upon publication' is weaker than the manuscript's reproducibility framing (per-fold split indices, machine-readable evidence registry) warrants. Depositing the pipeline configuration and split indices in a tagged public repository at revision would materially increase the paper's value.
- [§3.5, §4.3] §3.5: the ds004504 C-sweep was 'run on an earlier preprocessing of that cohort', so it bounds sensitivity rather than giving comparable values; this is disclosed, but the same caveat should be repeated where the swept values are cited in §4.3. Similarly, the IAF-adjusted classical sensitivity (82.1% vs 85.1%) uses an unnotched earlier run and should be labeled as internally matched only, as the text already does in §4.3 but not in the abstract-adjacent summary.
- [§4.6, dataset identity] §4.6 dataset-identity probe: the ds004504-vs-CAUEEG probe compares n=88 against n=1,187; a label-permutation null for the identity probe (analogous to the CHB-MIT permutation control) would confirm that AUROC 1.000 is not partly a probe-capacity artifact at extreme class imbalance, complementing the band-limited 0.9998 replication.
- [Throughout] Typesetting/consistency: several section headings show letterspacing artifacts ('T argeted controls', 'F requency-band ablation', 'F rozen linear probing'); Figure 1's '∆ = +9.3 pp' (vs classical) and the abstract's '+9.2 pp' (vs random-init) are both correct but should be labeled by comparator to avoid confusion; §4.6 '8 s-epochN=150' is missing a space.
Circularity Check
Empirical benchmark against external datasets and public checkpoints; no derivation that re-labels fitted inputs as predictions.
full rationale
This paper is a multi-model, multi-dataset empirical stress test of frozen EEG foundation-model probes, not a first-principles derivation. Its load-bearing claims (CAUEEG classical-over-REVE gap, near-ceiling dataset-identity decode, random-init beating pretrained REVE, CHB-MIT ictal positive with matched controls) are measured quantities on public cohorts and public checkpoints, compared against classical features, random initialisation, random features, label permutation, and scrambled-label adaptation. None of these results is algebraically forced by a fitted parameter renamed as a prediction, nor justified by a uniqueness theorem or ansatz imported from overlapping-author prior work. The companion manuscript [44] appears only as a bibliographic companion and is not used to underwrite the present conclusions. Pipeline-mismatch and recording-level-split concerns affect external validity, not circularity of the derivation chain. Score 0 with empty steps is therefore the correct finding.
Axiom & Free-Parameter Ledger
free parameters (4)
- Logistic-probe regularization C =
0.1 or 1.0 by task family
- PCA/GRP projection rank =
50 or 200 components
- Epoch length and labeling thresholds =
4.0 s; ≥50% overlap; 1:3 ratio
- Preprocessing band and normalization =
0.5–70 Hz; per-epoch z-score; clip ±8 SD
axioms (5)
- domain assumption Frozen linear-probe performance is a valid primary measure of representation quality for EEG foundation models.
- domain assumption Classical handcrafted spectral/Hjorth/entropy features are an appropriate minimal comparator for marginal value of pretrained embeddings.
- ad hoc to paper Where patient IDs are unavailable, recording-level CV plus an author patient-disjoint held-out split can bound leakage enough for a transfer-failure claim.
- domain assumption Same-session interictal negatives make CHB-MIT a test of ictal state decoding rather than session confounds.
- domain assumption Published pretraining corpus lists are sufficient to label tasks in-domain or out-of-domain for each encoder.
read the original abstract
Pretrained EEG foundation models are increasingly proposed for clinical decoding, but their transfer across populations and robustness to negative controls remain unclear. We benchmark six models (LaBraM, EEGMamba, CBraMod, REVE, BENDR, and BIOT) on five clinical tasks across four datasets using frozen linear probes with leave-one-subject-out, subject-grouped, or explicitly identified recording-level splits. Selected REVE findings are tested against random initialisation, random features, label permutation, scrambled-label fine-tuning, and projection sensitivity. On Korean dementia (CAUEEG, three-way), frozen REVE reaches 0.568 AUROC versus 0.769 for classical features; the ordering persists on a patient-disjoint held-out split (0.565 versus 0.768). Dataset identity is readily decoded from frozen embeddings (AUROC 1.000 at PCA-50; 0.9998 after band restriction and per-epoch z-scoring), whereas the same PCA-50 pipeline decodes Korean diagnosis at 0.528. A randomly initialised encoder also outperforms pretrained REVE on this task (0.659 versus 0.570). On Alzheimer's disease, Gaussian random projection and PCA of the same pretrained embeddings perform similarly, and classical features nominally exceed REVE at the subject level. The clearest controlled positive is cross-subject ictal detection on CHB-MIT (n=23), where REVE achieves 0.793 AUROC, 9.2 percentage points above a randomly initialised encoder. These results show that EEG foundation-model conclusions depend strongly on evaluation unit, dataset shift, comparator strength, and targeted controls.
Figures
Reference graph
Works this paper leans on
-
[1]
A., Grabowski, H
DiMasi, J. A., Grabowski, H. G., & Hansen, R. W. (2016). Innovation in the pharmaceutical industry: New estimates of R&D costs.Journal of Health Economics, 47, 20–33
2016
-
[2]
W., Craighead, J
Hay, M., Thomas, D. W., Craighead, J. L., et al. (2014). Clinical development success rates for investigational drugs.Nature Biotechnology, 32(1), 40–51
2014
-
[3]
& Barachant, A
Jayaram, V. & Barachant, A. (2018). MOABB: Trustworthy algorithm benchmarking for BCIs.Journal of Neural Engineering, 15(6), 066011
2018
-
[4]
Bommasani, R., et al. (2021). On the opportunities and risks of foundation models. arXiv:2108.07258
Pith/arXiv arXiv 2021
-
[5]
He, K., Chen, X., Xie, S., et al. (2022). Masked autoencoders are scalable vision learners. Proc. CVPR, pp. 16000–16009
2022
-
[6]
B., Zhao, L
Jiang, W. B., Zhao, L. M., & Lu, B. L. (2024). Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCI.Proc. ICLR(Spotlight)
2024
-
[7]
Wang, J., et al. (2025). EEGMamba: An EEG Foundation Model with Mamba.Neural Networks, 192, 107816
2025
-
[8]
Wang, J., et al. (2025). CBraMod: A Criss-Cross Brain Foundation Model for EEG Decoding. Proc. ICLR
2025
-
[9]
El Ouahidi, Y., Lys, J., Th¨ olke, P., Farrugia, N., Pasdeloup, B., Gripon, V., Jerbi, K., & Lioi, G. (2025). REVE: A Foundation Model for EEG, Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects.Advances in Neural Information Processing Systems (NeurIPS). arXiv:2510.21585
arXiv 2025
-
[10]
Wu, J., et al. (2025). AdaBrain-Bench: Benchmarking Brain Foundation Models for Brain- Computer Interface Applications.arXiv:2507.09882
Pith/arXiv arXiv 2025
-
[11]
Wang, Y., Huang, N., Mammone, N., Cecchi, M., & Zhang, X. (2025). LEAD: An EEG Foundation Model for Alzheimer’s Disease Detection.arXiv preprint arXiv:2502.01678 [cs.LG, eess.SP]
arXiv 2025
-
[12]
Kostas, D., Aroca-Ouellette, S., & Rudzicz, F. (2021). BENDR: Using transformers and a contrastive self-supervised learning task to learn from massive amounts of EEG data. Frontiers in Human Neuroscience, 15, 653659
2021
-
[13]
B., & Sun, J
Yang, C., Westover, M. B., & Sun, J. (2023). BIOT: Biosignal Transformer for Cross-data Learning in the Wild.Advances in Neural Information Processing Systems (NeurIPS), 36, 78240–78260
2023
-
[14]
Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020). A simple framework for contrastive learning of visual representations.Proc. ICML, pp. 1597–1607
2020
-
[15]
& Dao, T
Gu, A. & Dao, T. (2024). Mamba: Linear-time sequence modeling with selective state spaces.Proc. COLM
2024
-
[16]
Dao, T. & Gu, A. (2024). Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality.Proc. ICML
2024
-
[17]
Kuruppu, G., Wagh, N., Kremen, V., Pati, S., Worrell, G., & Varatharajah, Y. (2025). EEG Foundation Models: A Critical Review of Current Progress and Future Directions. arXiv:2507.11783. 21
arXiv 2025
-
[18]
Liu, D., Chen, Y., Chen, Z., Cui, Z., Wen, Y., An, J., Luo, J., & Wu, D. (2026). EEG Foundation Models: Progresses, Benchmarking, and Open Problems.arXiv:2601.17883
arXiv 2026
-
[19]
A., Laskaris, N., & Zafeiriou, S
Lee, N., Bakas, S., Barmpas, K., Panagakis, Y., Adamos, D. A., Laskaris, N., & Zafeiriou, S. (2025). Assessing the Capabilities of Large Brainwave Foundation Models.2025 IEEE 35th International Workshop on Machine Learning for Signal Processing (MLSP)
2025
-
[20]
Wang, X., Yang, Y., & Coyle, D. (2026). EEG-FM-Audit: A Systematic Evaluation and Analysis Pipeline for EEG Foundation Models.arXiv:2605.26910
Pith/arXiv arXiv 2026
-
[21]
Kastrati, A., B¨ urki, J., Lauer, J., Xuan, C., Iaquinto, R., & Wattenhofer, R. (2025). EEG- Bench: A Benchmark for EEG Foundation Models in Clinical Applications.Foundation Models for the Brain and Body Workshop (BrainBodyFM), NeurIPS 2025. arXiv:2512.08959
arXiv 2025
-
[22]
G., Karaiskou, A.-I., Gagliardi, G., Strypsteen, T., Badiei, M
Kontras, K., Osselaer, T., Mouslech, S. G., Karaiskou, A.-I., Gagliardi, G., Strypsteen, T., Badiei, M. H., Rani, A., Vanmarcke, M., Bhagubai, M., Ekbote, C., Hwang, J., Chatzichristos, C., Liang, P. P., & De Vos, M. (2026). NeuroAtlas: Benchmarking Foundation Models for Clinical EEG and Brain-Computer Interfaces.arXiv:2605.14698
Pith/arXiv arXiv 2026
-
[23]
Xiong, W., Li, J., Li, J., Zhu, K., & Jiang, C. (2026). EEG-FM-Bench: A Comprehensive Benchmark for the Systematic Evaluation and Diagnostic Analyses of EEG Foundation Models.International Conference on Machine Learning (ICML 2026). arXiv:2508.17742
arXiv 2026
-
[24]
Lu, Z., Li, Z., Shen, X., Lou, K., Xin, Y., Chen, X., Wang, S., Chen, X., Fan, J., Huang, C., Xu, X., Hou, Z., Wei, C., & Liu, Q. (2026). OmniEEG-Bench: A Standardized Evaluation Benchmark for EEG Foundation Models.arXiv:2606.00815
Pith/arXiv arXiv 2026
-
[25]
ˇSirca, U., Alimardani, M., Zafeiriou, S., & Barmpas, K. (2026). Beyond Accuracy: Robust- ness, Interpretability and Expressiveness of EEG Foundation Models.arXiv:2605.17562
Pith/arXiv arXiv 2026
-
[26]
T., et al
Schirrmeister, R. T., et al. (2017). Deep learning with convolutional neural networks for EEG decoding and visualization.Human Brain Mapping, 38(11), 5391–5420
2017
-
[27]
Dosovitskiy, A., et al. (2021). An image is worth 16x16 words: Transformers for image recognition at scale.Proc. ICLR
2021
-
[28]
Shoeb, A. H. (2009). Application of machine learning to epileptic seizure onset detection and treatment.PhD thesis, MIT
2009
-
[29]
& Picone, J
Obeid, I. & Picone, J. (2016). The Temple University Hospital EEG data corpus.Frontiers in Neuroscience, 10, 196
2016
-
[30]
Miltiadous, A., et al. (2023). A dataset of scalp EEG recordings of Alzheimer’s disease, frontotemporal dementia and healthy subjects from routine EEG.Data, 8(6), 95
2023
-
[31]
Kemp, B., et al. (2000). Analysis of a sleep-dependent neuronal feedback loop.IEEE Trans. Biomedical Engineering, 47(9), 1185–1194
2000
-
[32]
C., & Paik, J
Kim, M.-J., Youn, Y. C., & Paik, J. (2023). Deep learning-based EEG analysis to classify normal, mild cognitive impairment, and dementia: Algorithms and dataset.NeuroImage, 272, 120054
2023
-
[33]
Gramfort, A., et al. (2013). MEG and EEG data analysis with MNE-Python.Frontiers in Neuroscience, 7, 267
2013
-
[34]
Welch, P. (1967). The use of fast Fourier transform for the estimation of power spectra. IEEE Trans. Audio and Electroacoustics, 15(2), 70–73. 22
1967
-
[35]
Hjorth, B. (1970). EEG analysis based on time domain properties.Electroencephalography and Clinical Neurophysiology, 29(3), 306–310
1970
-
[36]
J., et al
Donoghue, T., Haller, M., Peterson, E. J., et al. (2020). Parameterizing neural power spectra into periodic and aperiodic components.Nature Neuroscience, 23(12), 1655–1665
2020
-
[37]
E., Li, C., & Rabinovic, A
Johnson, W. E., Li, C., & Rabinovic, A. (2007). Adjusting batch effects in microarray expression data using empirical Bayes methods.Biostatistics, 8(1), 118–127
2007
-
[38]
I., et al
Fortin, J.-P., Cullen, N., Sheline, Y. I., et al. (2018). Harmonization of cortical thickness measurements across scanners and sites.NeuroImage, 167, 104–120
2018
-
[39]
Botvinik-Nezer, R., et al. (2020). Variability in the analysis of a single neuroimaging dataset by many teams.Nature, 582, 84–88. DOI: 10.1038/s41586-020-2314-9
-
[40]
van Dijk, H., van Wingen, G., Denys, D., Olbrich, S., van Ruth, R., & Arns, M. (2022). The two decades brainclinics research archive for insights in neurophysiology (TDBRAIN) database.Scientific Data, 9, 333
2022
-
[41]
Cavanagh, J. F. (2021). EEG: 3-Stim Auditory Oddball and Rest in Parkinson’s Disease. OpenNeuro, dataset ds003490
2021
-
[42]
J., Shen, Y., Wallis, P., et al
Hu, E. J., Shen, Y., Wallis, P., et al. (2022). LoRA: Low-Rank Adaptation of Large Language Models.Proc. ICLR
2022
-
[43]
Kaplan, J., et al. (2020). Scaling laws for neural language models.arXiv:2001.08361
Pith/arXiv arXiv 2020
-
[44]
Zare, M. (2026). Foundation Models for EEG Are Blind to Long-Range Temporal Cor- relations: A Spectral–Temporal Dissociation Behind Their Cross-Population Fragility. Companion manuscript, under review, 2026
2026
-
[45]
L., et al
Goldberger, A. L., et al. (2000). PhysioBank, PhysioToolkit, and PhysioNet.Circulation, 101(23), e215–e220. 23
2000
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.