REVIEW 3 major objections 5 minor 14 references
The paper argues that EEG foundation models, whatever their front end, do not preserve long-range temporal correlations in their frozen embeddings: the alpha-envelope scaling exponent is absent even where the static 1/f spectrum is strongly
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 06:53 UTC pith:OMO7SZOF
load-bearing objection Careful audit with real controls that overstates its headline: the clean dissociation is spectral-input-only, while raw-waveform models are simply undecodable with the probes used. the 3 major comments →
Foundation Models for EEG Are Blind to Long-Range Temporal Correlations: A Spectral-Temporal Dissociation Behind Their Cross-Population Fragility
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that EEG foundation-model embeddings do not encode long-range temporal correlations even when they encode the static spectral shape of the signal. For the two spectral-input models, a probe on the frozen embedding recovers the aperiodic 1/f slope (R²=0.59–0.73) but not the alpha-envelope DFA exponent across cohorts; order-shuffling shows the residual DFA signal is order-independent, hence a static proxy rather than LRTC. An explicit DFA feature recovers the exponent at about half the measured reliability ceiling. The same frozen embeddings decode the recording site almost perfectly (0.98–1.00), while the DFA exponent is comparatively site-robust (0.71). On zero-s
What carries the argument
Central object: the DFA exponent of the alpha-band (8–13 Hz) amplitude envelope, computed over 0.5–30 s (or a shorter 0.5–2 s range in the short-cohort), a dimensionless slope summarizing power-law temporal correlations; it is invariant to amplitude, reference, and gain and exists only in the order of samples. The argument rides on three controls: (1) an order-preserving pre-pool probe comparing true token order vs. per-subject shuffle, which separates objective-level blindness from mean-pooling artifacts; (2) residualization of DFA on the 1/f slope, which rules out aperiodic shadowing; and (3) a leak-free two-site decoding task plus label-free harmonization, which separates site-identity le
Load-bearing premise
The load-bearing premise is that a learned probe on frozen embeddings is an adequate readout for all five architectures; the paper's own pre-pool control shows this probe is uninformative for two of the five raw-waveform models (both ordered and shuffled R² are negative), so the claim that none of the five represents LRTC is only as strong as that probe-adequacy premise.
What would settle it
Train any one of the five encoders with the paper's proposed LRTC-aware auxiliary loss, freeze it, and fit the same three-probe readout to the alpha-envelope DFA exponent; if R² does not move from about zero to above roughly 0.3 against the measured 0.64 reliability ceiling, the remedy fails. For the raw-waveform claim, show that a probe can recover the simpler 1/f slope from those same frozen embeddings; if it cannot, the 'none of five' blindness is confounded by probe weakness.
If this is right
- A downstream user who fits a light probe to a frozen EEG foundation-model embedding and targets an LRTC-linked outcome (e.g., dementia staging or depression) is building on a representation from which the relevant signal has already been removed.
- For spectral-input models, the apparent DFA recovery on one cohort is an order-independent static proxy: shuffling the token order does not reduce it, so it should not be read as temporal LRTC.
- The DFA exponent and the 1/f slope are empirically orthogonal (r=-0.06), so the dissociation is real rather than a relabeling of one axis.
- Recording-site identity is near-perfectly decodable from all five frozen embeddings (0.98–1.00 vs. 0.500 chance), whereas the discarded DFA exponent is site-robust (0.71), making site dominance a universal property of frozen EEG-FM embeddings.
- An LRTC-aware auxiliary loss, combining a scale-freeness term and an exponent-fidelity term, is proposed to move DFA recovery from about zero toward >0.3 and to improve zero-shot cross-population transfer.
Where Pith is reading between the lines
- The paper's own probe-adequacy caveats mean the 'none of five' statement is safest for the spectral-input models; for the three raw-waveform models, failing to recover even 1/f (R²≤0.12) leaves open that the readout, not the representation, is the bottleneck, so the raw-waveform blindness is better read as unproven than established.
- If the mechanism generalizes, it predicts that any feature living in the ordering of windows across tens of seconds—not just alpha-envelope DFA—will be absent from frozen embeddings trained on the same objectives; the paper gestures at this as a general caution.
- A direct test of the proposed fix is to freeze an encoder after adding the surrogate LRTC loss and probe for DFA; the paper's fixed-target negative control predicts that a non-per-recording slope target will not improve recovery, so the per-recording exponent-fidelity term is the load-bearing part of the remedy.
- The site-dominance finding suggests that harmonizing embeddings addresses only a first/second-moment batch effect; the paper shows harmonization cannot conjure a disease axis that was never encoded, so site removal is necessary but not sufficient for cross-population transfer.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper tests whether five frozen EEG foundation models (REVE, LaBraM, BENDR, CBraMod, BIOT) preserve long-range temporal correlations of the alpha-band envelope, measured by the DFA exponent, in their embeddings. Across two out-of-distribution dementia cohorts, the authors report that spectral-input models (CBraMod, BIOT) recover the static 1/f slope strongly (R² = 0.59–0.73) but not the DFA exponent across cohorts, while raw-waveform models (REVE, LaBraM, BENDR) recover neither target with the probes used. A classical hand-computed DFA feature recovers the exponent at R² = 0.32–0.38 against a split-half reliability ceiling of 0.64. Additional controls show that for REVE and CBraMod the failure is order-independent, and that all five embeddings decode recording site near-perfectly while the DFA feature is comparatively less site-decodable. A zero-shot transfer arm shows the frozen REVE embedding at chance while the DFA exponent transfers directionally but not at family-wise significance. The authors propose a differentiable LRTC-aware auxiliary pretraining loss as a remedy.
Significance. If the core dissociation holds, the finding is significant for EEG foundation-model evaluation: it identifies a disease-relevant temporal-scaling property that current pretraining objectives do not preserve, and it offers a concrete, falsifiable remedy. The paper has notable methodological strengths: a split-half reliability ceiling (R² = 0.64), an N = 770 replication, leak-free site-decoding controls with matched Gaussian nulls, exclusion of TDBRAIN due to REVE pretraining leakage, honest Holm correction over the full transfer family, and public code. The authors also explicitly flag many limitations, including the uninformative probe for raw-waveform models and the underpowered transfer arm. However, the title and abstract state a stronger universal claim than the evidence supports, and this overstatement is load-bearing for the proposed motivation.
major comments (3)
- [Abstract; §4.1, Table 2; §7] The headline claim 'None of the five FMs represented the LRTC in the temporal order' is not supported for the three raw-waveform models. In Table 2, REVE, LaBraM, and BENDR fail to recover both the DFA exponent and the 1/f slope (R² ≤ 0.12), and §7 states that for these three 'the probe is uninformative.' A null result from a probe that cannot recover a known spectral property cannot distinguish 'the embedding does not encode LRTC' from 'the readout cannot access it.' The clean spectral–temporal dissociation is established only for CBraMod and BIOT, which recover 1/f (R² = 0.59–0.73) but not DFA across cohorts. Because the abstract's universal blindness claim motivates the LRTC-aware objective in §5, the title and abstract should be revised to state the supported claim: spectral-input FMs encode the static 1/f slope but not the temporal DFA exponent, while raw-waveform FMs are not decoda
- [§4.2, Table 3] The order-preserving control does not resolve the objective-vs-pooling question for LaBraM and BENDR: both ordered and shuffled R² are negative, and LaBraM's paired gap is significantly negative (-0.090 ± 0.019, p = 0.009), which the paper correctly reads as probe instability. Consequently, the claim that the blindness is 'in the objective, not the pooling' can be asserted only for REVE and CBraMod; for LaBraM and BENDR the pre-pool representation remains untested. The paper mostly respects this in §4.2 and §7, but the Contribution 2 framing and the mechanistic discussion in §2/§5 should avoid implying the objective-level conclusion for all five architectures. Please add an explicit sentence in the Discussion that the objective-level cause is directly supported for REVE and CBraMod only.
- [Significance; §4.5, Table 5] The statement 'the discarded exponent transfers directionally where the frozen embedding is at chance' is not true for CBraMod. Table 5 shows CBraMod's raw W→K AUROC is 0.59 with bootstrap lower bound 0.55, above chance, and harmonization leaves it essentially unchanged; the paper itself identifies CBraMod as an unexplained exception. The Significance paragraph should say 'four of the five FMs' or otherwise explicitly exclude CBraMod from the claim that the frozen embedding is at chance. This matters because the abstract and Significance currently generalize a contrast that the paper's own Table 5 shows is not uniform.
minor comments (5)
- [§4.4, Table 4 caption] The phrase 'above the rule' in the caption is cryptic. Since the table lists all six pairwise and three leave-one-cohort-out directions, simply state that explicitly and delete 'above the rule.'
- [§5, Scale budget] The dyadic window sizes s ∈ {2,4,8,16,32} on about 60 tokens give only 3 or 1 non-overlapping windows at s = 16 and 32, respectively. The paper's §3 caveat about the noisiest top-scale point applies even more severely here; please state the minimum number of windows required for stable log F(s) estimates at the largest scales.
- [Table 2] BENDR's CAUEEG columns use N = 200 while the other models use N = 770. The text says the near-zero recovery is unaffected by sample size, but a sentence explicitly stating that the BENDR null was also checked at N = 770 (or explaining why it cannot be) would strengthen the table's comparability.
- [§4.5] The matched null for the DFA site-decoding probe (balanced accuracy 0.71 vs. '1.4× its matched null') is not described with the same detail as the embedding null. Please specify the null construction (e.g., Gaussian noise shaped like the 19-d DFA vector, same pipeline) so the reader can assess whether 0.71 is meaningfully above the null.
- [§4.4 Scope note] The companion manuscript is mentioned in the text but is not listed in the references. If it is under review elsewhere, please add a citation or state its status to avoid ambiguity about the source of the predicted random-initialisation control.
Circularity Check
No circularity: the DFA null is an empirical probe of frozen embeddings; the paper's own limitations narrow the claim but no result is equivalent to its input.
full rationale
The derivation chain is self-contained. The DFA target is an external scalar computed from raw EEG (alpha-band envelope DFA), and the five embeddings are frozen pretrained representations; no target information is used in pretraining or embedding extraction. The encoding probes are 5-fold cross-validated ridge/GBM/RF fits, so the null (DFA R^2≈0) is an empirical measurement. The classical DFA/1/f feature is explicitly labeled a positive control, not a prediction ('A classical DF A/1/f feature vector serves as the positive control (it must recover the target, and does)' — §3 Encoding probe), so its near-self-prediction is not a load-bearing derived result. The pre-pool order control is a shuffle-based comparison rather than a fitted input. The paper itself flags the probative-adequacy limits of its null: the Abstract states 'for these three the probe is uninformative,' and §7 Limitations states 'The pre-pool order control is uninformative for LaBraM and BENDR.' These caveats weaken the breadth of the 'none of five' phrasing, but they are inference limitations, not instances of an output being constructed from its input. The only self-referential passage is the companion-manuscript scope note (§4.4), which is explicitly not used as evidence ('Here we use transfer only to establish the cost of the encoding failure'); it is not load-bearing. No self-citation chain, uniqueness theorem, or ansatz-by-citation carries the argument.
Axiom & Free-Parameter Ledger
free parameters (3)
- DFA scale range =
0.5–30 s (16 log-spaced scales); 0.5–2 s (8 scales) for BrainLat
- Alpha band and filter order =
8–13 Hz, zero-phase order-4 Butterworth
- Readout/probe hyperparameters =
ridge, gradient boosting, random forest (pooled); 1D-CNN (pre-pool); PCA-50 for site probe
axioms (4)
- domain assumption The alpha-envelope DFA exponent is a valid, dimensionless measure of LRTC and is not shadowed by the 1/f slope (r=-0.06 in CAUEEG).
- domain assumption All non-Western test cohorts are out-of-distribution for all five FMs.
- domain assumption A frozen-embedding readout that fails to decode a known feature indicates the representation lacks that feature.
- domain assumption DFA measurement reliability (split-half ceiling R²=0.64) bounds achievable probe R².
read the original abstract
Objective. Electroencephalography (EEG) foundation models (FMs) are trained to reconstruct or contrastively align short patches, then pooled into a fixed embedding. We tested whether these embeddings retained the long-range temporal correlations (LRTC) quantified by the detrended-fluctuation-analysis (DFA) exponent of the alpha-band envelope, and whether it governs cross-population transfer. Approach. We probed five EEG FMs spanning raw-waveform and spectral-input architectures (REVE, LaBraM, BENDR, CBraMod, BIOT) on two out-of-distribution cohorts, comparing recovery of the DFA exponent against the static 1/f aperiodic slope. Order-preserving and residualization controls tested for pooling or aperiodic shadowing. A montage-harmonized, zero-shot transfer task compared the frozen embedding with the DFA exponent across three cohorts (adding a Western reference). Main results. None of the five FMs represented the LRTC in the temporal order. Raw-waveform models (REVE, LaBraM, BENDR) recovered neither the DFA exponent nor the 1/f slope (R^2 <= 0.12); for these three the probe is uninformative, so the dissociation is specific to the spectral-input models (CBraMod, BIOT), which recovered 1/f strongly (R^2 = 0.59-0.73) but not DFA across cohorts. A classical DFA feature recovered the exponent (R^2 = 0.32-0.38 against a 0.64 reliability ceiling), and LRTC was orthogonal to the aperiodic slope (r = -0.06). On cross-population transfer, the frozen REVE embedding did not beat chance (W to K, 0.45) and the dimensionless DFA exponent transferred directionally but not at family-wise significance; the other four did not uniformly replicate it. All five were dominated by a recording-site axis (decodable at 0.98-1.00 vs. 0.500 chance), whereas the DFA exponent they discard is site-robust (0.71).
Figures
Reference graph
Works this paper leans on
-
[1]
Jasmeet Singh Bindra and Siddharth Panwar. A spectral audit framework reveals task-dependent aperiodic reliance across eeg and ecg deep learning.arXiv preprint arXiv:2606.08583,
-
[4]
Ard Kastrati, Josua B¨ urki, Jonas Lauer, Cheng Xuan, Raffaele Iaquinto, and Roger Wattenhofer
contrastive SSL needs structure-preserving regularization to retain temporal/spatial similarity. Ard Kastrati, Josua B¨ urki, Jonas Lauer, Cheng Xuan, Raffaele Iaquinto, and Roger Wattenhofer. Eeg-bench: A benchmark for eeg foundation models in clinical applications.arXiv preprint arXiv:2512.08959,
-
[8]
ASZED-153 / Nigerian Schizophrenia EEG Dataset; OAUTHC Ile-Ife. cf. NSzED arXiv:2311.18484. Pavel Prado, Vicente Medel, Raul Gonzalez-Gomez, Agust ´ ın Sainz-Ballesteros, Victor Vidal, Her- nando Santamar ´ ıa-Garc ´ ıa, Sebastian Moguilner, Jhonny Mejia, Andrea Slachevsky, Maria Isabel Behrens, et al. The brainlat project, a multimodal neuroimaging datas...
-
[9]
Urban ˇSirca, Maryam Alimardani, Stefanos Zafeiriou, and Konstantinos Barmpas. Beyond accu- racy: Robustness, interpretability and expressiveness of EEG foundation models.arXiv preprint arXiv:2605.17562,
-
[10]
argues poor head-only performance is largely a mean-pooling artifact recoverable with token-level embeddings; probes task accuracy, not temporal LRTC. Jianwei Tai. Pretrained, frozen, still leaking: Auditing cross-encoder attribute transfer in eeg foundation models.arXiv preprint arXiv:2606.09189,
-
[11]
What do eeg foundation models capture from human brain signals?arXiv preprint arXiv:2605.11410,
Ling Tang, Qian Chen, Jilin Mei, Houshi Xu, Quanshi Zhang, Jing Shao, Na Zou, Xia Hu, and Dongrui Liu. What do eeg foundation models capture from human brain signals?arXiv preprint arXiv:2605.11410,
-
[12]
Ye Tao, Bradley T. Baker, Yu Wu, Anand D. Sarwate, Sandeep Panta, Sergey Plis, and Vince D. Calhoun. Batch effects in brain foundation model embeddings.arXiv preprint arXiv:2604.14441,
-
[14]
arXiv:2412.07236. 20 Chaoqi Yang, M. Brandon Westover, and Jimeng Sun. Biot: Biosignal transformer for cross-data learning in the wild. InAdvances in Neural Information Processing Systems (NeurIPS),
-
[2021]
William Lehn-Schiøler, Magnus Ruud Kjær, Rahul Thapa, Magnus Guldberg Pedersen, Anton Mos- quera Storgaard, Nick Williams, Radu Gatej, Tue Lehn-Schiøler, Andreas Brink-Kjær, Sadasivan Puthusserypady, S´ andor Beniczky, James Zou, and Lars Kai Hansen. Mechanistic interpretabil- ity of EEG foundation models via sparse autoencoders.arXiv preprint arXiv:2605.13930,
-
[2022]
Jiquan Wang, Sha Zhao, Zhiling Luo, Yangxuan Zhou, Haiteng Jiang, Shijian Li, Tao Li, and Gang Pan
DOI 10.1038/s41597-022-01409-z. Jiquan Wang, Sha Zhao, Zhiling Luo, Yangxuan Zhou, Haiteng Jiang, Shijian Li, Tao Li, and Gang Pan. Cbramod: A criss-cross brain foundation model for eeg decoding. InInternational Conference on Learning Representations (ICLR),
-
[2023]
CAUEEG: Chung-Ang University Hospital EEG dataset. Aditya Kommineni, Emily Zhou, Kleanthis Avramidis, Simon Bock Segaard, Jeppe Roden M¨ unster, Andreas Peter Juhl Hansen, Takfarinas Medani, Tiantian Feng, Richard Leahy, and Shrikanth Narayanan. Aperiodic and low-frequency spectral bias in reconstruction-based eeg foundation models.arXiv preprint arXiv:26...
-
[2024]
Yiru Jiao, Sander van Cranenburgh, Simeon Calvert, and Hans van Lint
spotlight; arXiv:2405.18765. Yiru Jiao, Sander van Cranenburgh, Simeon Calvert, and Hans van Lint. Structure-preserving contrastive learning for spatial time series.arXiv preprint arXiv:2502.06380,
-
[2025]
Richard Hardstone, Simon-Shlomo Poil, Giuseppina Schiavone, Rick Jansen, Vadim V
arXiv:2510.21585. Richard Hardstone, Simon-Shlomo Poil, Giuseppina Schiavone, Rick Jansen, Vadim V. Nikulin, Huibert D. Mansvelder, and Klaus Linkenkaer-Hansen. Detrended fluctuation analysis: A scale- free view on neuronal oscillations.Frontiers in Physiology, 3:450,
-
[2026]
19 Jun-You Lin, Ying Choon Wu, and Tzyy-Ping Jung
amplitude recoverable from frozen EEG-FM embeddings but phase not, attributed to time- translation invariance of the pretraining objective. 19 Jun-You Lin, Ying Choon Wu, and Tzyy-Ping Jung. The identity trap in eeg foundation models: A diagnostic audit.arXiv preprint arXiv:2606.06647,
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.