REVIEW 3 major objections 5 minor
The paper argues that EEG foundation models, whatever their front end, do not preserve long-range temporal correlations in their frozen embeddings: the alpha-envelope scaling exponent is absent even where the static 1/f spectrum is strongly
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 06:53 UTC pith:OMO7SZOF
load-bearing objection Careful audit with real controls that overstates its headline: the clean dissociation is spectral-input-only, while raw-waveform models are simply undecodable with the probes used. the 3 major comments →
Cross-Cohort Spectral-Temporal Dissociation in Frozen EEG Foundation-Model Representations
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that EEG foundation-model embeddings do not encode long-range temporal correlations even when they encode the static spectral shape of the signal. For the two spectral-input models, a probe on the frozen embedding recovers the aperiodic 1/f slope (R²=0.59–0.73) but not the alpha-envelope DFA exponent across cohorts; order-shuffling shows the residual DFA signal is order-independent, hence a static proxy rather than LRTC. An explicit DFA feature recovers the exponent at about half the measured reliability ceiling. The same frozen embeddings decode the recording site almost perfectly (0.98–1.00), while the DFA exponent is comparatively site-robust (0.71). On zero-s
What carries the argument
Central object: the DFA exponent of the alpha-band (8–13 Hz) amplitude envelope, computed over 0.5–30 s (or a shorter 0.5–2 s range in the short-cohort), a dimensionless slope summarizing power-law temporal correlations; it is invariant to amplitude, reference, and gain and exists only in the order of samples. The argument rides on three controls: (1) an order-preserving pre-pool probe comparing true token order vs. per-subject shuffle, which separates objective-level blindness from mean-pooling artifacts; (2) residualization of DFA on the 1/f slope, which rules out aperiodic shadowing; and (3) a leak-free two-site decoding task plus label-free harmonization, which separates site-identity le
Load-bearing premise
The load-bearing premise is that a learned probe on frozen embeddings is an adequate readout for all five architectures; the paper's own pre-pool control shows this probe is uninformative for two of the five raw-waveform models (both ordered and shuffled R² are negative), so the claim that none of the five represents LRTC is only as strong as that probe-adequacy premise.
What would settle it
Train any one of the five encoders with the paper's proposed LRTC-aware auxiliary loss, freeze it, and fit the same three-probe readout to the alpha-envelope DFA exponent; if R² does not move from about zero to above roughly 0.3 against the measured 0.64 reliability ceiling, the remedy fails. For the raw-waveform claim, show that a probe can recover the simpler 1/f slope from those same frozen embeddings; if it cannot, the 'none of five' blindness is confounded by probe weakness.
If this is right
- A downstream user who fits a light probe to a frozen EEG foundation-model embedding and targets an LRTC-linked outcome (e.g., dementia staging or depression) is building on a representation from which the relevant signal has already been removed.
- For spectral-input models, the apparent DFA recovery on one cohort is an order-independent static proxy: shuffling the token order does not reduce it, so it should not be read as temporal LRTC.
- The DFA exponent and the 1/f slope are empirically orthogonal (r=-0.06), so the dissociation is real rather than a relabeling of one axis.
- Recording-site identity is near-perfectly decodable from all five frozen embeddings (0.98–1.00 vs. 0.500 chance), whereas the discarded DFA exponent is site-robust (0.71), making site dominance a universal property of frozen EEG-FM embeddings.
- An LRTC-aware auxiliary loss, combining a scale-freeness term and an exponent-fidelity term, is proposed to move DFA recovery from about zero toward >0.3 and to improve zero-shot cross-population transfer.
Where Pith is reading between the lines
- The paper's own probe-adequacy caveats mean the 'none of five' statement is safest for the spectral-input models; for the three raw-waveform models, failing to recover even 1/f (R²≤0.12) leaves open that the readout, not the representation, is the bottleneck, so the raw-waveform blindness is better read as unproven than established.
- If the mechanism generalizes, it predicts that any feature living in the ordering of windows across tens of seconds—not just alpha-envelope DFA—will be absent from frozen embeddings trained on the same objectives; the paper gestures at this as a general caution.
- A direct test of the proposed fix is to freeze an encoder after adding the surrogate LRTC loss and probe for DFA; the paper's fixed-target negative control predicts that a non-per-recording slope target will not improve recovery, so the per-recording exponent-fidelity term is the load-bearing part of the remedy.
- The site-dominance finding suggests that harmonizing embeddings addresses only a first/second-moment batch effect; the paper shows harmonization cannot conjure a disease axis that was never encoded, so site removal is necessary but not sufficient for cross-population transfer.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper tests whether five frozen EEG foundation models (REVE, LaBraM, BENDR, CBraMod, BIOT) preserve long-range temporal correlations of the alpha-band envelope, measured by the DFA exponent, in their embeddings. Across two out-of-distribution dementia cohorts, the authors report that spectral-input models (CBraMod, BIOT) recover the static 1/f slope strongly (R² = 0.59–0.73) but not the DFA exponent across cohorts, while raw-waveform models (REVE, LaBraM, BENDR) recover neither target with the probes used. A classical hand-computed DFA feature recovers the exponent at R² = 0.32–0.38 against a split-half reliability ceiling of 0.64. Additional controls show that for REVE and CBraMod the failure is order-independent, and that all five embeddings decode recording site near-perfectly while the DFA feature is comparatively less site-decodable. A zero-shot transfer arm shows the frozen REVE embedding at chance while the DFA exponent transfers directionally but not at family-wise significance. The authors propose a differentiable LRTC-aware auxiliary pretraining loss as a remedy.
Significance. If the core dissociation holds, the finding is significant for EEG foundation-model evaluation: it identifies a disease-relevant temporal-scaling property that current pretraining objectives do not preserve, and it offers a concrete, falsifiable remedy. The paper has notable methodological strengths: a split-half reliability ceiling (R² = 0.64), an N = 770 replication, leak-free site-decoding controls with matched Gaussian nulls, exclusion of TDBRAIN due to REVE pretraining leakage, honest Holm correction over the full transfer family, and public code. The authors also explicitly flag many limitations, including the uninformative probe for raw-waveform models and the underpowered transfer arm. However, the title and abstract state a stronger universal claim than the evidence supports, and this overstatement is load-bearing for the proposed motivation.
major comments (3)
- [Abstract; §4.1, Table 2; §7] The headline claim 'None of the five FMs represented the LRTC in the temporal order' is not supported for the three raw-waveform models. In Table 2, REVE, LaBraM, and BENDR fail to recover both the DFA exponent and the 1/f slope (R² ≤ 0.12), and §7 states that for these three 'the probe is uninformative.' A null result from a probe that cannot recover a known spectral property cannot distinguish 'the embedding does not encode LRTC' from 'the readout cannot access it.' The clean spectral–temporal dissociation is established only for CBraMod and BIOT, which recover 1/f (R² = 0.59–0.73) but not DFA across cohorts. Because the abstract's universal blindness claim motivates the LRTC-aware objective in §5, the title and abstract should be revised to state the supported claim: spectral-input FMs encode the static 1/f slope but not the temporal DFA exponent, while raw-waveform FMs are not decoda
- [§4.2, Table 3] The order-preserving control does not resolve the objective-vs-pooling question for LaBraM and BENDR: both ordered and shuffled R² are negative, and LaBraM's paired gap is significantly negative (-0.090 ± 0.019, p = 0.009), which the paper correctly reads as probe instability. Consequently, the claim that the blindness is 'in the objective, not the pooling' can be asserted only for REVE and CBraMod; for LaBraM and BENDR the pre-pool representation remains untested. The paper mostly respects this in §4.2 and §7, but the Contribution 2 framing and the mechanistic discussion in §2/§5 should avoid implying the objective-level conclusion for all five architectures. Please add an explicit sentence in the Discussion that the objective-level cause is directly supported for REVE and CBraMod only.
- [Significance; §4.5, Table 5] The statement 'the discarded exponent transfers directionally where the frozen embedding is at chance' is not true for CBraMod. Table 5 shows CBraMod's raw W→K AUROC is 0.59 with bootstrap lower bound 0.55, above chance, and harmonization leaves it essentially unchanged; the paper itself identifies CBraMod as an unexplained exception. The Significance paragraph should say 'four of the five FMs' or otherwise explicitly exclude CBraMod from the claim that the frozen embedding is at chance. This matters because the abstract and Significance currently generalize a contrast that the paper's own Table 5 shows is not uniform.
minor comments (5)
- [§4.4, Table 4 caption] The phrase 'above the rule' in the caption is cryptic. Since the table lists all six pairwise and three leave-one-cohort-out directions, simply state that explicitly and delete 'above the rule.'
- [§5, Scale budget] The dyadic window sizes s ∈ {2,4,8,16,32} on about 60 tokens give only 3 or 1 non-overlapping windows at s = 16 and 32, respectively. The paper's §3 caveat about the noisiest top-scale point applies even more severely here; please state the minimum number of windows required for stable log F(s) estimates at the largest scales.
- [Table 2] BENDR's CAUEEG columns use N = 200 while the other models use N = 770. The text says the near-zero recovery is unaffected by sample size, but a sentence explicitly stating that the BENDR null was also checked at N = 770 (or explaining why it cannot be) would strengthen the table's comparability.
- [§4.5] The matched null for the DFA site-decoding probe (balanced accuracy 0.71 vs. '1.4× its matched null') is not described with the same detail as the embedding null. Please specify the null construction (e.g., Gaussian noise shaped like the 19-d DFA vector, same pipeline) so the reader can assess whether 0.71 is meaningfully above the null.
- [§4.4 Scope note] The companion manuscript is mentioned in the text but is not listed in the references. If it is under review elsewhere, please add a citation or state its status to avoid ambiguity about the source of the predicted random-initialisation control.
Circularity Check
No circularity: the DFA null is an empirical probe of frozen embeddings; the paper's own limitations narrow the claim but no result is equivalent to its input.
full rationale
The derivation chain is self-contained. The DFA target is an external scalar computed from raw EEG (alpha-band envelope DFA), and the five embeddings are frozen pretrained representations; no target information is used in pretraining or embedding extraction. The encoding probes are 5-fold cross-validated ridge/GBM/RF fits, so the null (DFA R^2≈0) is an empirical measurement. The classical DFA/1/f feature is explicitly labeled a positive control, not a prediction ('A classical DF A/1/f feature vector serves as the positive control (it must recover the target, and does)' — §3 Encoding probe), so its near-self-prediction is not a load-bearing derived result. The pre-pool order control is a shuffle-based comparison rather than a fitted input. The paper itself flags the probative-adequacy limits of its null: the Abstract states 'for these three the probe is uninformative,' and §7 Limitations states 'The pre-pool order control is uninformative for LaBraM and BENDR.' These caveats weaken the breadth of the 'none of five' phrasing, but they are inference limitations, not instances of an output being constructed from its input. The only self-referential passage is the companion-manuscript scope note (§4.4), which is explicitly not used as evidence ('Here we use transfer only to establish the cost of the encoding failure'); it is not load-bearing. No self-citation chain, uniqueness theorem, or ansatz-by-citation carries the argument.
Axiom & Free-Parameter Ledger
free parameters (3)
- DFA scale range =
0.5–30 s (16 log-spaced scales); 0.5–2 s (8 scales) for BrainLat
- Alpha band and filter order =
8–13 Hz, zero-phase order-4 Butterworth
- Readout/probe hyperparameters =
ridge, gradient boosting, random forest (pooled); 1D-CNN (pre-pool); PCA-50 for site probe
axioms (4)
- domain assumption The alpha-envelope DFA exponent is a valid, dimensionless measure of LRTC and is not shadowed by the 1/f slope (r=-0.06 in CAUEEG).
- domain assumption All non-Western test cohorts are out-of-distribution for all five FMs.
- domain assumption A frozen-embedding readout that fails to decode a known feature indicates the representation lacks that feature.
- domain assumption DFA measurement reliability (split-half ceiling R²=0.64) bounds achievable probe R².
read the original abstract
Objective. We tested whether frozen representations from five EEG foundation models support decoding of long-range temporal correlations, measured as the detrended-fluctuation-analysis (DFA) exponent of the alpha-band amplitude envelope. Approach. REVE, LaBraM, BENDR, CBraMod, and BIOT were evaluated in CAUEEG and BrainLat. A common 240 s estimator used 8-13 Hz filtering, DFA over 2-23.8 s, artifact masking, and quality control. One fixed nested-cross-validation readout predicted DFA and a fixed-mode aperiodic exponent. Controls tested pre-pool order sensitivity and aperiodic residualization. Results. CAUEEG included 764 recordings and BrainLat 79. BIOT decoded DFA in CAUEEG (R-squared = 0.232; conditional subject-bootstrap 95 percent interval, 0.121-0.310), and CBraMod was positive but imprecise (R-squared = 0.121; 0.003-0.214). Neither replicated in BrainLat, where all five point estimates were negative. In contrast, CBraMod and BIOT decoded the aperiodic exponent in both cohorts (R-squared = 0.459-0.757). BIOT remained positive after removal of the measured linear aperiodic association in matched CAUEEG data (R-squared = 0.240). The post-hoc order control was batch- and configuration-sensitive. Because chronological EEG epochs are not exchangeable, it was descriptive, not an LRTC-specific test. No revised DFA transfer direction passed source-label permutation testing. Cohort membership was near-ceiling decodable from all five embeddings, but this is not a pure site effect. Significance. CBraMod and BIOT show a replicated, model-specific spectral-temporal dissociation: aperiodic decoding is present in both cohorts, whereas alpha-envelope DFA decoding is cohort-dependent. These findings bound the evaluated readouts; they do not establish representational absence or an architectural cause. Transfer and clinical associations remain exploratory.
Figures
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.