{"id":"0d7192ab-02fe-484d-89cd-0dde4d8c4bb5","arxiv_id":"2508.15716","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that organizes EEG foundation-model research into five output-modality categories: native EEG, text, vision, audio, and multimodal fusion, with a claim to be the first such comprehensive taxonomy.","lead":"This survey reviews recent work using foundation models pretrained on non-EEG data (text, vision, audio) for EEG decoding, grouping papers into five modality-based categories. It is a useful reading map, but the paper's stated scope conflicts with some of the models it includes.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Scope violation undercuts central cross-domain claim: the survey's own Section II/Table II includes EEG-pretrained and EEG-fine-tuned models (BENDR, CBraMod, EEGM2, Large transformers), contradicting the 'exclusively non-EEG' inclusion criterion.","rationale":"The reader's conditional verdict is appropriate. I agree with their weakest_assumption: the inclusion criterion in Section I is the crux. This is a genuine internal inconsistency, not a matter of disputable boundaries. The paper's contribution is a readable taxonomy, but the headline claim of 'first comprehensive ... cross-domain' is falsified by its own corpus. Citation errors are secondary; the scope violation is the single most load-bearing issue because it changes what the survey claims to cover. The recommended test is cheap: an audit of Tables II-VI against the pretraining/fine-tuning rule. The fix is straightforward: narrow the claim to 'EEG foundation models, including but not limited to non-EEG pretrained,' or remove the EEG-pretrained rows. The contribution survives as a taxonomy but with corrected scope. No change to the reader's CONDITIONAL verdict is needed.","tokens_in":19351,"tokens_out":4062,"duration_ms":43610,"concrete_test":"Build an inclusion audit matrix for every row in Tables II-VI. For each cited model, record (i) the pretraining data domain (EEG vs non-EEG), as stated in the cited paper, and (ii) whether the cited method fine-tunes or adapts the pretrained model on EEG data. Confirm whether BENDR [29], CBraMod [32], EEGM2 [39], and Large transformers [44] satisfy the 'exclusively non-EEG, no EEG fine-tuning' rule; then check all other entries for the same two criteria. If any violation is found, the required edit is either to remove/reclassify EEG-pretrained entries or to revise the Abstract/Introduction to say the survey includes both EEG- and non-EEG-pretrained foundation models, deleting 'exclusively' and 'cross-domain' qualifiers as appropriate. A secondary pass should re-verify all reference keys against the bibliography (e.g., duplicate [117] entries) to avoid citation-induced misclassification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the first comprehensive modality-oriented taxonomy of foundation models in EEG analysis built on the scope statement in Section I: 'exclusively on foundation models that have been pre-trained on large-scale non-EEG data and directly applied to EEG analysis tasks' and 'exclude models fine-tuned on EEG datasets.' The premise is load-bearing: if the reviewed corpus includes models pretrained on EEG, the paper is not a map of cross-domain non-EEG transfer and the 'first comprehensive' novelty claim overreaches. The text itself violates the premise. Section II (unimodal decoding) includes BENDR [29], CBraMod [32], EEGM2 [39], and 'Large transformers are better EEG learners' [44]; each is pretrained on EEG data (self-supervised contrastive/Transformer/Mamba objectives) and, in several cases, fine-tuned on downstream EEG tasks. Section VII then describes the standard pipeline as 'self- or unsupervised pretraining ... from large-scale unlabeled EEG data,' directly acknowledging EEG-domain pretraining as part of the reviewed paradigm. This is not an external disagreement about what counts as a foundation model; it is an internal inconsistency between the paper's stated inclusion criterion and its own included corpus. The taxonomy's five output-modality categories may still be organizing, but the 'cross-domain' and 'exclusively non-EEG' qualifiers cannot stand as-is.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of foundation model applications in EEG analysis. It proposes a modality-oriented taxonomy with five categories: unimodal EEG decoding, EEG-to-text, EEG-to-vision, EEG-to-audio, and multimodal EEG fusion. The stated scope (Section I) is that only foundation models pretrained on large-scale non-EEG data and directly applied to EEG tasks are covered, and that models fine-tuned on EEG datasets are excluded. The paper describes representative works in each category, discusses architectural components such as EEG encoders and cross-modal alignment modules, and closes with challenges and future directions.","tokens_in":19745,"tokens_out":4422,"duration_ms":47772,"significance":"If the stated scope were consistently applied, the survey would fill a useful niche: it collects a large and recent body of work, organizes it into an intuitive five-category structure, and highlights important caveats such as the reliability concerns raised by Jo et al. [67] and the scaling-law observations of Banville et al. [82]. The taxonomy is easy to grasp and could serve as an entry point for researchers new to EEG foundation models. However, the central contribution is framed as a comprehensive map of cross-domain (non-EEG pretrained) transfer, and that framing is contradicted by the paper's own included corpus. The organizing value of the taxonomy survives, but the paper's main claim of an exclusively cross-domain survey does not.","major_comments":[{"comment":"The scope statement in Section I is load-bearing: the survey claims to cover 'exclusively' models pretrained on non-EEG data and to exclude models fine-tuned on EEG datasets. Section II contradicts this by including BENDR [29], which is pretrained via contrastive self-supervised learning on massive EEG data; CBraMod [32], an EEG-pretrained foundation model; EEGM2 [39], a self-supervised Mamba model for long-sequence EEG; and 'Large transformers are better EEG learners' [44], which is pretrained on EEG. These are in-domain, EEG-pretrained models, not cross-domain non-EEG models. The contradiction is internal, not a matter of divergent interpretations. The authors must either broaden the scope to include EEG-pretrained foundation models and adjust the 'cross-domain' claims accordingly, or remove/relabel the offending entries and re-audit the entire corpus against the stated criterion.","section":"§I vs. §II; Table II"},{"comment":"The discussion section explicitly describes the standard pipeline as 'self- or unsupervised pretraining ... from large-scale unlabeled EEG data' followed by 'task-specific fine-tuning,' and it lists adapters [44] and prompt tuning [34] as acceptable parameter-efficient adaptations. This directly acknowledges EEG-domain pretraining and EEG fine-tuning as part of the surveyed paradigm, contradicting the Section I exclusion. If the authors intend to distinguish parameter-efficient adaptation from full fine-tuning, that distinction is never made explicit, and in any case [44] and [34] are EEG-pretrained/fine-tuned systems. The boundary between 'cross-domain transfer' and 'EEG foundation models' needs to be redefined or the paper must be repositioned as covering both.","section":"§VII, training methodology paragraph"},{"comment":"The paper claims to be 'the first and latest comprehensive taxonomy' but provides no explicit methodology for literature selection: no search databases, no inclusion/exclusion criteria beyond the contradicted scope sentence, and no screening process. Given the rapid growth of the field, the 'comprehensive' claim is unverifiable and arguably overreaching. A short methodology paragraph or a table of inclusion decisions, especially explaining why some EEG-pretrained models are included while others are excluded, would address this. This is not purely cosmetic: the central contribution is the corpus/taxonomy, so the selection protocol is part of the evidence.","section":"§I and overall methodology"}],"minor_comments":[{"comment":"The reference list contains two identical entries for 'CineBrain' (J. Gao et al., arXiv:2503.06940). In Section VI.C, 'Wu et al. [117]' is cited for a hypergraph-based language model, but [117] is the duplicate CineBrain entry. This citation is incorrect and must be fixed.","section":"References [116] and [117]"},{"comment":"Several typos and formatting issues: 'demonstrating superior performance' (Section I) should be 'demonstrated'; 'Reconstrution' in the Section IV.B heading; stray spaces in arXiv identifiers such as 'arXiv: 2411.15395' and 'arXiv: 2501.17489'; inconsistent capitalization of 'modality oriented' in the index terms.","section":"Throughout"},{"comment":"The BART paper is by Lewis et al., not 'M. Lewi.' Please correct the author name.","section":"Reference [57]"},{"comment":"The abstract and introduction promise a 'rigorous' analysis of 'theoretical foundations' and 'architectural innovations,' but the body is mostly descriptive. The only equation is the InfoNCE-style contrastive loss in Section III.A. Consider tempering the wording or adding a comparative analysis that goes beyond citing representative architectures.","section":"§I and abstract"},{"comment":"The five categories are not mutually exclusive: EEG-to-vision reconstruction can be considered a multimodal perception task (Section VI.A), and EEG-to-text generation is also a cross-modal task. The paper would benefit from explicit rules for assigning a paper to exactly one category, especially because the taxonomy is the main contribution.","section":"Taxonomy definition"}],"recommendation":"major_revision","confidential_remarks":"The scope inconsistency is the key issue. The paper's own Section II and Section VII undermine the 'exclusively non-EEG' and 'exclude models fine-tuned on EEG' claims, and this is the central axis of the survey. The taxonomy itself may still be salvageable, but the framing needs substantial revision. I would also ask the authors to make the literature selection process explicit; as it stands, the 'comprehensive' claim cannot be checked."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Chris, quick read of arXiv:2508.15716. It's a survey that tries to organize EEG analysis with foundation models into five output-modality categories: native EEG decoding, EEG-text, EEG-vision, EEG-audio, multimodal fusion. That taxonomy is actually useful — it gives a newcomer a reasonable map of a messy, fast-moving field, and the tables are dense with recent work. I also give them credit for including the Jo et al. critique of EEG-to-text reliability and the Banville scaling-law result; they're not just cheerleading.\n\nBut the central claim is shaky. The intro says they include only models pretrained on large-scale non-EEG data and exclude models fine-tuned on EEG. Then Section II (unimodal decoding) includes BENDR, CBraMod, EEGM2, and 'Large transformers are better EEG learners' — all pretrained on EEG. Section VII even describes the standard pipeline as self/supervised pretraining on large-scale unlabeled EEG data. That's not a quibble about definitions; it's an internal contradiction between the stated scope and the actual corpus. The 'cross-domain exclusively non-EEG' qualifier can't stand, and the 'first comprehensive' novelty claim overreaches as written. The fix is straightforward: either drop the EEG-pretrained models from the corpus or broaden the claim. But as is, the paper's main contribution is mislabeled.\n\nMinor issues: there's a duplicated reference for CineBrain ([116] and [117]), and the AudioLDM2 citation in Section V.C points at [100] when [105] is the correct one.\n\nBottom line: the survey is a decent organizing effort, not a deep contribution, and the scope problem is real but correctable. I'd send it to review rather than desk reject — a referee can force the reconciliation and the taxonomy is worth having in the literature. I wouldn't cite it in its current form, but I'd use it as a starting spreadsheet for my own reading.","headline":"A genuinely useful taxonomy of EEG foundation-model work, undermined by an internal scope contradiction that needs fixing before the 'cross-domain, non-EEG' claim can stand.","tokens_in":20165,"tokens_out":2571,"would_cite":false,"duration_ms":26903,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes a five-way, output-modality taxonomy — native EEG decoding, EEG-text, EEG-vision, EEG-audio, and multimodal fusion — as the organizing frame for foundation models applied to EEG analysis.","keywords":["EEG analysis","foundation models","cross-domain transfer","modality-oriented taxonomy","EEG-to-text","EEG-to-vision","EEG-to-audio","multimodal fusion"],"falsifier":"A reader could audit every entry in Tables II through VI against the scope rule; any model pre-trained on EEG data would falsify the claim that the survey covers strictly non-EEG-pretrained models. A second test: take a representative EEG-to-text pipeline and replace the pretrained language model with a randomly initialized version of the same architecture; if downstream decoding does not degrade, then cross-domain pretraining is not doing the work the taxonomy attributes to it.","tokens_in":19324,"feed_emoji":"🧠","tokens_out":6634,"duration_ms":71992,"temperature":0.7,"pith_summary":"This paper tries to establish that the exploding literature on foundation models applied to electroencephalography (EEG) has a simple organizing principle: what the pretrained model produces. It proposes the first modality-oriented taxonomy, sorting work into five output-modality families: native EEG decoding, EEG-text, EEG-vision, EEG-audio, and multimodal fusion. The payoff of getting the classification right is practical: researchers gain a shared map for comparing models, the recurring roles a foundation model can play (feature extractor, alignment tool, generative backbone) become visible, and open problems such as interpretability, cross-subject generalization, and missing benchmarks can be attacked systematically. The scope is deliberately limited to models pre-trained on non-EEG data and applied without EEG fine-tuning, and the survey's usefulness depends on that boundary being held consistently.","feed_headline":"EEG foundation models sorted into five output families","feed_subtitle":"A survey groups brain-decoding research by what the pretrained model emits: native EEG, text, vision, audio, or multimodal fusion.","key_machinery":"The governing device is a function-driven, modality-oriented taxonomy: five output-modality families (native EEG decoding, EEG-text, EEG-vision, EEG-audio, and multimodal fusion) plus a three-role account of foundation models as feature extractor, cross-modal alignment tool, or generative backbone. The taxonomy does the organizing work: it converts a scattered literature into a grid in which each surveyed paper can be located by what it emits, and it makes the shared EEG-encoder / alignment / decoder pipeline visible across otherwise dissimilar papers.","core_discovery":"The paper's central claim is that the heterogeneous and rapidly growing body of work using pretrained foundation models for EEG analysis is best understood by the modality of the model's output, not by the EEG task alone. It sorts the field into five output-modality families: native unimodal EEG decoding; EEG-to-text alignment, generation, and domain-specific understanding; EEG-to-vision retrieval, reconstruction, and video/3D generation; EEG-to-audio decoding, generation, and reconstruction; and multimodal EEG fusion. Within each family, the paper argues, the foundation model plays one of a small set of recurring roles — feature extractor, cross-modal alignment bridge, or generative backbon","pith_inferences":["The paper's stated exclusion of EEG-pretrained models is not consistently applied: the unimodal decoding section includes models trained on EEG itself (for example, a contrastive transformer, a Mamba-based encoder, and several large-transformers), so the actual scope is broader than the advertised cross-domain-only boundary.","If the taxonomy is right, a direct test suggests itself: compare same-architecture models with and without large-scale non-EEG pretraining on each of the five output families; the taxonomy predicts pretrained versions should win reliably in cross-modal families and less clearly in native EEG decoding.","The paper's call for EEG digital twins and X-to-EEG synthesis implies the taxonomy could be extended beyond output modalities to treat synthetic EEG as a new modality, something the paper does not spell out."],"forward_implications":["Researchers can position any EEG foundation-model paper by asking what modality it outputs, making the field's design space explicit.","The recurring three-stage pipeline (self-supervised pretraining, cross-modal alignment, task-specific fine-tuning) becomes a template for designing new methods.","The survey's named open problems — cross-subject generalization, authenticity of generated outputs, and missing standardized benchmarks — become the concrete agenda for the next wave of work.","The inverse direction (generating EEG from text, image, or audio) is identified as a gap, so future work can extend the taxonomy from decoding to synthesis."],"supporting_citations":[{"why":"Supplies the definition of foundation models that sets the survey's scope.","marker":"[6]"},{"why":"The earlier EEG-decoding survey that the authors position their modality-oriented taxonomy against.","marker":"[5]"},{"why":"The EEG-to-text review that defines that research direction for the survey.","marker":"[19]"},{"why":"The contrastive EEG-image alignment work that anchors the EEG-to-vision category.","marker":"[20]"},{"why":"The non-invasive speech-decoding study that anchors the EEG-to-audio category.","marker":"[21]"},{"why":"The critique showing LLMs can generate coherent text from irrelevant EEG, which drives the paper's authenticity discussion.","marker":"[67]"},{"why":"The scaling-law study cited to support data-scalability claims in visual reconstruction.","marker":"[82]"},{"why":"The benchmark reference cited as evidence that standardized evaluation protocols are missing.","marker":"[123]"}],"fun_headline_variants":["EEG foundation models: five output modes, one framework","Five output families organize the EEG model landscape","EEG model survey: output modality is the organizing principle","EEG foundation models grouped by what they output"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The central claim collapses if any surveyed model was actually trained on brain recordings (EEG) or fine-tuned on them, because the taxonomy is explicitly about non-EEG pretraining; the paper's own unimodal section contains examples of EEG-pretrained models, so this assumption is already under strain.","fun_headline_variants_meta":{"raw":{"variants":["EEG foundation models: five output modes, one framework","Five output families organize the EEG model landscape","EEG model survey: output modality is the organizing principle","EEG foundation models grouped by what they output"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000751,"raw_usage":{"total_tokens":3157,"prompt_tokens":697,"completion_tokens":2460,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":2398}},"tokens_in":441,"tokens_out":2460,"duration_ms":20898,"temperature":1.0,"reasoning_tokens":2398,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:42:55.657342+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could audit every entry in Tables II through VI against the scope rule; any model pre-trained on EEG data would falsify the claim that the survey covers strictly non-EEG-pretrained models. A second test: take a representative EEG-to-text pipeline and replace the pretrained language model with a randomly initialized version of the same architecture; if downstream decoding does not degrade, then cross-domain pretraining is not doing the work the taxonomy attributes to it.","supporting_citations":[],"review_version":1}