{"id":"399c00d7-14ec-45e1-a192-b57ec7155f28","arxiv_id":"2506.06353","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A taxonomy and review of studies applying large language models to EEG signals, organized into four domains and three adaptation strategies.","lead":"This survey maps recent research that combines large language models with brain-wave recordings, sorting the work into four categories: brain-signal foundation models, translating brain waves into text, generating images or 3D objects from brain waves, and clinical applications. It is a useful orientation for newcomers, but it contributes no new measurements, code, or experimental results.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's central taxonomy rests on an unstated, possibly non-representative study selection; without a stated review protocol, the four-domain structure and Fig. 3 percentages are not verifiable.","rationale":"The reader's weakest assumption identifies the same core concern: the survey's representativeness and completeness are not supported by any stated methodology. The paper explicitly claims to be 'systematic' and 'comprehensive' in the abstract and conclusion, yet Section 3 and Section 5 provide no search protocol, no inclusion/exclusion criteria, no number of screened studies, and no explanation of how the four-domain taxonomy was derived. The quantitative distribution in Fig. 3 is presented as percentages without a denominator, making it impossible to assess the relative emphasis of research areas. Since the central claim is that this taxonomy is a foundational organizational resource, the selection of studies is the load-bearing assumption. A biased or incomplete selection would not merely be an omission; it would change the taxonomy's structure and the paper's contribution. The duplicate EEGPT naming ([62] vs. [68]) is a genuine and concrete flaw that illustrates careless handling of the literature, but it is not the central threat to the main claim. The concern is substantive enough to justify the reader's CONDITIONAL verdict. It does not, however, warrant rejection, because the paper still provides a usable orientation to the field, and the taxonomy may survive even with modest additions; the issue is that this cannot be verified from the manuscript. Thus the verdict remains CONDITIONAL, pending the stated methodology and corrected ambiguity. The proposed concrete test — an independent systematic search — directly settles whether the selection is representative.","tokens_in":20961,"tokens_out":3215,"duration_ms":30691,"concrete_test":"Independently re-run a systematic literature search (e.g., Scopus, Web of Science, arXiv, PubMed, IEEE Xplore) for 2020-2025 using queries combining EEG/electroencephalography with LLM/GPT/BERT/transformer/large language model, apply explicit inclusion criteria (studies that use an LLM or LLM-style transformer for EEG analysis or generation), and compare the resulting study set against the papers cited in Table 2 and Fig. 8. If the search reveals a substantial cluster of missing studies (e.g., more than 5 additional papers on EEG-to-text or emotion recognition) or if more than 20% of the retrieved studies fall outside the four proposed domains, then the taxonomy's representativeness claim fails and the distribution in Fig. 3 would need revision. Also verify the counts implied by Fig.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that it provides a 'systematic review' and 'comprehensive' taxonomy of LLM-EEG research. For a survey, the load-bearing condition is that the reviewed study set is representative and sufficiently complete to support the four-domain organization and the quantitative distribution in Fig. 3. The manuscript never states a search strategy, inclusion/exclusion criteria, the databases queried, or the number of screened papers. Section 3 introduces the taxonomy directly, and Section 5 describes selected studies without explaining why these and not others were included. Figure 3 shows percentages (25%, 31%, 16%, 28%) with no denominator or study count, so a reader cannot tell whether the distribution reflects the field or the authors' informal sampling. If, for example, a substantial body of work on EEG-based speech decoding or emotion recognition with LLMs was omitted, the four-domain taxonomy would misrepresent the field's actual emphasis, undermining the paper's usefulness as a 'foundational resource.' The duplicate naming of two distinct models as 'EEGPT' (Ref. [62] masked spatio-temporal modeling; Ref. [68] autoregressive 1.1B model) is a concrete ambiguity that further reduces confidence in the survey's carefulness, though it is secondary to the selection-bias concern. Because the systematic-review claim and the numerical distribution both depend on an unstated selection protocol, the central contribution is not currently falsifiable from the manuscript alone.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a survey of recent work combining large language models (LLMs) with electroencephalography (EEG) analysis. The authors propose a four-domain taxonomy — LLM-inspired foundation models, EEG-to-language decoding, cross-modal generation, and clinical applications and dataset tools — and illustrate it with tables and figures summarizing selected studies, model types, tasks, and datasets. The paper also reviews adaptation strategies (fine-tuning, zero-shot, and few-shot learning) and outlines future directions. Its stated contribution is to serve as a systematic and comprehensive foundational resource for this emerging interdisciplinary field.","tokens_in":21228,"tokens_out":3588,"duration_ms":31553,"significance":"If the taxonomy and the quantitative distribution of studies are representative of the field, this survey provides a genuinely useful organizational framework for researchers at the intersection of LLMs and EEG. The paper's strengths include the breadth of the covered topics, the detailed tables (Tables 2 and 3) that summarize models, tasks, and datasets, and the clear diagrams of the taxonomy and adaptation workflows. As a survey, it makes no new technical claims or fitted predictions, so the standard circularity concerns for derivations do not apply. However, its value as a 'foundational resource' depends on the transparency and representativeness of the study selection, which the current manuscript does not establish.","major_comments":[{"comment":"The abstract and Section 3 describe the paper as a 'systematic review' and present a quantitative distribution of studies (25%, 31%, 16%, 28%), but the manuscript never states the search strategy, the databases queried, the inclusion/exclusion criteria, or the number of papers screened. Without a stated methodology and a denominator for Figure 3, the percentages are not verifiable and the four-domain taxonomy cannot be assessed for representativeness. Please add a methodology subsection detailing the literature search, screening process, and exact study counts, and report counts in addition to percentages in Figure 3.","section":"Section 3 and Figure 3"},{"comment":"Two distinct models named 'EEGPT' are discussed: EEGPT [62], which uses electrode-wise masked modeling, and EEGPT [68], an autoregressive 1.1B-parameter model. The text and Table 2 do not disambiguate these two works, which is highly confusing for readers and weakens confidence in the survey's carefulness. Please rename or explicitly distinguish the two models (e.g., as EEGPT-MAE and EEGPT-AR) consistently in the text, Table 2, Figure 8, and the reference list.","section":"Section 5.1 and Table 2"},{"comment":"The taxonomy's boundaries are internally inconsistent: Section 3 defines 'Cross-Modal EEG Generation' as translation to images, text, or 3D objects, while 'EEG-to-Language Decoding' is a separate category that also covers text generation. In Figure 8, AdaCT, an EEG-to-text method, is placed under Cross-Modal Generation rather than under EEG-to-Language Decoding, whereas other EEG-to-text systems appear in the latter category. The authors should provide explicit operational criteria for assigning studies to the four domains, or revise the category definitions to remove this overlap.","section":"Section 3 taxonomy vs. Figure 8 and Table 2"}],"minor_comments":[{"comment":"The text cites Vaswani et al. (2017) as [31], but reference [31] is listed as Parmar et al., 'Image transformer'; this appears to be a citation error, since the same work is also cited as [19].","section":"Section 2.2, reference [31]"},{"comment":"The task description 'EET-to-text and reading analysis' contains a typo and should read 'EEG-to-text and reading analysis'.","section":"Table 3, ZuCo 1.0 row"},{"comment":"The sentence 'crucial tasks given the field’s challenges with data heterogeneity' is missing a period or em-dash before 'crucial' and would benefit from a grammatical fix to clarify that the tools address crucial tasks.","section":"Section 5.4.5"},{"comment":"The table lists only 'ClinicalBERT' for reference [63], but Section 5.4.5 states that Meditron-7B and BioMistral are also used in that work; please make the table consistent with the text.","section":"Table 2, reference [63]"}],"recommendation":"major_revision","confidential_remarks":"The central concern is that the survey's 'systematic review' claim is not supported by a stated methodology, making the taxonomy and Figure 3 unverifiable. This is fixable within the scope of the paper by adding a literature-search protocol and clearer taxonomy boundary criteria. The duplicate 'EEGPT' naming is a concrete symptom of insufficient cross-checking and should be resolved. I would not recommend rejection, as the paper's organizational framework is potentially valuable, but the revision needs to address the methodological and consistency issues described in the major comments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful survey for someone entering the LLM-EEG area. The four-domain taxonomy (foundation models, EEG-to-language decoding, cross-modal generation, clinical tools) is a sensible way to organize a very recent and scattered literature, and Tables 2 and 3 do real work: they map the models, datasets, and tasks so a newcomer can see who did what without reading twenty papers. I got value from the dataset table in particular.\n\nThe soft spots are real but bounded. The abstract calls it a 'systematic review,' but there is no methodology: no search strategy, databases, inclusion/exclusion criteria, or screening counts. Figure 3 shows percentages (25/31/16/28) with no study count or denominator, so the distribution is not auditable. That matters, because the paper's own integrative claim hangs on the reviewed set being representative. If the selection is biased, the taxonomy could still be right, but the paper currently gives the reader no way to tell. The stress-test note is on target, and it is not a manufactured flaw; the text itself asserts the systematic/comprehensive framing.\n\nOne concrete error reinforces the carelessness concern: two distinct models in the same reference list are both called EEGPT—[62] is a masked spatio-temporal model from NeurIPS 2024, [68] is an autoregressive 1.1B model from a different group. Section 5.1 presents them in different subsections, and Table 2 lists both, but neither the table nor the text flags the name collision. That is a minor fix but it is confusing to the exact audience the survey is supposed to help.\n\nWhat is not wrong: the summaries of individual papers are broadly consistent with the cited work, I did not find obvious mis-citations, and the adaptation-strategy section (fine-tuning, zero-shot, few-shot) is a legitimate framing device. There are no fitted parameters or new claims to falsify, so circularity concerns do not apply.\n\nWho is this for? A grad student or early postdoc getting into LLM-EEG, as a starting map. It is not a critical scholarly contribution—it is a synthesis. But the field is young enough that a decent synthesis is worth having.\n\nRecommendation: a serious editor should send it to review, with a clear request to add a methodology paragraph (search, selection, counts) and fix the EEGPT naming. If the authors do that, it becomes a solid orientation resource.","headline":"A useful orientation survey of the LLM-EEG area whose 'systematic' and quantitative claims outrun the stated methodology; the taxonomy itself is reasonable and the tables are handy.","tokens_in":21736,"tokens_out":2454,"would_cite":false,"duration_ms":21175,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey organizes EEG-LLM research into four domains: foundation models, brain-to-language decoding, cross-modal generation, and clinical tools.","keywords":["EEG","large language models","brain-computer interfaces","EEG-to-language decoding","foundation models","cross-modal generation","clinical applications"],"falsifier":"A systematic replication that defines a search query, inclusion criteria, and screening count, then classifies every eligible paper, would settle the claim: if a substantial fifth category emerges or the four-domain distribution in Figure 3 shifts materially, the taxonomy misrepresents the field.","tokens_in":20773,"feed_emoji":"🧠","tokens_out":8816,"duration_ms":70788,"temperature":0.7,"pith_summary":"The paper sets out to show that the growing convergence between large language models and EEG research has produced a body of work that can be organized into four domains: LLM-inspired foundation models for EEG representation learning, EEG-to-language decoding, cross-modal generation (images and 3D objects), and clinical applications plus dataset tools. It argues that transformer-based architectures, adapted through fine-tuning, few-shot, and zero-shot learning, let EEG models perform complex tasks such as generating natural language from brain signals, semantic interpretation, and diagnostic assistance. A reader should care because the field currently lacks a shared structure, and this taxonomy gives researchers a common vocabulary for placing new work.","feed_headline":"Four domains organize EEG-LLM research, survey shows","feed_subtitle":"The four-domain map shows how language models decode, generate, and diagnose from brain signals.","key_machinery":"The central object is the taxonomy itself: four categories with a hierarchical breakdown of modeling strategies. The machinery underneath is the transformer/LLM architecture family, specifically masked (BERT-style) and autoregressive (GPT-style) pretraining, plus three adaptation strategies—fine-tuning (including prefix tuning and adapters), zero-shot prompting, and few-shot in-context learning—that let text-trained models ingest EEG signals and emit text, images, or diagnostic labels.","core_discovery":"The paper's central claim is that the convergence of LLMs and EEG is not a scattering of isolated experiments but a field that can be organized into four functional domains: foundation models that learn transferable EEG representations through masked or autoregressive pretraining; EEG-to-language decoding that generates natural language from brain signals through decoders, semantic alignment, or instruction tuning; cross-modal generation that turns brain activity into images or 3D objects; and clinical and dataset work covering emotion recognition, mental health diagnosis, motor imagery classification, reading analysis, and data tools. The survey presents this taxonomy as a foundational resource, with tables and diagrams showing which LLMs are used, how they are adapted, and which datasets support each line of work. If the map is right, future research can be positioned, compared, and extended within a common structure.","pith_inferences":["The Figure 3 distribution implies the four domains are not equally developed, with cross-modal generation at 16 percent the thinnest cluster; that area may absorb future work as diffusion and 3D generative models improve.","The paper's own framing suggests a testable next step: if EEG foundation models mature into standardized pretrained backbones, researchers should rely less on training task-specific EEG models from scratch, mirroring the shift that pretrained LLMs caused in NLP.","Because the survey reports no search strategy or inclusion criteria, a reader cannot yet tell whether the four-domain structure would survive a formal systematic review; that is the natural next test of the taxonomy.","An implication the authors leave implicit is that all four domains inherit EEG's noisy-signal and inter-subject variability problems, so progress in preprocessing and artifact handling likely gates progress in each category."],"forward_implications":["New EEG-LLM work can be positioned within one of four domains, giving the field a shared reference structure for comparing methods.","EEG-to-language decoding is the largest surveyed cluster at 31 percent of studies (Figure 3), so brain-to-text communication is the near-term application most likely to mature.","Fine-tuning and few-shot or zero-shot adaptation mean existing text-trained LLMs can be repurposed for EEG tasks without large labeled EEG datasets.","Cross-modal generation indicates that brain activity can guide image and 3D synthesis, pointing toward richer brain-computer interfaces.","Clinical tools such as lightweight emotion copilots and dataset harmonization frameworks suggest LLMs can support real-time feedback and data standardization, not just offline analysis."],"supporting_citations":[{"why":"Supplies the instruction-tuned LLM pipeline used for open-vocabulary EEG-to-text decoding.","marker":"[9]"},{"why":"Provides the LLM semantic-scaffold approach behind the EEG-to-image generation category.","marker":"[10]"},{"why":"Provides the EEG-to-3D object reconstruction method that defines the 3D branch of cross-modal generation.","marker":"[11]"},{"why":"Gives the multimodal few-shot mental health classifier that anchors the clinical applications category.","marker":"[12]"},{"why":"Supplies the prefix-tuning method that bridges EEG encoders to GPT-style text decoders.","marker":"[39]"},{"why":"Supports zero-shot contrastive alignment and retrieval via a shared EEG-text embedding space.","marker":"[42]"},{"why":"Provides the discrete contrastive alignment framework for EEG-to-text translation.","marker":"[50]"},{"why":"Supplies the dataset harmonization tool that grounds the dataset-management branch.","marker":"[53]"},{"why":"Supplies the cross-dataset foundation model representative of the masked-pretraining branch.","marker":"[60]"},{"why":"Supplies the transformer and self-attention architecture that all adapted models build on.","marker":"[31]"}],"fun_headline_variants":["LLM-EEG survey maps four research domains","Survey: four domains link LLMs and EEG","EEG-LLM taxonomy: four domains, one map","Study: LLMs decode EEG in four domains","LLM-EEG convergence framed in four domains"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The taxonomy's four-domain structure assumes the studies selected for review are representative and complete enough to map the field, but the paper reports no search strategy, inclusion criteria, or screening count, so a biased sample would distort the map.","fun_headline_variants_meta":{"raw":{"variants":["LLM-EEG survey maps four research domains","Survey: four domains link LLMs and EEG","EEG-LLM taxonomy: four domains, one map","Study: LLMs decode EEG in four domains","LLM-EEG convergence framed in four domains"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000146,"raw_usage":{"total_tokens":1144,"prompt_tokens":868,"completion_tokens":276,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":484,"completion_tokens_details":{"reasoning_tokens":200}},"tokens_in":484,"tokens_out":276,"duration_ms":2787,"temperature":1.0,"reasoning_tokens":200,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:28:05.568319+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic replication that defines a search query, inclusion criteria, and screening count, then classifies every eligible paper, would settle the claim: if a substantial fifth category emerges or the four-domain distribution in Figure 3 shifts materially, the taxonomy misrepresents the field.","supporting_citations":[{"cited_title":"Thought2text: Text generation from eeg signal using large language models (llms),","cited_arxiv_id":null,"evidence_quote":"Supplies the instruction-tuned LLM pipeline used for open-vocabulary EEG-to-text decoding."}],"review_version":1}