{"id":"4ba951e4-4140-4f30-a7a8-752d5e5c92b6","arxiv_id":"2412.19999","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A PRISMA-style review of EEG-to-output decoding claims to analyze 1,800 studies but omits the flow diagram, study list, and quantitative synthesis needed to back that claim.","lead":"This paper is a review of studies that use generative AI to turn EEG brain signals into images, videos, and audio. It claims to screen 1,800 papers, narrows the field to 95, and sketches trends, challenges, and a research roadmap.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PRISMA-based analysis of 1,800 studies is unverifiable: no flow diagram, included-study list, or extraction data; the reported funnel and Table 1 are internally inconsistent.","rationale":"The paper's only substantive contribution is the claimed systematic review of 1,800 EEG-to-output studies. If that claim cannot be verified, the entire paper reduces to an essay on three selected papers plus generic discussion, which does not support a systematic-review status. The reader's weakest assumption—that the three case studies are representative of the 95-study pool and that the qualitative findings follow from the screening—is closely related but slightly downstream: the more fundamental issue is that no evidence of the screening itself is present. The manuscript provides exact counts for the funnel and a corrupted Table 1, both of which can be checked. If the authors can supply the search details and the full list, the systematic-review claim might become plausible; without them, the claim is not merely unproven but internally inconsistent. This aligns with the reader's REJECT verdict, so no change is recommended. The proposed concrete test is a direct verification step: reproduce the funnel independently and compare against the reported numbers.","tokens_in":11712,"tokens_out":3688,"duration_ms":36395,"concrete_test":"Independently reconstruct the PRISMA funnel by requesting the authors' search protocol and included-study list: run the same Boolean query across PubMed, IEEE Xplore, Scopus, JSTOR, and Google Scholar for 2015–2024; verify the initial pool of 1,800, the deduplication count, and that the 95 final studies meet the stated inclusion criteria (generative model, public dataset, quantitative metrics). At the same time, recompute Table 1 from the per-database retrieval counts. If the initial pool, the funnel counts, or the listed studies cannot be reproduced, the systematic-review claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that the authors followed PRISMA and analyzed 1,800 studies to identify trends—cannot be substantiated from the manuscript itself. Section 2 reports exact funnel counts (1,800 → approximately 295 → 150 → 95) but omits every artifact that would make such a funnel verifiable: the PRISMA flow diagram, per-database search results, the Boolean query string, deduplication counts, exclusion reasons, and the complete list of the 95 included studies. No data-extraction form, no aggregated metric tables, and no quantitative trend analysis appear anywhere in the Findings or Discussion sections. Moreover, the only quantitative summary offered, Table 1, is corrupted: entries such as \"Google Scholar 1765 39 201 3920 .10 22 .21 57 115 .26 13 .86\" are unparseable, and the numbers do not reconcile with the reported funnel (Google Scholar alone would account for 1765 of the 1800 initial pool). The three case studies (EEG2Image, Chen's musicality evaluation, EEG2Video) are descriptions of individual papers, not syntheses of 95 studies. Section 4's strengths-and-limitations discussion is generic material that does not cite the 95 studies and could apply to any deep-learning field. Thus the load-bearing assumption—that a systematic screening was actually performed and that the findings derive from it—is unsupported and internally contradicted by the manuscript's own reporting.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript claims to be a systematic review following PRISMA guidelines, analyzing 1,800 studies to survey EEG-to-output decoding across images, video, and audio. Section 2.1 describes a four-stage screening funnel (1,800 → approximately 295 → 150 → 95) and includes Table 1 as a summary of academic metrics. Section 3 presents three case studies: EEG2Image, Chen's musicality evaluation, and EEG2Video, with equations quoted from the original papers. Section 4 discusses the strengths and limitations of GANs, VAEs, and Transformers, dataset scarcity, cross-subject variability, ethics, and future directions, and Section 5 concludes with a proposed roadmap. The abstract's central claim is that PRISMA-based analysis of 1,800 studies identifies key trends, challenges, and opportunities.","tokens_in":11958,"tokens_out":6295,"duration_ms":61130,"significance":"If the systematic screening were substantiated, this paper would provide a useful, reproducible map of EEG-to-output research and a reliable trend analysis. The manuscript has some strengths: the three case-study descriptions are readable, the equations are recognizable from the cited sources, and the paper explicitly names the PRISMA 2020 guideline. I found no circular-reasoning issue, because the equations are quoted from cited papers and are not used as evidence for the review's own conclusions. However, the central contribution is currently unverifiable: the claimed 1,800-study PRISMA analysis is not supported by the required reporting artifacts, and the 'Findings' section consists of single-paper case studies rather than a synthesis of the 95 allegedly included studies. For these reasons, the present significance is limited to a narrative overview of a few methods.","major_comments":[{"comment":"The central claim that the authors 'analyze 1800 studies' using PRISMA cannot be verified from the manuscript. PRISMA 2020 requires a flow diagram, database-specific search strings, per-database yields, deduplication counts, exclusion reasons, and a list of included studies. None of these are supplied. The text reports only the final funnel numbers (1,800, approximately 295, 150, 95) with a conceptual query notation, leaving no way to check which studies were screened, why they were excluded, or whether the 95-study set exists. Because this claim is the paper's headline contribution, the omission is load-bearing.","section":"Abstract and Section 2.1"},{"comment":"Table 1 is corrupted and cannot be used as evidence. For example, the Google Scholar row reads '1765 39 201 3920 .10 22 .21 57 115 .26 13 .86', which cannot be parsed under the stated column headers (Papers, Citations, Cites/Year, Cites/Paper, h-index, g-index, hI-index). Although the Papers column sums to 1,800 across the five sources, the remaining fields are unparseable, and the table is never discussed in the text. A corrupted table cannot substantiate the claimed systematic screening process.","section":"Table 1"},{"comment":"Section 3 does not report findings from the 95 included studies. It presents three single-paper case studies—EEG2Image, Chen's musicality evaluation, and EEG2Video—with equations and results quoted from those individual papers. There is no data-extraction table, no aggregate distribution of methods, datasets, or metrics across the 95 studies, and no quantitative trend analysis. Consequently, the abstract's claim that the review 'identifies key trends, challenges, and opportunities' is not demonstrably derived from the systematic screening described in Section 2.","section":"Section 3, Findings"},{"comment":"The Discussion is generic and disconnected from the alleged 95-study pool. Statements such as GANs being 'notoriously difficult to train' and Transformers requiring 'large datasets for optimal performance' are textbook observations supported by general references rather than by an analysis of the screened EEG-to-output literature. No synthesis, evidence table, or citation of the 95 studies connects the discussion to the claimed systematic review, so Section 4 does not deliver on the promise of a field-level trend analysis.","section":"Section 4, Discussion"},{"comment":"The statement 'We followed a similar methodology as Jain et al. [21]' is problematic because reference [21] is arXiv:2312.10057, 'Generative AI in writing research papers: A new type of algorithmic bias and uncertainty in scholarly work.' That paper is not a systematic-review methodology paper, so it cannot serve as a methodological precedent for the PRISMA process described here. This further weakens the credibility of the claimed methodology.","section":"Section 2.1, Reference [21]"}],"minor_comments":[{"comment":"The final sentence mixes grammatical forms: 'to improve decoding accuracy and broadening real-world applications' should be 'to improve decoding accuracy and broaden real-world applications.'","section":"Abstract"},{"comment":"The funnel counts are reported with inconsistent precision: 'exactly 1800 studies' is stated as an exact number, while the next step is 'approximately 295 papers.' The source of the exactness of 1,800 should be explained.","section":"Section 2.1"},{"comment":"Table 1 is not referenced or discussed anywhere in the text; if it is retained, a caption and a clear explanation of its construction are needed.","section":"Table 1"},{"comment":"Several reference entries contain line-break artifacts that will break URLs in the compiled PDF, including references 24, 25, and 42; these should be cleaned before any resubmission.","section":"References"},{"comment":"The bilinear model equation f(X_s^m) = w_1^T X_s^m w_2 + b is presented without specifying the rank of the bilinear term or how the projection vectors are constrained; a brief clarifying sentence would improve the summary.","section":"Section 3.2"}],"recommendation":"reject","confidential_remarks":"The combination of missing PRISMA artifacts, the corrupted Table 1, and the citation of an arXiv paper about generative AI in academic writing as a 'similar methodology' raises concerns about whether the claimed 1,800-study screening was actually performed. The editor may wish to request the search logs, PRISMA flow diagram, and the complete list of 95 included studies if any revision is entertained. In its current form, the manuscript does not meet the standards of a systematic review and the central contribution is unverifiable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Plainly: the central claim of this review—that the authors followed PRISMA and analyzed 1,800 studies down to 95—is not supported by the manuscript. There is no PRISMA flow diagram, no list of the 95 included studies, no extraction forms, and no aggregated quantitative findings. Table 1 is garbled (the per-source counts roughly sum to 1,800, but the remaining columns are unreadable), so we cannot verify even the raw search numbers.\n\nWhat the paper does reasonably well is summarize three specific methods. The EEG2Image, Chen musicality, and EEG2Video case studies are described with correct equations and faithful attributions; a newcomer could learn the basic ideas from Sections 3.1-3.3. The reference list is broad enough to serve as a starting point.\n\nThe soft spots are load-bearing. Section 3 is three paper summaries, not a synthesis of 95 studies. Section 4 is generic commentary on GANs, VAEs, transformers, and dataset challenges that could be written without any screening. The methodology cites a paper about AI-generated research papers as its template, which does not substitute for a PRISMA protocol. The abstract's '1,800 studies' figure is effectively unverifiable.\n\nThis might work as a short non-systematic introduction to three representative papers, but as a 'comprehensive review' it is not credible. I would desk reject; the flaws are not fixable with minor revision, since the missing artifacts require the actual screening to be documented.","headline":"A review whose PRISMA claim of screening 1,800 studies is unverifiable from the manuscript; the case-study descriptions are accurate, but the systematic-review scaffolding does not hold.","tokens_in":12483,"tokens_out":3332,"would_cite":false,"duration_ms":33390,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic screen of 1,800 EEG studies positions generative models as the core of brain-signal decoding.","keywords":["EEG decoding","image reconstruction","video synthesis","audio decoding","generative models","systematic review","brain-computer interfaces","cross-subject generalization"],"falsifier":"Reproduce the search and screening using the paper's stated inclusion criteria and see whether the resulting 95-study corpus would place GANs and image reconstruction as the dominant trend and would include the three case studies among the most influential papers; if the corpus differs materially, or if the case studies fall outside the top papers by the paper's own citation and venue criteria, the review's central characterization of the field would be called into question.","tokens_in":91,"feed_emoji":"🧠","tokens_out":8478,"duration_ms":137615,"temperature":0.7,"pith_summary":"This paper is a systematic review of efforts to reconstruct images, videos, and audio directly from EEG brain recordings. The authors screen 1,800 records down to 95 core studies and present three representative systems: EEG2Image for image generation, EEG2Video for dynamic visual decoding, and a bilinear-model framework for scoring the musicality of machine-composed music. The review's central claim is that generative models—GANs, VAEs, and transformers—are the state of the art in EEG-to-output decoding, and that the field's progress is limited mainly by scarce standardized datasets, cross-subject variability, and inconsistent evaluation metrics. If the review's map is correct, it gives researchers a single entry point to the field and points investment toward shared benchmarks and multimodal integration.","feed_headline":"1,800 studies screened: brain signals to images, video, and audio","feed_subtitle":"Review finds GANs, VAEs, and transformers drive EEG decoding, but dataset scarcity blocks real-world use.","key_machinery":"The argument rests on three case studies that stand in for the field. EEG2Image couples a contrastive triplet-loss feature extractor with a conditional GAN modified by mode-seeking regularization to produce $128 \\times 128$ reconstructions. EEG2Video uses a Seq2Seq architecture for temporal alignment, a text-embedding semantic predictor, and a dynamic-aware noise-adding (DANA) process to steer a diffusion model toward coherent video. The musicality framework applies a bilinear model to EEG feature matrices, with Gamma-band and DC components as the discriminative signals, to rank compositions. Around these, the review's screening pipeline—identification, screening, eligibility, inclusion, with citation and venue filters—is the mechanism that converts a search of 1,800 records into a 95-study corpus.","core_discovery":"On its own terms, the paper establishes that EEG-to-output decoding has become a distinct research area with measurable achievements: images can be reconstructed at $128 \\times 128$ resolution using a conditional GAN with triplet-loss features and mode-seeking regularization; dynamic video can be decoded with a Seq2Seq model and a dynamic-aware noise-adding process, yielding a reported SSIM of 0.256 and 15.9% semantic accuracy on a 40-class task; and auditory musicality can be ranked by a bilinear model drawing on Gamma-band EEG. The review asserts that applying a systematic screening protocol to 1,800 initial records yields 95 core studies, and that these studies cluster around generative architectures while sharing the weaknesses of small subject pools, limited cross-subject transfer, and no universal benchmark. The paper's proposed roadmap treats standardized open datasets, transfer and federated learning, multimodal recording, and explainable AI as the necessary next steps for practical brain-computer interfaces.","pith_inferences":["The review's screening yields 95 studies, but the paper presents no quantitative trend data connecting those studies to its conclusions; the trends it reports are illustrative and would need a data-extraction pass to be verified.","Because the three case studies are the only substantive examples, the review's characterization of the field is only as reliable as the representativeness of those three systems; a different selection could change the apparent balance among image, video, and audio lines of work.","A testable extension of the musicality result is to apply the same bilinear EEG-scoring method to other generative outputs (e.g., synthesized speech or deepfake video) to see whether Gamma-band responses track perceptual quality across modalities.","If EEG2Video's reported metrics are reproducible on a held-out set, its dynamic-aware noise-adding mechanism could be transferred to other modality-decoding diffusion models; if not, the case for EEG-based video decoding weakens."],"forward_implications":["If the review's map is correct, the immediate bottleneck is data, not model architecture: standardized public datasets and shared benchmarks would likely raise reconstruction quality and cross-subject transfer more than further architecture tweaks.","Hybrid approaches that combine VAEs with GANs, and transformer-based temporal models, are the most promising near-term designs for closing the image and video quality gap.","Multimodal integration (e.g., EEG with fNIRS or MEG) and federated learning could mitigate the privacy and data-scarcity problems the review identifies, making real-world brain-computer interfaces more feasible.","The reported EEG2Video numbers (SSIM 0.256, 15.9% semantic accuracy) suggest dynamic video decoding is still early-stage; the roadmap implies that reaching usable video BCIs will require new datasets with naturalistic stimuli and standardized evaluation."],"supporting_citations":[{"why":"Supplies the systematic-review protocol that defines the paper's method and four screening stages.","marker":"[20]"},{"why":"Provides the methodology template the review says it follows for analyzing the literature.","marker":"[21]"},{"why":"The EEG2Image case study, the main evidence for GAN-based image reconstruction with triplet loss and mode-seeking regularization.","marker":"[24]"},{"why":"The EEG2Video case study, the main evidence for Seq2Seq and DANA-based video synthesis with reported SSIM and semantic accuracy.","marker":"[42]"},{"why":"The bilinear-model case study for musicality evaluation using Gamma-band EEG features.","marker":"[35]"},{"why":"A public EEG-to-image dataset, supporting the eligibility criterion that studies must use public datasets.","marker":"[25]"},{"why":"Defines SSIM, the evaluation metric used for EEG2Video and generally for reconstruction quality.","marker":"[47]"}],"fun_headline_variants":["1,800 EEG studies reviewed: brain waves to images, video, audio","EEG decoding review: images and sound, but dataset gaps remain","GANs, VAEs, transformers lead EEG-to-media decoding, review finds","95 core studies on EEG-to-output: progress meets data scarcity","From brain waves to media: 1,800-study review highlights limits"],"cache_read_input_tokens":14592,"weakest_assumption_plain":"The review's field-level findings rest on the assumption that the three hand-chosen case studies represent the 95 screened studies, and that the qualitative discussion follows from the screening; the paper does not show the data extraction that would connect them.","fun_headline_variants_meta":{"raw":{"variants":["1,800 EEG studies reviewed: brain waves to images, video, audio","EEG decoding review: images and sound, but dataset gaps remain","GANs, VAEs, transformers lead EEG-to-media decoding, review finds","95 core studies on EEG-to-output: progress meets data scarcity","From brain waves to media: 1,800-study review highlights limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000775,"raw_usage":{"total_tokens":3402,"prompt_tokens":890,"completion_tokens":2512,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":2417}},"tokens_in":506,"tokens_out":2512,"duration_ms":18711,"temperature":1.0,"reasoning_tokens":2417,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:41:43.307548+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Reproduce the search and screening using the paper's stated inclusion criteria and see whether the resulting 95-study corpus would place GANs and image reconstruction as the dominant trend and would include the three case studies among the most influential papers; if the corpus differs materially, or if the case studies fall outside the top papers by the paper's own citation and venue criteria, the review's central characterization of the field would be called into question.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the systematic-review protocol that defines the paper's method and four screening stages."},{"cited_title":"Generative AI in Writing Research Papers: A New Type of Algorithmic Bias and Uncertainty in Scholarly Work","cited_arxiv_id":"2312.10057","evidence_quote":"Provides the methodology template the review says it follows for analyzing the literature."},{"cited_title":"EEG2IMAGE: Image Reconstruction from EEG Brain Signals","cited_arxiv_id":"2302.10121","evidence_quote":"The EEG2Image case study, the main evidence for GAN-based image reconstruction with triplet loss and mode-seeking regularization."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The EEG2Video case study, the main evidence for Seq2Seq and DANA-based video synthesis with reported SSIM and semantic accuracy."},{"cited_title":"In: Proceedi ngs of the 25th ACM international conference on Multimedia","cited_arxiv_id":null,"evidence_quote":"The bilinear-model case study for musicality evaluation using Gamma-band EEG features."}],"review_version":1}