{"id":"d0a78ffb-2ff7-43fb-a67a-6ced36a6ab70","arxiv_id":"2508.18571","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Alljoined-1.6M is a 1.6-million-trial EEG-image dataset recorded on a 32-channel consumer headset, showing that semantic decoding and EEG-to-image reconstruction work on affordable hardware at scale.","lead":"A new open dataset records 1.67 million EEG responses to images from 20 people using a ~$2,200 consumer headset, far cheaper than the research-grade systems used for the leading benchmark. The authors show that semantic category information and even image reconstructions can be decoded from this noisier data, and that performance keeps improving with more trials.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reconstruction and scaling claims rest entirely on ENIGMA, an anonymous in-review method referenced to a nonexistent Appendix B, making the paper's central quantitative results unauditable.","rationale":"The reader's verdict is CONDITIONAL, and I agree with that verdict, but for a different primary reason. The reader's weakest assumption concerns wireless trigger jitter. That is a real risk, and the paper's only defense is an unquantified claim of 'millisecond-accurate triggers' plus a 0–0.6% trial-exclusion rate (Section 3). However, the internal evidence partly mitigates this risk: the averaged ERP shows clear P1/N200 peaks and the cluster permutation tests yield 16/21 significant category contrasts, which would be smeared away by large jitter. A jitter problem would degrade, but not necessarily destroy, the central decoding claim; models could even learn the jitter distribution. By contrast, the absence of ENIGMA is a hard blocker: the manuscript points to an Appendix B that does not exist, so Table 1 and Figure 7A cannot be reproduced or even understood. The reconstruction and scaling claims are not peripheral; they are highlighted in the abstract ('effective EEG-to-Image reconstruction', 'log-linear decoding performance with increasing data volume'). The paper itself flags the missing support ('see Appendix B'), so under the review rule this must be weighed explicitly. The dataset release itself is a genuine contribution, and the basic LDA decoding and ERP analyses provide credible evidence that category information is present in consumer-grade EEG, so REJECT would be too harsh. UNVERDICTED might be defensible for the reconstruction claims specifically, but the dataset's value and the simpler decoding benchmarks support a CONDITIONAL verdict pending release or removal of the ENIGMA-dependent claims. The selection-filtering issue (Appendix A.3: 10/48 excluded for quality or cooperativeness, including four removed as 'difficult or unpleasant to work with') further limits the real-world deployability claim, but is secondary to the reproducibility blocker.","tokens_in":18057,"tokens_out":8285,"duration_ms":78991,"concrete_test":"Inspect the released GitHub and HuggingFace repositories for an implementation or trained weights of ENIGMA; if absent, re-run the Table 1 evaluation (or at least the human 2AFC and CLIP/Incep 2WC metrics) on Alljoined-1.6M using only ATM-S and Perceptogram, and re-plot the scaling curve using only those public methods. If the public methods still yield above-chance identification accuracy (2AFC > 60%) and monotone improvement with log-sample count, the dataset's core utility survives without ENIGMA; if they do not, the paper's reconstruction and scaling conclusions stand only once ENIGMA is released and documented.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 ('EEG-to-Image Reconstruction') states: 'we took all publicly available EEG-to-Image reconstruction methods (ENIGMA [2], ATM-S [3], and Perceptogram [4]) and reproduced their methods on our dataset.' Reference [2] is 'Anonymous. Enigma... In Review, see Appendix B., 2025.' No Appendix B exists in this preprint (appendices run A.1–A.10), and the ENIGMA architecture is not otherwise described. Every quantitative benchmark in Table 1 (Alljoined rows) and the scaling curves in Figure 7A are produced with ENIGMA; the log-linear scaling claim in Section 4 ('Scaling Analysis') is specifically ENIGMA's learning curve. Thus the two headline assertions — 'effective EEG-to-Image reconstruction' and 'log-linear decoding performance with increasing data volume' — cannot be audited or reproduced by any reader. This is a missing reference / omitted proof admitted by the manuscript itself ('see Appendix B'). If ENIGMA were removed, ATM-S and Perceptogram (Table 1) do show above-chance 2AFC identification (~60–65%) on Alljoined-1.6M, but those methods are not used for the scaling analysis and their reconstructions are markedly below the paper's 'comparable to THINGS-EEG2' language. The central claim therefore depends on an unverifiable model.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Alljoined-1.6M, an EEG dataset of over 1.6 million visual stimulus trials recorded from 20 participants with a 32-channel Emotiv Flex 2 consumer-grade headset, using the THINGS image set and a rapid serial visual presentation paradigm. The authors report ERP analyses with cluster-based permutation tests, time-resolved pair-wise LDA decoding across seven meta-categories, EEG-to-image reconstruction benchmarks using ENIGMA, ATM-S, and Perceptogram, saliency maps, a scaling analysis of reconstruction performance against training set size, and a channel-count ablation. The central claims are that consumer-grade hardware can support high-level semantic decoding and effective EEG-to-image reconstruction, and that decoding performance scales log-linearly with data volume without saturation. The dataset and benchmark code are planned for public release.","tokens_in":18210,"tokens_out":5421,"duration_ms":50098,"significance":"A public dataset of this scale recorded on affordable hardware would be a valuable community resource for studying the trade-off between hardware cost, signal fidelity, and data volume in visual EEG decoding. The basic decoding evidence is solid: time-resolved LDA crosses chance with significant clusters around 100, 220, and 400 ms; 16 of 21 meta-category contrasts reach significance in cluster-based permutation tests; and the ERP morphology is consistent with expected P1/N200 structure. The paper also provides behavioral attention checks with AUC-based scoring and a large human-rater evaluation of reconstructions, which are welcome additions. However, the two headline claims—effective reconstruction and log-linear scaling—rest almost entirely on ENIGMA, an anonymous in-review model with no architecture or training details and a pointer to a nonexistent Appendix B. The reconstruction table also shows large drops relative to THINGS-EEG2 that contradict the \"comparable\" wording. The dataset contribution is likely to be significant, but the manuscript in its current form does not provide auditable support for the reconstruction and scaling claims.","major_comments":[{"comment":"The reconstruction benchmark, scaling analysis, channel-count analysis, and saliency maps are all obtained with ENIGMA, cited as \"Anonymous. Enigma: A unified lightweight eeg-to-image model for multi-subject visual decoding. In Review, see Appendix B., 2025.\" The preprint's appendices run A.1 through A.10; no Appendix B exists and no architectural, training, or implementation details of ENIGMA are provided. This is a missing reference / omitted proof that the manuscript itself flags with \"see Appendix B.\" Because Table 1 (Alljoined rows), Figure 7A, Figure 7B, and the saliency analysis are all produced with this model, the paper's central quantitative results cannot be audited or reproduced by any reader. Please include a complete description of ENIGMA (architecture, training procedure, hyperparameters, and any code release) in the paper or an appendix, or alternatively remove or substantially qualify the reconstruction and scaling claims. The scaling claim in Section 4 is specifically ENIGMA's learning curve, so without this information the \"no sign of saturating\" conclusion is unsupported.","section":"Section 4, Ref [2], Appendices"},{"comment":"The text states that reconstructions on Alljoined-1.6M \"produced reconstructions with quantitative scores comparable to those of THINGS-EE2,\" but the numbers in Table 1 do not support this wording. For ENIGMA, on Alljoined-1.6M versus THINGS-EEG2: AlexNet(2) drops from 81.89% to 63.62%, CLIP from 78.90% to 62.91%, Top-1 retrieval from 27.60% to 6.00%, and human identification accuracy from 83.06% to 65.43%. ATM-S and Perceptogram also show large declines on most metrics. While a 65.43% human identification accuracy is above chance and indicates some preserved information, the reconstruction quality is markedly lower on Alljoined-1.6M than on THINGS-EEG2. The phrase \"comparable\" is misleading and should be replaced with a quantitative description of the performance gap, along with appropriate significance tests or confidence intervals for the metric differences.","section":"Section 4, Table 1"},{"comment":"The paper states that \"Millisecond-accurate triggers delivered through the Emotiv API\" aligned image onset with the EEG timeline, and reports discarding 0-0.6% of trials for synchronization mismatches. However, no quantification of residual trigger jitter or latency is given for the wireless Bluetooth 5.2 connection. All time-resolved analyses—ERP peaks, time-resolved LDA decoding, and saliency maps—depend on precise stimulus-to-EEG alignment. Please provide a validation of trigger timing (e.g., a photodiode or analog stimulus channel recorded simultaneously with EEG, or a distribution of trigger delays across trials) and discuss how any residual jitter might affect the reported temporal effects. Without this, the millisecond-level temporal claims rest on an unverified assumption.","section":"Section 3, Hardware and Recording Setup"},{"comment":"The paper's own saliency analysis concludes that \"the model relies almost entirely on low-level visual cues,\" with \"virtually identical occipital P1/N1 footprint across categories.\" This directly qualifies the abstract's claim that the paper demonstrates \"decoding of high-level semantic information from EEG of seen images.\" The cluster-based ERP contrasts and LDA decoding may also be driven by low-level image statistics correlated with the meta-categories rather than abstract semantic content. Please reconcile this apparent contradiction: either temper the \"high-level semantic decoding\" claim, or provide additional analyses (e.g., controlling for low-level image features such as spatial frequency, luminance, or entropy) that isolate semantic content from low-level confounds.","section":"Section 4, Saliency Maps"}],"minor_comments":[{"comment":"The footnote says \"Additional details on the metrics used are in Appendix A.3,\" but the metrics are described in Appendix A.5; Appendix A.3 is titled \"Data Collection Details.\" Please correct the cross-reference.","section":"Section 4, Table 1 footnote"},{"comment":"There is a typo in \"comparable to those of THINGS-EE2\" — the dataset name should be THINGS-EEG2.","section":"Section 4, EEG-to-Image Reconstruction"},{"comment":"The phrase \"difficult or unpleasant to work with\" used to describe why four participants were excluded is subjective and potentially stigmatizing; consider rewording to describe the behavioral criteria more neutrally, e.g., \"inconsistent with experimental instructions.\"","section":"Appendix A.3"},{"comment":"The sentence \"This corresponds to Layout 1 in 8)\" has a malformed reference; it should refer to Figure 8 or the appropriate subpanel.","section":"Appendix A.2"},{"comment":"The text says \"each participant completed 4 x 20,880 = 83,520 image trials\" and later that training images were shown 4-5 times and test images 80 times. Please make the repetition counts explicit for the training and test sets so readers can verify the trial arithmetic.","section":"Section 3, Dataset Scale"},{"comment":"The caption says reconstructions were \"selected ... with the highest scores on all of the image feature metrics in Table 1,\" but if the selections are based on all metrics jointly, this could induce selection bias in the qualitative display; please clarify how the exemplars were chosen.","section":"Figure 5 caption"}],"recommendation":"major_revision","confidential_remarks":"If ENIGMA is the authors' own in-review work, the anonymous citation to a nonexistent Appendix B is a serious transparency problem for a dataset paper that will be used as a benchmark. The authors should either describe the model fully, release its code, or explicitly state that the reconstruction and scaling results are preliminary. The \"comparable\" language in Section 4 is also not supported by Table 1 and should be corrected before publication. The dataset itself and the basic decoding analyses are worth publishing, but the current manuscript does not meet the standard of auditability expected for a benchmark resource."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a genuinely useful dataset paper, and the headline claim that a ~$2.2k 32-channel consumer EEG system can support large-scale visual decoding research mostly holds up for the basic analyses. The dataset is new, twice the size of THINGS-EEG2, released with code and preprocessed data, and the cheap-hardware angle matters for democratization. The ERP morphology looks right, the cluster-permutation tests show 16/21 meta-category contrasts with significant effects, and the time-resolved LDA decodes above chance with sensible temporal structure. Those analyses are simple, transparent, and credible.\n\nThe problem is that the reconstruction and scaling sections rest entirely on ENIGMA, an anonymous in-review manuscript that the paper cross-references to an Appendix B that does not exist. Nothing about the architecture or training is given. So the two strongest quantitative claims—effective EEG-to-image reconstruction and log-linear scaling without saturation—cannot be audited by any reader. That is a load-bearing flaw in the current preprint. The other two reconstruction methods (ATM-S, Perceptogram) do show above-chance identification, but they are not used for the scaling curves, and their scores are far below the paper's 'comparable to THINGS-EEG2' language. Table 1 shows Alljoined metrics 10-20 points lower across the board. That is not 'comparable' in any ordinary sense.\n\nTwo smaller issues. First, the trigger-timing section reports 0-0.6% discarded trials but never quantifies residual Bluetooth jitter, so users cannot assess the timing precision of epochs. The fact that P1/N200 peaks appear at the expected latencies suggests the alignment is roughly correct, but it deserves explicit measurement. Second, participant retention: 20 of 48 initial recruits, and four exclusions were for being 'difficult or unpleasant to work with.' The paper is transparent about this in the appendix, but the real-world deployability conclusion should be softened.\n\nThe meta-category groupings and the electrode montage were chosen on THINGS-EEG2, so they are not independent validation, but the paper does not pretend otherwise. The behavioral 2AFC experiment is a nice addition.\n\nBottom line: the dataset is a real contribution and the basic decoding evidence is solid. The reconstruction and scaling claims need the ENIGMA model fully documented or removed, and the comparison language needs correction. Send it to peer review with a request for major revisions.","headline":"A genuinely useful consumer-EEG dataset with solid basic decoding evidence, but the headline reconstruction and scaling claims rest on an unauditable anonymous model and need major revision before publication.","tokens_in":18886,"tokens_out":2178,"would_cite":true,"duration_ms":21128,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Cheap EEG decodes seen images when given 1.6 million trials","keywords":["EEG dataset","consumer-grade EEG","brain-computer interface","visual decoding","EEG-to-Image reconstruction","semantic decoding","scaling laws","THINGS"],"falsifier":"Measure the actual trigger latency of the Emotiv Flex 2 by presenting a photodiode-verified stimulus marker through the same Emotiv API while recording EEG, and compare the observed P1 latency variance against the advertised sub-4 ms temporal precision; if jitter exceeds a few milliseconds, or if the 0-0.6% discarded-trial rate masks systematic delays, the ERP and decoding timing conclusions would need revision.","tokens_in":17734,"feed_emoji":"🧠","tokens_out":4970,"duration_ms":51220,"temperature":0.7,"pith_summary":"This paper presents Alljoined-1.6M, a new open EEG dataset of more than 1.6 million image-viewing trials from 20 people, recorded with a 32-channel consumer headset costing about $2.2k. The authors set out to test whether such affordable hardware, despite its lower signal-to-noise ratio, can support the same deep-learning decoding tasks normally run on systems that are roughly 27 times more expensive. They report that high-level semantic information about viewed images can be decoded from the data, that EEG-to-Image reconstruction models trained on it produce usable reconstructions, and that decoding performance grows log-linearly with trial count with no sign of saturation. If true, this weakens the assumption that brain-computer interface research requires lab-grade EEG, and it makes large-scale data collection feasible for small labs and real-world deployments.","feed_headline":"Cheap EEG decodes seen images when given 1.6 million trials","feed_subtitle":"A $2.2k headset plus big data yields semantic decoding, image reconstruction, and log-linear scaling without saturation.","key_machinery":"The load-bearing object is the dataset itself: 1.6 million trials of 250 Hz EEG epochs spanning -200 ms to 1000 ms around image onset, from 20 subjects, four sessions each, over 16,740 THINGS images with a train/test split that separates both images and object categories. Its power comes from combining a large trial count with repeated presentations of the same test images (80 times per subject), allowing within-subject averaging to raise signal-to-noise ratio, plus questionnaire metadata that makes trait and state confounds explicit. The decoding analyses run through three mechanisms: time-resolved linear discriminant analysis for pairwise category decoding, the ENIGMA encoder, which maps EEG trials to CLIP image-language embeddings for retrieval and reconstruction, and a subsampling protocol that fits ENIGMA on progressively larger trial subsets to measure scaling.","core_discovery":"The central claim is that data volume can compensate for hardware quality in EEG-based visual decoding. Using the Emotiv Flex 2, a 32-channel wireless system roughly 27 times cheaper than the 64-channel research-grade amplifier used in THINGS-EEG2, the authors collected 1.6 million stimulus-locked trials across 20 subjects and report above-chance pairwise category decoding, significant category-selective ERP clusters in 16 of 21 comparisons, and EEG-to-Image reconstructions from ENIGMA, ATM-S, and Perceptogram that human raters identify correctly 62-65% of the time in a two-alternative forced-choice task. They further claim that reconstruction quality improves log-linearly with training data and has not saturated at the full dataset size, and that reducing the montage to about 24 channels costs little performance. The paper's point is that the binding constraint on EEG decoding research is no longer hardware price but dataset scale.","pith_inferences":["The log-linear scaling result, if it extends beyond the single encoder tested, implies that the cost-performance frontier could be crossed by crowdsourced at-home recordings, where low hardware cost makes 10^7-trial collections plausible.","Because the electrode montage was chosen by ablating decoding performance on THINGS-EEG2, a natural follow-up is to run the same channel ablation on Alljoined-1.6M itself to see whether the optimal occipital layout shifts under lower signal-to-noise conditions.","The saliency result, which shows the ENIGMA model relying mainly on early occipital cues at 160-300 ms, suggests that reconstruction on consumer hardware may be driven by low-level visual regularities; a testable extension is to compare retrieval accuracy across meta-categories matched for low-level image statistics.","The dataset's repeated test-image presentations (80 per subject) enable a direct estimate of single-trial versus averaged-trial decoding ceilings, which could quantify how much of the gap to research-grade EEG is noise rather than missing neural information."],"forward_implications":["Other groups can run semantic decoding and EEG-to-Image reconstruction research with equipment costing around $2.2k instead of roughly $60k, lowering the entry barrier for small labs.","Collecting more data on consumer hardware is a reliable route to better decoding: the reported log-linear scaling with no saturation implies that moving toward 10 million trials would continue to buy accuracy.","A 32-channel montage is not the decisive limit; the channel ablation suggests that tests with 24 or fewer channels can still be worthwhile, which matters for portable headsets.","The dataset provides a realistic low-SNR benchmark in which architecture choices become visible, since the more complex ATM-S underperforms relative to simpler linear and multi-subject models on this data.","The dataset can support models that generalize across subjects and categories because training and test images and categories do not overlap, reducing the confounds that plagued earlier visual EEG datasets."],"supporting_citations":[{"why":"Supplies the THINGS-EEG2 comparison benchmark, the identical 16,740 stimulus set, the 250 Hz data format, and the prior decoding and reconstruction numbers the paper measures itself against.","marker":"[21]"},{"why":"Provides the earlier THINGS-EEG1 dataset and establishes the experimental lineage of large-scale visual EEG collection within the THINGS initiative.","marker":"[24]"},{"why":"Supplies the subsampling protocol used for the scaling analysis and the prior evidence that decoding performance improves log-linearly with data volume.","marker":"[5]"},{"why":"Supplies the ENIGMA EEG-to-image model, the main architecture used for reconstruction, retrieval, saliency, and scaling analyses in this paper.","marker":"[2]"},{"why":"Supplies the Perceptogram EEG-to-image reconstruction method that is reproduced and benchmarked on the new dataset.","marker":"[19]"},{"why":"Supplies the ATM-S visual decoding and reconstruction method that is reproduced and benchmarked on the new dataset.","marker":"[37]"},{"why":"Supplies the THINGS image database, the source of the 16,740 naturalistic object images and their 1,854 category labels.","marker":"[26]"},{"why":"Provides the large-scale repeated-measurement study design precedent (NSD) for collecting multiple sessions per participant over matched stimuli.","marker":"[1]"},{"why":"Supplies the cluster-based permutation testing procedure used to establish the statistical significance of category-selective ERP and decoding effects.","marker":"[41]"}],"fun_headline_variants":["Data scales cheap EEG to visual decoding","1.6M trials: cheap EEG decodes visual semantics","$2.2k EEG + 1.6M trials = image decoding","Budget EEG image decoding via massive data","Affordable EEG: 1.6M trials enable reconstruction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole timing analysis rests on the claim that the wireless trigger stream aligns image onset to the EEG timeline with millisecond accuracy; if Bluetooth or software delays jitter the triggers, the P1 and N200 peaks and the stimulus-locked decoding results would be smeared or shifted.","fun_headline_variants_meta":{"raw":{"variants":["Data scales cheap EEG to visual decoding","1.6M trials: cheap EEG decodes visual semantics","$2.2k EEG + 1.6M trials = image decoding","Budget EEG image decoding via massive data","Affordable EEG: 1.6M trials enable reconstruction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001392,"raw_usage":{"total_tokens":5691,"prompt_tokens":1064,"completion_tokens":4627,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":680,"completion_tokens_details":{"reasoning_tokens":4548}},"tokens_in":680,"tokens_out":4627,"duration_ms":33423,"temperature":1.0,"reasoning_tokens":4548,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:56:31.228104+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the actual trigger latency of the Emotiv Flex 2 by presenting a photodiode-verified stimulus marker through the same Emotiv API while recording EEG, and compare the observed P1 latency variance against the advertised sub-4 ms temporal precision; if jitter exceeds a few milliseconds, or if the 0-0.6% discarded-trial rate masks systematic delays, the ERP and decoding timing conclusions would need revision.","supporting_citations":[{"cited_title":"Gifford, Kshitij Dwivedi, Gemma Roig, and Radoslaw M","cited_arxiv_id":null,"evidence_quote":"Supplies the THINGS-EEG2 comparison benchmark, the identical 16,740 stimulus set, the 250 Hz data format, and the prior decoding and reconstruction numbers the paper measures itself against."},{"cited_title":"Robinson, Michael N","cited_arxiv_id":null,"evidence_quote":"Provides the earlier THINGS-EEG1 dataset and establishes the experimental lineage of large-scale visual EEG collection within the THINGS initiative."},{"cited_title":"Enigma: A unified lightweight eeg-to-image model for multi-subject visual decoding","cited_arxiv_id":null,"evidence_quote":"Supplies the ENIGMA EEG-to-image model, the main architecture used for reconstruction, retrieval, saliency, and scaling analyses in this paper."},{"cited_title":"Visual Decoding and Reconstruction via EEG Embeddings with Guided Diffusion","cited_arxiv_id":null,"evidence_quote":"Supplies the ATM-S visual decoding and reconstruction method that is reproduced and benchmarked on the new dataset."},{"cited_title":"Hebart, Adam H","cited_arxiv_id":null,"evidence_quote":"Supplies the THINGS image database, the source of the 16,740 naturalistic object images and their 1,854 category labels."},{"cited_title":"Allen, Ghislain St-Yves, Yihan Wu, Jesse L","cited_arxiv_id":null,"evidence_quote":"Provides the large-scale repeated-measurement study design precedent (NSD) for collecting multiple sessions per participant over matched stimuli."},{"cited_title":"Nonparametric statistical testing of eeg-and meg-data","cited_arxiv_id":null,"evidence_quote":"Supplies the cluster-based permutation testing procedure used to establish the statistical significance of category-selective ERP and decoding effects."}],"review_version":2}