Pith. sign in

REVIEW 4 major objections 6 minor 69 references

Alljoined-1.6M: A Million-Trial EEG-Image Dataset for Evaluating Affordable Brain-Computer Interfaces

T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Cheap EEG decodes seen images when given 1.6 million trials

desk verdict A genuinely useful consumer-EEG dataset with solid basic decoding evidence, but the headline reconstruction and scaling claims rest on an unauditable anonymous model and need major revision before publication. read the letter →

arxiv 2508.18571 v2 pith:Z2ONI4LR submitted 2025-08-26 q-bio.NC

classification q-bio.NC
keywords EEGdatasetconsumer-gradebrain-computerinterfacevisualdecodingEEG-to-ImagereconstructionsemanticscalinglawsTHINGS
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper presents Alljoined-1.6M, a new open EEG dataset of more than 1.6 million image-viewing trials from 20 people, recorded with a 32-channel consumer headset costing about $2.2k. The authors set out to test whether such affordable hardware, despite its lower signal-to-noise ratio, can support the same deep-learning decoding tasks normally run on systems that are roughly 27 times more expensive. They report that high-level semantic information about viewed images can be decoded from the data, that EEG-to-Image reconstruction models trained on it produce usable reconstructions, and that decoding performance grows log-linearly with trial count with no sign of saturation. If true, this weakens the assumption that brain-computer interface research requires lab-grade EEG, and it makes large-scale data collection feasible for small labs and real-world deployments.

What carries the argument

The load-bearing object is the dataset itself: 1.6 million trials of 250 Hz EEG epochs spanning -200 ms to 1000 ms around image onset, from 20 subjects, four sessions each, over 16,740 THINGS images with a train/test split that separates both images and object categories. Its power comes from combining a large trial count with repeated presentations of the same test images (80 times per subject), allowing within-subject averaging to raise signal-to-noise ratio, plus questionnaire metadata that makes trait and state confounds explicit. The decoding analyses run through three mechanisms: time-resolved linear discriminant analysis for pairwise category decoding, the ENIGMA encoder, which maps EEG trials to CLIP image-language embeddings for retrieval and reconstruction, and a subsampling protocol that fits ENIGMA on progressively larger trial subsets to measure scaling.

What would settle it

Measure the actual trigger latency of the Emotiv Flex 2 by presenting a photodiode-verified stimulus marker through the same Emotiv API while recording EEG, and compare the observed P1 latency variance against the advertised sub-4 ms temporal precision; if jitter exceeds a few milliseconds, or if the 0-0.6% discarded-trial rate masks systematic delays, the ERP and decoding timing conclusions would need revision.

Watch

Extended reading notes

Core claim

The central claim is that data volume can compensate for hardware quality in EEG-based visual decoding. Using the Emotiv Flex 2, a 32-channel wireless system roughly 27 times cheaper than the 64-channel research-grade amplifier used in THINGS-EEG2, the authors collected 1.6 million stimulus-locked trials across 20 subjects and report above-chance pairwise category decoding, significant category-selective ERP clusters in 16 of 21 comparisons, and EEG-to-Image reconstructions from ENIGMA, ATM-S, and Perceptogram that human raters identify correctly 62-65% of the time in a two-alternative forced-choice task. They further claim that reconstruction quality improves log-linearly with training data and has not saturated at the full dataset size, and that reducing the montage to about 24 channels costs little performance. The paper's point is that the binding constraint on EEG decoding research is no longer hardware price but dataset scale.

Load-bearing premise

The whole timing analysis rests on the claim that the wireless trigger stream aligns image onset to the EEG timeline with millisecond accuracy; if Bluetooth or software delays jitter the triggers, the P1 and N200 peaks and the stimulus-locked decoding results would be smeared or shifted.

Editorial extensions

If this is right

  • Other groups can run semantic decoding and EEG-to-Image reconstruction research with equipment costing around $2.2k instead of roughly $60k, lowering the entry barrier for small labs.
  • Collecting more data on consumer hardware is a reliable route to better decoding: the reported log-linear scaling with no saturation implies that moving toward 10 million trials would continue to buy accuracy.
  • A 32-channel montage is not the decisive limit; the channel ablation suggests that tests with 24 or fewer channels can still be worthwhile, which matters for portable headsets.
  • The dataset provides a realistic low-SNR benchmark in which architecture choices become visible, since the more complex ATM-S underperforms relative to simpler linear and multi-subject models on this data.
  • The dataset can support models that generalize across subjects and categories because training and test images and categories do not overlap, reducing the confounds that plagued earlier visual EEG datasets.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The log-linear scaling result, if it extends beyond the single encoder tested, implies that the cost-performance frontier could be crossed by crowdsourced at-home recordings, where low hardware cost makes 10^7-trial collections plausible.
  • Because the electrode montage was chosen by ablating decoding performance on THINGS-EEG2, a natural follow-up is to run the same channel ablation on Alljoined-1.6M itself to see whether the optimal occipital layout shifts under lower signal-to-noise conditions.
  • The saliency result, which shows the ENIGMA model relying mainly on early occipital cues at 160-300 ms, suggests that reconstruction on consumer hardware may be driven by low-level visual regularities; a testable extension is to compare retrieval accuracy across meta-categories matched for low-level image statistics.
  • The dataset's repeated test-image presentations (80 per subject) enable a direct estimate of single-trial versus averaged-trial decoding ceilings, which could quantify how much of the gap to research-grade EEG is noise rather than missing neural information.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces Alljoined-1.6M, an EEG dataset of over 1.6 million visual stimulus trials recorded from 20 participants with a 32-channel Emotiv Flex 2 consumer-grade headset, using the THINGS image set and a rapid serial visual presentation paradigm. The authors report ERP analyses with cluster-based permutation tests, time-resolved pair-wise LDA decoding across seven meta-categories, EEG-to-image reconstruction benchmarks using ENIGMA, ATM-S, and Perceptogram, saliency maps, a scaling analysis of reconstruction performance against training set size, and a channel-count ablation. The central claims are that consumer-grade hardware can support high-level semantic decoding and effective EEG-to-image reconstruction, and that decoding performance scales log-linearly with data volume without saturation. The dataset and benchmark code are planned for public release.

Significance. A public dataset of this scale recorded on affordable hardware would be a valuable community resource for studying the trade-off between hardware cost, signal fidelity, and data volume in visual EEG decoding. The basic decoding evidence is solid: time-resolved LDA crosses chance with significant clusters around 100, 220, and 400 ms; 16 of 21 meta-category contrasts reach significance in cluster-based permutation tests; and the ERP morphology is consistent with expected P1/N200 structure. The paper also provides behavioral attention checks with AUC-based scoring and a large human-rater evaluation of reconstructions, which are welcome additions. However, the two headline claims—effective reconstruction and log-linear scaling—rest almost entirely on ENIGMA, an anonymous in-review model with no architecture or training details and a pointer to a nonexistent Appendix B. The reconstruction table also shows large drops relative to THINGS-EEG2 that contradict the "comparable" wording. The dataset contribution is likely to be significant, but the manuscript in its current form does not provide auditable support for the reconstruction and scaling claims.

major comments (4)
  1. [Section 4, Ref [2], Appendices] The reconstruction benchmark, scaling analysis, channel-count analysis, and saliency maps are all obtained with ENIGMA, cited as "Anonymous. Enigma: A unified lightweight eeg-to-image model for multi-subject visual decoding. In Review, see Appendix B., 2025." The preprint's appendices run A.1 through A.10; no Appendix B exists and no architectural, training, or implementation details of ENIGMA are provided. This is a missing reference / omitted proof that the manuscript itself flags with "see Appendix B." Because Table 1 (Alljoined rows), Figure 7A, Figure 7B, and the saliency analysis are all produced with this model, the paper's central quantitative results cannot be audited or reproduced by any reader. Please include a complete description of ENIGMA (architecture, training procedure, hyperparameters, and any code release) in the paper or an appendix, or alternatively remove or substantially qualify the reconstruction and scaling claims. The scaling claim in Section 4 is specifically ENIGMA's learning curve, so without this information the "no sign of saturating" conclusion is unsupported.
  2. [Section 4, Table 1] The text states that reconstructions on Alljoined-1.6M "produced reconstructions with quantitative scores comparable to those of THINGS-EE2," but the numbers in Table 1 do not support this wording. For ENIGMA, on Alljoined-1.6M versus THINGS-EEG2: AlexNet(2) drops from 81.89% to 63.62%, CLIP from 78.90% to 62.91%, Top-1 retrieval from 27.60% to 6.00%, and human identification accuracy from 83.06% to 65.43%. ATM-S and Perceptogram also show large declines on most metrics. While a 65.43% human identification accuracy is above chance and indicates some preserved information, the reconstruction quality is markedly lower on Alljoined-1.6M than on THINGS-EEG2. The phrase "comparable" is misleading and should be replaced with a quantitative description of the performance gap, along with appropriate significance tests or confidence intervals for the metric differences.
  3. [Section 3, Hardware and Recording Setup] The paper states that "Millisecond-accurate triggers delivered through the Emotiv API" aligned image onset with the EEG timeline, and reports discarding 0-0.6% of trials for synchronization mismatches. However, no quantification of residual trigger jitter or latency is given for the wireless Bluetooth 5.2 connection. All time-resolved analyses—ERP peaks, time-resolved LDA decoding, and saliency maps—depend on precise stimulus-to-EEG alignment. Please provide a validation of trigger timing (e.g., a photodiode or analog stimulus channel recorded simultaneously with EEG, or a distribution of trigger delays across trials) and discuss how any residual jitter might affect the reported temporal effects. Without this, the millisecond-level temporal claims rest on an unverified assumption.
  4. [Section 4, Saliency Maps] The paper's own saliency analysis concludes that "the model relies almost entirely on low-level visual cues," with "virtually identical occipital P1/N1 footprint across categories." This directly qualifies the abstract's claim that the paper demonstrates "decoding of high-level semantic information from EEG of seen images." The cluster-based ERP contrasts and LDA decoding may also be driven by low-level image statistics correlated with the meta-categories rather than abstract semantic content. Please reconcile this apparent contradiction: either temper the "high-level semantic decoding" claim, or provide additional analyses (e.g., controlling for low-level image features such as spatial frequency, luminance, or entropy) that isolate semantic content from low-level confounds.
minor comments (6)
  1. [Section 4, Table 1 footnote] The footnote says "Additional details on the metrics used are in Appendix A.3," but the metrics are described in Appendix A.5; Appendix A.3 is titled "Data Collection Details." Please correct the cross-reference.
  2. [Section 4, EEG-to-Image Reconstruction] There is a typo in "comparable to those of THINGS-EE2" — the dataset name should be THINGS-EEG2.
  3. [Appendix A.3] The phrase "difficult or unpleasant to work with" used to describe why four participants were excluded is subjective and potentially stigmatizing; consider rewording to describe the behavioral criteria more neutrally, e.g., "inconsistent with experimental instructions."
  4. [Appendix A.2] The sentence "This corresponds to Layout 1 in 8)" has a malformed reference; it should refer to Figure 8 or the appropriate subpanel.
  5. [Section 3, Dataset Scale] The text says "each participant completed 4 x 20,880 = 83,520 image trials" and later that training images were shown 4-5 times and test images 80 times. Please make the repetition counts explicit for the training and test sets so readers can verify the trial arithmetic.
  6. [Figure 5 caption] The caption says reconstructions were "selected ... with the highest scores on all of the image feature metrics in Table 1," but if the selections are based on all metrics jointly, this could induce selection bias in the qualitative display; please clarify how the exemplars were chosen.

Circularity Check

1 steps flagged · score 6.0 of 10

EEG-to-image reconstruction and scaling claims rest on ENIGMA, an anonymous in-review method whose only specification is a cross-reference to a nonexistent Appendix B.

  1. self citation load bearing [Section 4, 'EEG-to-Image Reconstruction'; Reference [2]; Appendix A.10]
    "For our analysis, we took all publicly available EEG-to-Image reconstruction methods (ENIGMA [2], ATM-S [3], and Perceptogram [4]) and reproduced their methods on our dataset. ... [2] Anonymous. Enigma: A unified lightweight eeg-to-image model for multi-subject visual decoding. In Review, see Appendix B., 2025."

    Every headline reconstruction number (Table 1, Alljoined-1.6M ENIGMA rows), the log-linear scaling curves (Fig. 7A), the channel-count analysis, and the saliency maps are produced with ENIGMA. The only specification given for ENIGMA is reference [2], which directs the reader to 'Appendix B' of this manuscript, yet the paper's appendices run only A.1-A.10 and contain no ENIGMA architecture, loss, training, or evaluation details. The central quantitative claims therefore rest on a self-referential ghost citation: the reader cannot determine whether the stated results follow from a reproducible model or reduce to some construction hidden in the missing appendix.

full rationale

The dataset contribution itself, the ERP analyses, the cluster-based permutation tests, and the pairwise LDA decoding are self-contained: they use only the released EEG recordings, standard preprocessing, and transparent decoders, and they provide independent evidence that category information is decodable from the Emotiv recordings. The human-rated 2AFC identification scores for ATM-S and Perceptogram on Alljoined-1.6M (60.31% and 62.00%) are also grounded in published, code-released methods, and they support the weaker claim that the dataset supports above-chance reconstruction identification. What prevents a low score is that the two headline quantitative results—'effective EEG-to-Image reconstruction' at a level comparable to THINGS-EEG2 and 'log-linear decoding performance with increasing data volume'—are carried by ENIGMA, whose only specification is reference [2]: 'In Review, see Appendix B,' while the appendices run only A.1-A.10. The manuscript itself points to the nonexistent appendix, admitting the missing support. Because ENIGMA's architecture, training, and evaluation cannot be audited, the reconstruction and scaling conclusions are load-bearing on a self-referential citation rather than on a reproducible derivation. This is partial circularity rather than full self-definition, because independent baselines still substantiate the dataset's basic usability.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The main choices tuned by hand or external data are the recording montage, chosen by decoding ablations on THINGS-EEG2, and the meta-category taxonomy, created by LLMs plus manual curation. The central model ENIGMA is an unauditable external assumption and is listed as an axiom rather than a parameter. No new physical entities are introduced.

free parameters (2)
  • Electrode montage selection = 32-channel layout 1 (occipital/central 10-20 subset)
    The montage was chosen by running decoding ablations on THINGS-EEG2 (Appendix A.2), so the recording configuration is tuned on external data and may favor paradigms similar to THINGS-EEG2.
  • Meta-category groupings = 7 semantic groups, e.g., Toys/Games/Musical Instruments merged
    Created with ChatGPT 4o and Gemini 2.5 Pro Preview plus manual review (Appendix A.8); buildings and outdoor scenes were dropped because they did not appear in the test set. This is a hand-made label taxonomy, not a natural grouping.
assumptions (5)
  • domain assumption Emotiv Flex 2 delivers millisecond-accurate stimulus triggers over Bluetooth.
    Section 3 (Hardware and Recording Setup) states 'Millisecond-accurate triggers delivered through the Emotiv API.' All ERP timing, LDA time courses, and saliency maps rely on this; residual jitter is not quantified, and 0-0.6% of trials were dropped for synchronization mismatches.
  • domain assumption Each EEG epoch is dominated by the currently displayed image despite the 100 ms image + 100 ms blank RSVP.
    Section 3 and Appendix A.9 acknowledge overlapping neural responses to consecutive images; decoding analyses nonetheless label each epoch with the current stimulus and do not model the previous image's trailing activity.
  • ad hoc to paper ENIGMA (Ref [2]) is a valid, correctly implemented model that can serve as a benchmark.
    Reference [2] is an anonymous in-review manuscript with an Appendix B cited but absent from this preprint; the model's architecture and training are unavailable, yet ENIGMA drives Table 1, Figures 5-7, and the saliency/channel analyses.
  • domain assumption The cost model ($50/hour collection, ~$2.2k hardware) is a fair basis for the cost-performance comparison.
    Figure 1C and the discussion use these rates, stated without source data; conclusions about the financial investment needed to reach a given decoding benchmark depend on them.
  • domain assumption Participant screening to 20 of 48 volunteers does not materially bias the decoding conclusions.
    Section 3 says participants were 'filtered for participants who had high behavioral scores and high task engagement,' and Appendix A.3 reports 6 exclusions for EEG quality and 4 for behavioral/interpersonal factors. If retained participants are unrepresentative of real-world users, the feasibility claim may be overstated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Alljoined-1.6M: A Million-Trial EEG-Image Dataset for Evaluating Affordable Brain-Computer Interfaces." pith.science (2026). https://pith.science/paper/Z2ONI4LR

@misc{pith2026250818571,
  author       = {Pith},
  title        = {Pith review of: Alljoined-1.6M: A Million-Trial EEG-Image Dataset for Evaluating Affordable Brain-Computer Interfaces},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Z2ONI4LR}},
  note         = {Machine review of arXiv:2508.18571}
}
abstract

We present a new large-scale electroencephalography (EEG) dataset as part of the THINGS initiative, comprising over 1.6 million visual stimulus trials collected from 20 participants, and totaling more than twice the size of the most popular current benchmark dataset, THINGS-EEG2. Crucially, our data was recorded using a 32-channel consumer-grade wet electrode system costing ~$2.2k, around 27x cheaper than research-grade EEG systems typically used in cognitive neuroscience labs. Our work is one of the first open-source, large-scale EEG resource designed to closely reflect the quality of hardware that is practical to deploy in real-world, downstream applications of brain-computer interfaces (BCIs). We aim to explore the specific question of whether deep neural network-based BCI research and semantic decoding methods can be effectively conducted with such affordable systems, filling an important gap in current literature that is extremely relevant for future research. In our analysis, we not only demonstrate that decoding of high-level semantic information from EEG of visualized images is possible at consumer-grade hardware, but also that our data can facilitate effective EEG-to-Image reconstruction even despite significantly lower signal-to-noise ratios. In addition to traditional benchmarks, we also conduct analyses of EEG-to-Image models that demonstrate log-linear decoding performance with increasing data volume on our data, and discuss the trade-offs between hardware cost, signal fidelity, and the scale of data collection efforts in increasing the size and utility of currently available datasets. Our contributions aim to pave the way for large-scale, cost-effective EEG research with widely accessible equipment, and position our dataset as a unique resource for the democratization and development of effective deep neural models of visual cognition.

Figures

Figures reproduced from arXiv: 2508.18571 by the authors.

Figure 1
Figure 1. A: A picture of the collection setup for the Alljoined-1.6M dataset. B: A comparison between THINGS-EEG2 [21] and Alljoined-1.6M (ours) in terms of hardware cost and the number of subjects in the dataset. C: A bifurcated plot displaying decoding performance against collection cost. Costs include the purchase of the EEG headset and amplifier, and then a $50/hour collection cost to compensate participants and technici… view at source ↗
Figure 2
Figure 2. Experimental paradigm for Alljoined-1.6M. Details can be found in Section [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Average Event Related Potential (ERP) across all 20 subjects and all 4 sessions for a total [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (32 more)
Figure 4
Figure 4. Figure 4: Average pair-wise decoding across meta-category combinations. Fig. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Qualitative comparison of EEG-to-Image reconstruction methods on Alljoined-1.6M. [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Spatiotemporal saliency for ENIGMA. Warm colors indicate electrodes whose activity diverges most strongly from the grand-average VEP and therefore pushes the model toward the CLIP centroid of that category. All four rows reveal a shared peak over occipital sensors arou…
Figure 7
Figure 7. Figure 7: (A) Scaling analysis of model performance for Alljoined-1.6M and THINGS-EEG2. The number of training samples are plotted on a log-scale X-axis, and the normalized average of feature metrics presented in [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Various electrode positions and subsets from the 10-20 layout, compared in ablations on [PITH_FULL_IMAGE:figures/full_fig_p015_8.png]
Figure 9
Figure 9. Figure 9: High and low level performance for image reconstruction on various 32 channel electrode [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]
Figure 10
Figure 10. Figure 10: AUC values for each of the 20 participants based on their behavioral responses in the [PITH_FULL_IMAGE:figures/full_fig_p017_10.png]
Figure 11
Figure 11. Figure 11: An example of the 2 alternative forced choice task used in our behavioral experiment [PITH_FULL_IMAGE:figures/full_fig_p019_11.png]
Figure 12
Figure 12. Figure 12: Cluster analyses results for the contrast between Animals and Foods/Plants. Yellow shaded [PITH_FULL_IMAGE:figures/full_fig_p020_12.png]
Figure 13
Figure 13. Figure 13: Cluster analyses results for the contrast between Animals and Body Parts/Apparel. Yellow [PITH_FULL_IMAGE:figures/full_fig_p020_13.png]
Figure 14
Figure 14. Figure 14: Cluster analyses results for the contrast between Animals and Household Items/Furniture. [PITH_FULL_IMAGE:figures/full_fig_p021_14.png]
Figure 15
Figure 15. Figure 15: Cluster analyses results for the contrast between Animals and Tools. Yellow shaded area [PITH_FULL_IMAGE:figures/full_fig_p021_15.png]
Figure 16
Figure 16. Figure 16: Cluster analyses results for the contrast between Animals and Toys/Games. Yellow shaded [PITH_FULL_IMAGE:figures/full_fig_p022_16.png]
Figure 17
Figure 17. Figure 17: Cluster analyses results for the contrast between Animals and Vehicles. Yellow shaded [PITH_FULL_IMAGE:figures/full_fig_p022_17.png]
Figure 18
Figure 18. Figure 18: Cluster analyses results for the contrast between Body Parts/Apparel and Household [PITH_FULL_IMAGE:figures/full_fig_p023_18.png]
Figure 19
Figure 19. Figure 19: Cluster analyses results for the contrast between Body Parts/Apparel and Toys/Games. [PITH_FULL_IMAGE:figures/full_fig_p023_19.png]
Figure 20
Figure 20. Figure 20: Cluster analyses results Foods/Plants and Body Parts/Apparel. Yellow shaded area show [PITH_FULL_IMAGE:figures/full_fig_p024_20.png]
Figure 21
Figure 21. Figure 21: Cluster analyses results for the contrast between Foods/Plants and Household [PITH_FULL_IMAGE:figures/full_fig_p024_21.png]
Figure 22
Figure 22. Figure 22: Cluster analyses results for the contrast between Foods/Plants and Tools. Yellow shaded [PITH_FULL_IMAGE:figures/full_fig_p025_22.png]
Figure 23
Figure 23. Figure 23: Cluster analyses results for the contrast between Foods/Plants and Toys/Games. Yellow [PITH_FULL_IMAGE:figures/full_fig_p025_23.png]
Figure 24
Figure 24. Figure 24: Cluster analyses results for the contrast between Foods/Plants and Vehicles. Yellow shaded [PITH_FULL_IMAGE:figures/full_fig_p026_24.png]
Figure 25
Figure 25. Figure 25: Cluster analyses results for the contrast between Tools and Body Parts/Apparel. Yellow [PITH_FULL_IMAGE:figures/full_fig_p026_25.png]
Figure 26
Figure 26. Figure 26: Cluster analyses results for the contrast between Tools and Household Items/Furniture. [PITH_FULL_IMAGE:figures/full_fig_p027_26.png]
Figure 27
Figure 27. Figure 27: Cluster analyses results for the contrast between Tools and Toys/Games. Yellow shaded [PITH_FULL_IMAGE:figures/full_fig_p027_27.png]
Figure 28
Figure 28. Figure 28: Cluster analyses results for the contrast between Vehicles and Body Parts/Apparel. Yellow [PITH_FULL_IMAGE:figures/full_fig_p028_28.png]
Figure 29
Figure 29. Figure 29: Cluster analyses results for the contrast between Vehicles and Household Items/Furniture. [PITH_FULL_IMAGE:figures/full_fig_p028_29.png]
Figure 30
Figure 30. Figure 30: Cluster analyses results for the contrast between Vehicles and Tools. Yellow shaded area [PITH_FULL_IMAGE:figures/full_fig_p029_30.png]
Figure 31
Figure 31. Figure 31: Cluster analyses results for the contrast between Vehicles and Toys/Games. Yellow shaded [PITH_FULL_IMAGE:figures/full_fig_p029_31.png]
Figure 32
Figure 32. Figure 32: Cluster analyses results for the contrast between Household Items/Furniture and [PITH_FULL_IMAGE:figures/full_fig_p030_32.png]
Figure 33
Figure 33. Figure 33: Raw Integrated-Gradients saliency matrix for the foods and plants embedding. Color [PITH_FULL_IMAGE:figures/full_fig_p031_33.png]
Figure 34
Figure 34. Figure 34: The same saliency matrix after convolution with an 11-sample boxcar kernel. Temporal [PITH_FULL_IMAGE:figures/full_fig_p032_34.png]
Figure 35
Figure 35. Figure 35: Activation-maximization mask produced by directly optimizing the input EEG to maximize [PITH_FULL_IMAGE:figures/full_fig_p032_35.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

69 extracted references · 49 canonical work pages

  1. [1]

    Allen, Ghislain St-Yves, Yihan Wu, Jesse L

    Emily J. Allen, Ghislain St-Yves, Yihan Wu, Jesse L. Breedlove, Jacob S. Prince, Logan T. Dow- dle, Matthias Nau, Brad Caron, Franco Pestilli, Ian Charest, J. Benjamin Hutchinson, Thomas Naselaris, and Kendrick Kay. A massive 7T fMRI dataset to bridge cognitive neuroscience and artificial intelligence. Nature Neuroscience, 25(1):116–126, January 2022

  2. [2]

    Enigma: A unified lightweight eeg-to-image model for multi-subject visual decoding

    Anonymous. Enigma: A unified lightweight eeg-to-image model for multi-subject visual decoding. In Review, see Appendix B., 2025

  3. [3]

    Validation of the emotiv epoc® eeg gaming system for measuring research quality auditory erps

    Nicholas A Badcock, Petroula Mousikou, Yatin Mahajan, Peter De Lissa, Johnson Thie, and Genevieve McArthur. Validation of the emotiv epoc® eeg gaming system for measuring research quality auditory erps. PeerJ, 1:e38, 2013

  4. [4]

    Validation of the emotiv epoc eeg system for research quality auditory event-related potentials in children

    Nicholas A Badcock, Kathryn A Preece, Bianca de Wit, Katharine Glenn, Nora Fieder, Johnson Thie, and Genevieve McArthur. Validation of the emotiv epoc eeg system for research quality auditory event-related potentials in children. PeerJ, 3:e907, 2015

  5. [5]

    Scaling laws for decoding images from brain activity

    Hubert Banville, Yohann Benchetrit, Stéphane d’Ascoli, Jérémy Rapin, and Jean-Rémi King. Scaling laws for decoding images from brain activity. arXiv preprint arXiv:2501.15322, 2025

  6. [6]

    Uncovering the structure of clinical eeg signals with self-supervised learning

    Hubert Banville, Yohann Benchetrit, Stéphane d’Ascoli, Jérémy Rapin, and Jean-Rémi King. Uncovering the structure of clinical eeg signals with self-supervised learning. Journal of Neural Engineering, 18(4):046020, 2021

  7. [7]

    Electrophysio- logical studies of face perception in humans

    Shlomo Bentin, Truett Allison, Aina Puce, Erik Perez, and Gregory McCarthy. Electrophysio- logical studies of face perception in humans. Journal of cognitive neuroscience, 8(6):551–565, 1996

  8. [8]

    The acceptability, feasibility, and utility of portable electroencephalography to study resting-state neurophysiology in rural communities

    Supriya Bhavnani, Dhanya Parameshwaran, Kamal Kant Sharma, Debarati Mukherjee, Gauri Divan, Vikram Patel, and Tara C Thiagarajan. The acceptability, feasibility, and utility of portable electroencephalography to study resting-state neurophysiology in rural communities. Frontiers in human neuroscience, 16:802764, 2022

Show all 69 references
  1. [9]

    Language models are few-shot learners

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020

  2. [10]

    Unsupervised learning of visual features by contrasting cluster assignments

    Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. CoRR, abs/2006.09882, 2020

  3. [11]

    Pyles, Austin Marcus, Abhinav Gupta, Michael J

    Nadine Chang, John A. Pyles, Austin Marcus, Abhinav Gupta, Michael J. Tarr, and Elissa M. Aminoff. BOLD5000, a public fMRI dataset while viewing 5000 visual images. Scientific Data, 6(1):49, May 2019. Number: 1 Publisher: Nature Publishing Group

  4. [12]

    Structure-preserved image reconstruction from brain recordings, 2023

    Zijiao Chen, Jonathan Xu, Jiaxin Qing, Ruilin Li, and Juan Helen Zhou. Structure-preserved image reconstruction from brain recordings, 2023

  5. [13]

    Resolving human object recognition in space and time

    Radoslaw M Cichy, Dimitrios Pantazis, and Aude Oliva. Resolving human object recognition in space and time. Nature Neuroscience, 17(3):455–462, 2014

  6. [14]

    Resolving human object recogni- tion in space and time

    Radoslaw Martin Cichy, Dimitrios Pantazis, and Aude Oliva. Resolving human object recogni- tion in space and time. Nature neuroscience, 17(3):455–462, 2014

  7. [15]

    Imagenet: A large- scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large- scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009

  8. [16]

    Bert: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human langu...

  9. [17]

    A p300-based quantitative compari- son between the emotiv epoc headset and a medical eeg device

    Matthieu Duvinage, Thierry Castermans, Thierry Dutoit, Mathieu Petieau, Thomas Hoellinger, Caty De Saedeleer, K Seetharaman, and G Cheron. A p300-based quantitative compari- son between the emotiv epoc headset and a medical eeg device. Biomedical Engineering, 765(1):2012–2764, 2012

  10. [18]

    The face-specific n170 component reflects late stages in the structural encoding of faces

    Martin Eimer. The face-specific n170 component reflects late stages in the structural encoding of faces. Neuroreport, 11(10):2319–2324, 2000

  11. [19]

    Teng Fei, Abhinav Uppal, Ian Jackson, Srinivas Ravishankar, David Wang, and Virginia R. de Sa. Perceptogram: Reconstructing Visual Percepts from EEG. arXiv preprint arXiv:2404.01250,

  12. [20]

    Effect of stimulus size in a visual erp-based bci under rsvp.Sensors, 22(23):9505, 2022

    Álvaro Fernández-Rodríguez, Aube Darves-Bornoz, Francisco Velasco-Álvarez, and Ricardo Ron-Angevin. Effect of stimulus size in a visual erp-based bci under rsvp.Sensors, 22(23):9505, 2022

  13. [21]

    Gifford, Kshitij Dwivedi, Gemma Roig, and Radoslaw M

    Alessandro T. Gifford, Kshitij Dwivedi, Gemma Roig, and Radoslaw M. Cichy. A large and rich eeg dataset for modeling human visual object recognition. NeuroImage, 264:119754, 2022

  14. [22]

    Gordon and Anil K

    Emma C. Gordon and Anil K. Seth. Ethical considerations for the use of brain–computer interfaces for cognitive enhancement. PLOS Biology, 22(10):1–15, 10 2024

  15. [23]

    Engemann, Daniel Strohmeier, Christian Brodbeck, and et al

    Alexandre Gramfort, Martin Luessi, Eric Larson, Denis A. Engemann, Daniel Strohmeier, Christian Brodbeck, and et al. Meg and eeg data analysis with mne-python. Frontiers in Neuroscience, 7:267, 2013

  16. [24]

    Robinson, Michael N

    Tijl Grootswagers, Ivy Zhou, Austin K. Robinson, Michael N. Hebart, and Thomas A. Carlson. Human eeg recordings for 1,854 concepts presented in rapid serial visual presentation streams. Scientific Data, 9:3, 2022

  17. [25]

    Multivariate pattern analysis for meg: A comparison of dissimilarity measures

    Matthias Guggenmos, Philipp Sterzer, and Radoslaw Martin Cichy. Multivariate pattern analysis for meg: A comparison of dissimilarity measures. Neuroimage, 173:434–447, 2018

  18. [26]

    Hebart, Adam H

    Michael N. Hebart, Adam H. Dickter, Alexis Kidder, Anna Corriveau, Cody Van Wicklin, and Chris I. Baker. Things: A database of 1,854 object concepts and more than 26,000 naturalistic object images. PLOS ONE, 14(10):e0223792, 2019

  19. [27]

    Denoising diffusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020

  20. [28]

    Contribution of image statistics and semantics in local vs

    Eric Lützow Holm, Diego Fernández Slezak, and Enzo Tagliazucchi. Contribution of image statistics and semantics in local vs. distributed eeg decoding of rapid serial visual presentation. bioRxiv, pages 2023–09, 2023

  21. [29]

    Generic decoding of seen and imagined objects using hierarchical visual features

    Tomoyasu Horikawa and Yukiyasu Kamitani. Generic decoding of seen and imagined objects using hierarchical visual features. Nature communications, 8(1):15037, 2017

  22. [30]

    Brain-optimized inference improves reconstructions of fMRI brain activity, December 2023

    Reese Kneeland, Jordyn Ojeda, Ghislain St-Yves, and Thomas Naselaris. Brain-optimized inference improves reconstructions of fMRI brain activity, December 2023. arXiv:2312.07705 [cs, q-bio]

  23. [31]

    Reconstructing seen images from human brain activity via guided stochastic search

    Reese Kneeland, Jordyn Ojeda, Ghislain St-Yves, and Thomas Naselaris. Reconstructing seen images from human brain activity via guided stochastic search. In Conference on Cognitive Computational Neuroscience, 2023

  24. [32]

    Second Sight: Using brain-optimized encoding models to align image distributions with human brain activity, June

    Reese Kneeland, Jordyn Ojeda, Ghislain St-Yves, and Thomas Naselaris. Second Sight: Using brain-optimized encoding models to align image distributions with human brain activity, June

  25. [33]

    Advancing wearable bci: Headphone eeg for cognitive load detection in lab and field

    Michael T Knierim, Christian Zimny, Gabriel Ivucic, and Tobias Röddiger. Advancing wearable bci: Headphone eeg for cognitive load detection in lab and field. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 9(1):1–26, 2025. 11

  26. [34]

    Imagenet classification with deep convolutional neural networks

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In F. Pereira, C.J. Burges, L. Bottou, and K.Q. Weinberger, editors, Advances in Neural Information Processing Systems, volume 25. Curran Associates, Inc., 2012

  27. [35]

    Reading senseless sentences: Brain potentials reflect semantic incongruity

    Marta Kutas and Steven A Hillyard. Reading senseless sentences: Brain potentials reflect semantic incongruity. Science, 207(4427):203–205, 1980

  28. [36]

    Lawhern, Alex J

    Vernon J. Lawhern, Alex J. Solon, Nicholas R. Waytowich, Stacey M. Gordon, Christine P. Hung, and Brent J. Lance. Eegnet: a compact convolutional neural network for eeg-based brain–computer interfaces. Journal of Neural Engineering, 15(5):056013, 2018

  29. [37]

    Visual Decoding and Reconstruction via EEG Embeddings with Guided Diffusion

    Dongyang Li, Chen Wei, Shiying Li, Jiachen Zou, and Quanying Liu. Visual Decoding and Reconstruction via EEG Embeddings with Guided Diffusion. InAdvances in Neural Information Processing Systems (NeurIPS), 2024

  30. [38]

    Training on the test set? an analysis of spampinato et al.[31]

    Ren Li, Jared S Johansen, Hamad Ahmed, Thomas V Ilyevsky, Ronnie B Wilbur, Hari M Bharadwaj, and Jeffrey Mark Siskind. Training on the test set? an analysis of spampinato et al.[31]. arXiv preprint arXiv:1812.07697, 2018

  31. [39]

    The perils and pitfalls of block design for eeg classification experiments

    Ren Li, Jared S Johansen, Hamad Ahmed, Thomas V Ilyevsky, Ronnie B Wilbur, Hari M Bharad- waj, and Jeffrey Mark Siskind. The perils and pitfalls of block design for eeg classification experiments. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(1):316–333, 2020

  32. [40]

    An introduction to the event-related potential technique

    Steven J Luck. An introduction to the event-related potential technique. MIT press, 2014

  33. [41]

    Nonparametric statistical testing of eeg-and meg-data

    Eric Maris and Robert Oostenveld. Nonparametric statistical testing of eeg-and meg-data. Journal of neuroscience methods, 164(1):177–190, 2007

  34. [42]

    Eeg microstates as a tool for studying the temporal dynamics of whole-brain neuronal networks: a review

    Christoph M Michel and Thomas Koenig. Eeg microstates as a tool for studying the temporal dynamics of whole-brain neuronal networks: a review. Neuroimage, 180:577–593, 2018

  35. [43]

    Al Wahedi

    Dmitry Mikhaylov, Muhammad Saeed, Mohamed Husain Alhosani, and Yasser F. Al Wahedi. Comparison of eeg signal spectral characteristics obtained with consumer-and research-grade devices. Sensors, 24(24):8108, 2024

  36. [44]

    Omega: the open meg archive

    Guiomar Niso, Christine Rogers, Jeremy T Moreau, Li-Yuan Chen, Cecile Madjar, Samir Das, Elizabeth Bock, François Tadel, Alan C Evans, Pierre Jolicoeur, et al. Omega: the open meg archive. Neuroimage, 124:1182–1187, 2016

  37. [45]

    The temple university hospital eeg data corpus

    Iyad Obeid and Joseph Picone. The temple university hospital eeg data corpus. Frontiers in neuroscience, 10:196, 2016

  38. [46]

    Natural scene reconstruction from fMRI signals using generative latent diffusion

    Furkan Ozcelik and Rufin VanRullen. Natural scene reconstruction from fMRI signals using generative latent diffusion. Scientific Reports, 13, 2023

  39. [47]

    Event-related eeg/meg synchronization and desyn- chronization: basic principles

    Gert Pfurtscheller and FH Lopes Da Silva. Event-related eeg/meg synchronization and desyn- chronization: basic principles. Clinical neurophysiology, 110(11):1842–1857, 1999

  40. [48]

    Updating p300: An integrative theory of p3a and p3b

    John Polich. Updating p300: An integrative theory of p3a and p3b. Clinical Neurophysiology, 118(10):2128–2148, 2007

  41. [49]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pa...

  42. [50]

    Falk, and Jocelyn Faubert

    Yannick Roy, Hubert Banville, Isabela Albuquerque, Alexandre Gramfort, Tiago H. Falk, and Jocelyn Faubert. Deep learning-based electroencephalography analysis: a systematic review. Journal of Neural Engineering, 16(5):051001, 2019

  43. [51]

    A scoping review on the use of consumer-grade eeg devices for research

    Joshua Sabio, Nikolas S Williams, Genevieve M McArthur, and Nicholas A Badcock. A scoping review on the use of consumer-grade eeg devices for research. Plos one, 19(3):e0291186, 2024. 12

  44. [52]

    Scaling law in neural data: Non-invasive speech decoding with 175 hours of eeg data

    Motoshige Sato, Kenichi Tomeoka, Ilya Horiguchi, Kai Arulkumaran, Ryota Kanai, and Shuntaro Sasai. Scaling law in neural data: Non-invasive speech decoding with 175 hours of eeg data. arXiv preprint arXiv:2407.07595, 2024

  45. [53]

    Scotti, Mihir Tripathy, Cesare Kadir Torrico Villanueva, Reese Kneeland, Tong Chen, Ashutosh Narang, Charan Santhirasegaran, Jonathan Xu, Thomas Naselaris, Kenneth A

    Paul S. Scotti, Mihir Tripathy, Cesare Kadir Torrico Villanueva, Reese Kneeland, Tong Chen, Ashutosh Narang, Charan Santhirasegaran, Jonathan Xu, Thomas Naselaris, Kenneth A. Nor- man, and Tanishq Mathew Abraham. Mindeye2: shared-subject models enable fmri-to-image with 1 hour...

  46. [54]

    Reconstructing the mind’s eye: fMRI-to-image with contrastive learning and diffusion priors

    Paul Steven Scotti, Atmadeep Banerjee, Jimmie Goode, Stepan Shabalin, Alex Nguyen, Cohen Ethan, Aidan James Dempster, Nathalie Verlinde, Elad Yundler, David Weisberg, Kenneth Norman, and Tanishq Mathew Abraham. Reconstructing the mind’s eye: fMRI-to-image with contrastive lear...

  47. [55]

    De- coding Natural Images from EEG for Object Recognition

    Yonghao Song, Bingchuan Liu, Xiang Li, Nanlin Shi, Yijun Wang, and Xiaorong Gao. De- coding Natural Images from EEG for Object Recognition. In Proceedings of the International Conference on Learning Representations (ICLR), 2024

  48. [56]

    Deep Learning Human Mind for Automated Visual Classification

    Carlo Spampinato, Sebastiano Palazzo, Ignazio Kavasidis, Daniele Giordano, Nada Souly, and Mubarak Shah. Deep Learning Human Mind for Automated Visual Classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 6809–6818, 2017

  49. [57]

    Axiomatic attribution for deep networks

    Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep networks. In International conference on machine learning, pages 3319–3328. PMLR, 2017

  50. [58]

    Rethinking the inception architecture for computer vision

    Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. CoRR, abs/1512.00567, 2015

  51. [59]

    High-resolution image reconstruction with latent diffusion models from human brain activity

    Yu Takagi and Shinji Nishimoto. High-resolution image reconstruction with latent diffusion models from human brain activity. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14453–14463, 2023

  52. [60]

    Improving visual image reconstruction from human brain activity using latent diffusion models via multiple decoded inputs, 2023

    Yu Takagi and Shinji Nishimoto. Improving visual image reconstruction from human brain activity using latent diffusion models via multiple decoded inputs, 2023

  53. [61]

    Mingxing Tan and Quoc V . Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, Califor...

  54. [62]

    Review of the bci competition iv

    Michael Tangermann, Klaus-Robert Müller, Ad Aertsen, Niels Birbaumer, Christoph Braun, Clemens Brunner, Robert Leeb, Carsten Mehring, Kai J Miller, Gernot R Müller-Putz, et al. Review of the bci competition iv. Frontiers in neuroscience, 6:55, 2012

  55. [63]

    The cambridge centre for ageing and neuroscience (cam-can) data repository: Structural and functional mri, meg, and cognitive data from a cross-sectional adult lifespan sample

    Jason R Taylor, Nitin Williams, Rhodri Cusack, Tibor Auer, Meredith A Shafto, Marie Dixon, Lorraine K Tyler, Richard N Henson, et al. The cambridge centre for ageing and neuroscience (cam-can) data repository: Structural and functional mri, meg, and cognitive data from a cross...

  56. [64]

    Speed of processing in the human visual system

    Simon Thorpe, Denis Fize, and Catherine Marlot. Speed of processing in the human visual system. nature, 381(6582):520–522, 1996

  57. [65]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017

  58. [66]

    Bovik, H.R

    Zhou Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4):600–612, April 2004. Conference Name: IEEE Transactions on Image Processing. 13

  59. [67]

    odd-ball

    Nikolas S Williams, Genevieve M McArthur, Bianca de Wit, George Ibrahim, and Nicholas A Badcock. A validation of emotiv epoc flex saline for eeg and erp research. PeerJ, 8:e9713, 2020. 14 A Appendix A.1 Demographics Table 2: Participant Demographics by Category Variable Catego...

  60. [2023]

    arXiv:2306.00927 [cs, q-bio]

  61. [2024]

    (extended version with additional analyses)

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.