REVIEW 1 major objections 1 minor 2 cited by
Neuro-3D: Towards 3D Visual Decoding from EEG Signals
T0 review · 1 major / 1 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Neuro-3D claims to be the first framework to decode 3D visual perception from EEG signals, reconstructing colored point clouds from 12 subjects' brain activity across 72 object categories.
desk verdict A genuinely useful first EEG-3D benchmark and a competent pipeline, but the 'high fidelity instance-specific reconstruction' claim needs a category-prototype baseline before it can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the Dynamic-Static EEG-Fusion Encoder followed by a decoupled colored point cloud decoder. The encoder uses temporal self-attention to embed static and dynamic EEG signals, then an attention-based aggregator that treats the static embedding as the query and the dynamic embedding as the key-value pair, adaptively blending the stable single-view response with the richer rotating-video response. The fused representation is split by separate MLP projections into geometry and appearance features, which are aligned to CLIP video features through contrastive and MSE losses and supervised by shape and color classification losses. The geometry feature conditions a point-voxel diffusion model that generates an 8192-point shape, and the appearance feature conditions a single-step coloring model that assigns dominant colors via majority voting.
What would settle it
Train a control model that maps EEG (or even no EEG) to the per-category mean point cloud and the most frequent dominant color, and evaluate it with the same Chamfer distance, F1, and N-way metrics on the EEG-3D test objects. If this category-prototype baseline matches Neuro-3D's reported numbers (Chamfer distance $5.35\times 10^{-2}$, F1 77.01 percent, 2-way top-1 55.81 percent), then the reconstruction scores do not prove instance-specific 3D decoding.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that a two-stage diffusion pipeline conditioned on EEG embeddings can reconstruct colored point clouds from brain signals at above-chance semantic fidelity, and that fusing static and dynamic EEG responses via an attention aggregator improves both shape and color recovery. The authors present this as the first demonstration of EEG-based 3D visual decoding, extending prior fMRI-based 3D reconstruction work while adding a benchmark dataset that pairs EEG with 3D shapes, videos, images, text captions, and color labels. The quantitative evidence includes 72-way EEG classification at 5.91 percent top-1 accuracy, reconstruction metrics of Chamfer distance $5.35\times 10^{-2}$, F1 score 77.01 percent, and 2-way top-1 accuracy of 55.81 percent averaged over five diffusion samples.
Load-bearing premise
The load-bearing premise is that the reported reconstruction metrics measure how close the generated point cloud is to the specific object the person viewed, rather than how typical it is of the object's category; the paper does not compare against a model that always outputs a per-category average shape.
Editorial extensions
If this is right
- If the results are correct, EEG-based decoding can be extended from 2D images to 3D objects, giving neuroscience a non-invasive tool to probe real-time 3D perception.
- The EEG-3D dataset becomes a benchmark for training and comparing EEG-driven 3D reconstruction models, filling the stated gap of paired EEG and 3D stimulus data.
- Fusing static and dynamic EEG responses improves both classification and reconstruction over either signal alone, indicating complementary neural codes for stable and motion-derived 3D information.
- Decoupling geometry from appearance features improves reconstruction, supporting the view that shape and color are separable at the level of decodable neural representations.
- The brain-region analyses show occipital and temporal electrodes matter most, aligning with known visual processing pathways and suggesting the reconstructions reflect genuine visual processing.
Reading between the lines
- Editor's inference: if instance-level decoding is confirmed, EEG could enable real-time closed-loop experiments where a participant sees an object and receives its reconstruction within seconds, a capability fMRI's temporal resolution cannot support.
- Editor's inference: the color modeling reduces object color to a few dominant colors by majority voting, so the reported color fidelity is category-level palette matching, not per-point texture; testing on objects with fine-grained texture would clarify how much appearance information EEG actually carries.
- Editor's inference: because each category contributes only 8 training and 2 test instances and the 72-way EEG classification top-1 is only 5.91 percent, the reconstruction scores could in part reflect category-level prototype generation; a decoding test on object categories unseen during training would separate category information from instance-specific 3D detail.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a new task, 3D visual decoding from EEG signals, along with a new dataset, EEG-3D, containing EEG recordings from 12 subjects viewing 72 categories of 3D objects rendered as both rotating videos and static images. The authors also propose Neuro-3D, a framework that fuses static and dynamic EEG features through an attention-based aggregator, then uses a diffusion-based colored point cloud decoder to reconstruct both shape and color. The paper reports classification results well above chance for object and color categories, qualitative reconstruction examples, and quantitative reconstruction metrics (Chamfer distance, F1 score, and N-way top-K accuracy).
Significance. The EEG-3D dataset is a potentially valuable resource: it is the first EEG dataset paired with 3D object stimuli, it includes dynamic and static conditions plus resting state, and it provides multimodal analysis data (videos, images, text, 3D shapes, color labels). If the reconstruction claim were fully supported, the work would open a new direction in real-time EEG-based 3D decoding. However, as presented, the evidence primarily supports category-level information being decodable from EEG; the instance-specific 'high fidelity' reconstruction claim in the abstract is not sufficiently established. The authors are commended for releasing code and data, and for including a brain-region analysis that aligns with known visual pathways.
major comments (1)
- [Section 5.3.1, Table 4; Supplementary Section 8] The color generation is reduced to a majority-voting dominant-color label (Section 4.3), yet the abstract claims reconstruction of 'colored 3D objects with high fidelity.' There is no quantitative evaluation of color accuracy beyond the six-way color-type classification in Table 2 and qualitative inspection. I request a quantitative color metric, such as dominant-color prediction accuracy on the reconstructed point clouds, or per-point color error against ground truth, so that the color claim can be assessed. If such a metric is not feasible, the abstract and conclusion should be tempered to 'dominant color style' rather than high-fidelity color reconstruction.
minor comments (1)
- [Section 3.4] In Section 5.2.1, the claim that 'all methods exceed chance-level performance by a significant margin' is not backed by statistical tests; please add significance testing or rephrase to avoid implying formal significance.
Circularity Check
Score 2/10: the EEG-to-3D reconstruction chain is evaluated with external, held-out benchmarks (Objaverse-trained PointNet++, ground-truth CD/F1) and no training loss equals an evaluation metric; only a minor, non-load-bearing self-citation ([65], [66] by co-author Yonghao Song) keeps the score above zero.
-
self citation load bearing
[Sec. 5.2.1 (Comparison with Related Methods) and Sec. 5.4 (Analysis of Brain Regions); references [65] and [66]]
"We re-implement several state-of-the-art EEG encoders [35, 60, 65, 66] for comparative analysis by training separate object and color classifiers. ... This finding aligns with the previous neuroscience discoveries regarding the brain's visual processing mechanisms [19, 36, 66]."
Two cited baselines, [65] EEG Conformer and [66] TSConv (ICLR 2024), are first-authored by Yonghao Song, a co-author of the present paper, and [66] is additionally cited in Sec. 5.4 as neuroscience support for the occipital/temporal electrode finding. The overlap is minor and non-load-bearing: the methods are re-implemented and benchmarked on the new EEG-3D dataset (Tab. 2), so they are externally falsifiable rather than assumed evidence, and the brain-region result is independently anchored by refs [19], [3], and [11]. No equation or reported number in the 3D decoding pipeline depends on these citations, so under the proportional rubric this single minor self-citation sets the score at 2 rather than 0.
full rationale
The derivation chain is EEG signals to a dynamic-static fusion encoder to decoupled geometry/appearance features fg and fa to diffusion-based shape generation plus a single-step color model (Secs. 4.2-4.3). The reconstruction benchmarks are external and held out: the N-way top-K metric uses a PointNet++ classifier pre-trained on Objaverse data with the paper's test instances excluded (Supp. Sec. 8), and Chamfer distance and F1 compare generated clouds directly to the ground-truth stimulus clouds (Sec. 5.3.1); none of these metrics is computed from the model's own fitted parameters or training losses. The training objectives are CLIP alignment to frozen video features (Eq. 4), category cross-entropy on EEG features (Eq. 6), and the diffusion reconstruction objective (Eq. 12), none of which equals an evaluation metric, and the evaluation classifier is never used during training. The nearest concern is that Eq. 6 explicitly trains fg to carry shape-category information while the N-way metric measures category-level semantic fidelity; but since the metric is external and average scores are only a few points above chance (55.81% vs 50% for 2-way top-1), no result is forced by construction, so the abstract's stronger 'high fidelity, instance-specific reconstruction' claim remains unproven without a category-prototype baseline - a correctness risk, not a circular step. The best-of-5 sample selection in Supp. Sec. 8 ('identify the optimal result based on the classifier's predicted scores') inflates the reported maximum but is transparently labeled and does not equate the prediction with its input. The authors themselves acknowledge the color simplification to dominant-color majority voting (Secs. 4.3 and 6), matching the flagged limitation passage. The only author-overlapping references, [65] and [66], are re-implemented baselines and confirmatory neuroscience context independently anchored by [19], [3], and [11], so they are not load-bearing. Overall, no self-definitional equation, no fitted-parameter-as-prediction, no uniqueness-import, no ansatz-smuggled-via-citation, and no renamed-known-result step were found; score 2 reflects only the minor self-citation.
Assumptions & free parameters
free parameters (5)
- loss_coefficient_alpha =
0.01
- loss_coefficient_gamma =
0.1
- num_video_frames_n =
4
- feature_dimension =
1024
- point_cloud_size_N =
8192
assumptions (3)
- domain assumption EEG signals from viewing rotating 3D videos and static images encode information about 3D object shape and color that a learned model can extract.
- domain assumption The mean of CLIP image embeddings over 4 video frames is a sufficient target for aligning EEG features to both geometry and appearance.
- domain assumption A PointNet++ classifier trained on Objaverse categories provides a valid fidelity measure for generated point clouds.
Cite this review
Pith. "Pith review of Neuro-3D: Towards 3D Visual Decoding from EEG Signals." pith.science (2026). https://pith.science/paper/N2CT63OQ
@misc{pith2026241112248,
author = {Pith},
title = {Pith review of: Neuro-3D: Towards 3D Visual Decoding from EEG Signals},
year = {2026},
howpublished = {\url{https://pith.science/paper/N2CT63OQ}},
note = {Machine review of arXiv:2411.12248}
}
read the original abstract
Human's perception of the visual world is shaped by the stereo processing of 3D information. Understanding how the brain perceives and processes 3D visual stimuli in the real world has been a longstanding endeavor in neuroscience. Towards this goal, we introduce a new neuroscience task: decoding 3D visual perception from EEG signals, a neuroimaging technique that enables real-time monitoring of neural dynamics enriched with complex visual cues. To provide the essential benchmark, we first present EEG-3D, a pioneering dataset featuring multimodal analysis data and extensive EEG recordings from 12 subjects viewing 72 categories of 3D objects rendered in both videos and images. Furthermore, we propose Neuro-3D, a 3D visual decoding framework based on EEG signals. This framework adaptively integrates EEG features derived from static and dynamic stimuli to learn complementary and robust neural representations, which are subsequently utilized to recover both the shape and color of 3D objects through the proposed diffusion-based colored point cloud decoder. To the best of our knowledge, we are the first to explore EEG-based 3D visual decoding. Experiments indicate that Neuro-3D not only reconstructs colored 3D objects with high fidelity, but also learns effective neural representations that enable insightful brain region analysis. The dataset and associated code will be made publicly available.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 2 Pith papers
-
3D-Telepathy: Reconstructing 3D Objects from EEG Signals
3D-Telepathy reconstructs 3D objects from EEG signals by combining a dual self-attention EEG encoder with stable diffusion and variational score distillation into a NeRF, and reports best 2D-frame metrics among compar...
-
Multimodal Brain-Computer Interfaces: AI-powered Decoding Methodologies
This review organizes multimodal brain-computer interface decoding into three algorithmic task types and surveys AI methods for visual, speech, and affective decoding.
Reference graph
Works this paper leans on
-
[1]
A massive 7t fmri dataset to bridge cognitive neuroscience and artificial intelligence
Emily J Allen, Ghislain St-Yves, Yihan Wu, Jesse L Breedlove, Jacob S Prince, Logan T Dowdle, Matthias Nau, Brad Caron, Franco Pestilli, Ian Charest, et al. A massive 7t fmri dataset to bridge cognitive neuroscience and artificial intelligence. Nature Neuroscience, 25(1):116–126, 2022. 2, 4
2022
-
[2]
Dreamdiffusion: Generating high- quality images from brain eeg signals
Yunpeng Bai, Xintao Wang, Yan-pei Cao, Yixiao Ge, Chun Yuan, and Ying Shan. Dreamdiffusion: Generating high- quality images from brain eeg signals. arXiv preprint arXiv:2306.16934, 2023. 2
arXiv 2023
-
[3]
A map of object space in primate inferotemporal cortex
Pinglei Bao, Liang She, Mason McGill, and Doris Y Tsao. A map of object space in primate inferotemporal cortex. Na- ture, 583(7814):103–108, 2020. 8
2020
-
[4]
From voxels to pixels and back: Self-supervision in natural-image reconstruction from fMRI
Roman Beliy, Guy Gaziv, Assaf Hoogi, Francesca Strappini, Tal Golan, and Michal Irani. From voxels to pixels and back: Self-supervision in natural-image reconstruction from fMRI. Advances in Neural Information Processing Systems, 32, 2019. 2
2019
-
[5]
BOLD5000, a public fMRI dataset while viewing 5000 visual images
Nadine Chang, John A Pyles, Austin Marcus, Abhinav Gupta, Michael J Tarr, and Elissa M Aminoff. BOLD5000, a public fMRI dataset while viewing 5000 visual images. Sci- entific Data, 6(1):49, 2019. 2, 4
2019
-
[6]
Representation of vestibular and visual cues to self-motion in ventral intraparietal cortex
Aihua Chen, Gregory C DeAngelis, and Dora E Angelaki. Representation of vestibular and visual cues to self-motion in ventral intraparietal cortex. Journal of Neuroscience, 31 (33):12036–12052, 2011. 8
2011
-
[7]
Eegformer: To- wards transferable and interpretable large-scale eeg founda- tion model
Yuqi Chen, Kan Ren, Kaitao Song, Yansen Wang, Yi- fan Wang, Dongsheng Li, and Lili Qiu. Eegformer: To- wards transferable and interpretable large-scale eeg founda- tion model. arXiv preprint arXiv:2401.10278, 2024. 1, 2
arXiv 2024
-
[8]
Seeing beyond the brain: Masked model- ing conditioned diffusion model for human vision decoding
Zijiao Chen, Jiaxin Qing, Tiange Xiang, Wan Lin Yue, and Juan Helen Zhou. Seeing beyond the brain: Masked model- ing conditioned diffusion model for human vision decoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 1, 2, 6
2023
Show all 88 references
-
[9]
Cinematic mindscapes: High-quality video reconstruction from brain activity
Zijiao Chen, Jiaxin Qing, and Juan Helen Zhou. Cinematic mindscapes: High-quality video reconstruction from brain activity. Advances in Neural Information Processing Sys- tems, 36, 2024. 2, 3, 6, 1
2024
-
[10]
Objaverse: A universe of annotated 3D objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3D objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...
2023
-
[11]
Untangling invariant object recognition
James J DiCarlo and David D Cox. Untangling invariant object recognition. Trends in Cognitive Sciences, 11(8):333– 341, 2007. 8
2007
-
[12]
fMRI of human visual cortex
Stephen A Engel, David E Rumelhart, Brian A Wandell, Adrian T Lee, Gary H Glover, Eduardo-Jose Chichilnisky, Michael N Shadlen, et al. fMRI of human visual cortex. Na- ture, 369(6481):525–525, 1994. 2
1994
-
[13]
fmri-3d: A comprehensive dataset for enhancing fmri-based 3d reconstruction
Jianxiong Gao, Yuqian Fu, Yun Wang, Xuelin Qian, Jianfeng Feng, and Yanwei Fu. fmri-3d: A comprehensive dataset for enhancing fmri-based 3d reconstruction. arXiv preprint arXiv:2409.11315, 2024. 3
2024 arXiv
-
[14]
Mind-3D: Reconstruct high- quality 3D objects in human brain
Jianxiong Gao, Yuqian Fu, Yun Wang, Xuelin Qian, Jian- feng Feng, and Yanwei Fu. Mind-3D: Reconstruct high- quality 3D objects in human brain. In European Conference on Computer Vision, 2024. 2, 3, 4
2024
-
[15]
Eeg variability: Task-driven or subject- driven signal of interest? NeuroImage, 252:119034, 2022
Erin Gibson, Nancy J Lobaugh, Steve Joordens, and An- thony R McIntosh. Eeg variability: Task-driven or subject- driven signal of interest? NeuroImage, 252:119034, 2022. 1
2022
-
[16]
A large and rich EEG dataset for mod- eling human visual object recognition
Alessandro T Gifford, Kshitij Dwivedi, Gemma Roig, and Radoslaw M Cichy. A large and rich EEG dataset for mod- eling human visual object recognition. NeuroImage, 264: 119754, 2022. 2, 4
2022
-
[17]
MEG and EEG data analysis with MNE-Python
Alexandre Gramfort, Martin Luessi, Eric Larson, Denis A Engemann, Daniel Strohmeier, Christian Brodbeck, Roman Goj, Mainak Jas, Teon Brooks, Lauri Parkkonen, et al. MEG and EEG data analysis with MNE-Python. Frontiers in Neu- roinformatics, 7:267, 2013. 4, 1
2013
-
[18]
The human vi- sual cortex
Kalanit Grill-Spector and Rafael Malach. The human vi- sual cortex. Annual Review of Neuroscience, 27(1):649–677,
-
[19]
The lateral occipital complex and its role in object recogni- tion
Kalanit Grill-Spector, Zoe Kourtzi, and Nancy Kanwisher. The lateral occipital complex and its role in object recogni- tion. Vision Research, 41(10-11):1409–1422, 2001. 8
2001
-
[20]
Human EEG recordings for 1,854 concepts presented in rapid serial visual presentation streams
Tijl Grootswagers, Ivy Zhou, Amanda K Robinson, Martin N Hebart, and Thomas A Carlson. Human EEG recordings for 1,854 concepts presented in rapid serial visual presentation streams. Scientific Data, 9(1):3, 2022. 2, 4
2022
-
[21]
Causal links between dorsal medial superior temporal area neurons and multisensory heading perception
Yong Gu, Gregory C DeAngelis, and Dora E Angelaki. Causal links between dorsal medial superior temporal area neurons and multisensory heading perception. Journal of Neuroscience, 32(7):2299–2313, 2012. 8
2012
-
[22]
Multivariate pattern analysis for meg: A comparison of dissimilarity measures
Matthias Guggenmos, Philipp Sterzer, and Radoslaw Martin Cichy. Multivariate pattern analysis for meg: A comparison of dissimilarity measures. NeuroImage, 173:434–447, 2018. 1
2018
-
[23]
The organization of behavior: A neu- ropsychological theory
Donald Olding Hebb. The organization of behavior: A neu- ropsychological theory. Psychology press, 2005. 2
2005
-
[24]
The perception of vi- sual information
William R Hendee and Peter NT Wells. The perception of vi- sual information. Springer Science & Business Media, 1997. 1
1997
-
[25]
Denoising diffu- sion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020. 2, 3
2020
-
[26]
Generic de- coding of seen and imagined objects using hierarchical vi- sual features
Tomoyasu Horikawa and Yukiyasu Kamitani. Generic de- coding of seen and imagined objects using hierarchical vi- sual features. Nature Communications, 8(1):15037, 2017. 2, 4
2017
-
[27]
Discrepancy between inter-and intra-subject variability in eeg-based motor imagery brain- computer interface: Evidence from multiple perspectives
Gan Huang, Zhiheng Zhao, Shaorong Zhang, Zhenxing Hu, Jiaming Fan, Meisong Fu, Jiale Chen, Yaqiong Xiao, Jun Wang, and Guo Dan. Discrepancy between inter-and intra-subject variability in eeg-based motor imagery brain- computer interface: Evidence from multiple perspectives. Fr...
2023
-
[28]
fmri-based decoding of visual information from hu- man brain activity: A brief review
Shuo Huang, Wei Shao, Mei-Ling Wang, and Dao-Qiang Zhang. fmri-based decoding of visual information from hu- man brain activity: A brief review. Machine Intelligence Research, 18(2):170–184, 2021. 2 9
2021
-
[29]
Large brain model for learning generic representations with tremendous eeg data in bci
Wei-Bang Jiang, Li-Ming Zhao, and Bao-Liang Lu. Large brain model for learning generic representations with tremendous eeg data in bci. In International Conference on Learning Representations, 2024. 1
2024
-
[30]
Brain2image: Con- verting brain signals into images
Isaak Kavasidis, Simone Palazzo, Concetto Spampinato, Daniela Giordano, and Mubarak Shah. Brain2image: Con- verting brain signals into images. In Proceedings of the 25th ACM International Conference on Multimedia, pages 1809– 1817, 2017. 2, 4
2017
-
[31]
Dif- fusionclip: Text-guided diffusion models for robust image manipulation
Gwanghyun Kim, Taesung Kwon, and Jong Chul Ye. Dif- fusionclip: Text-guided diffusion models for robust image manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2426– 2435, 2022. 3
2022
-
[32]
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 ,
-
[33]
A penny for your (visual) thoughts: Self-supervised reconstruction of natural movies from brain activity
Ganit Kupershmidt, Roman Beliy, Guy Gaziv, and Michal Irani. A penny for your (visual) thoughts: Self-supervised reconstruction of natural movies from brain activity. arXiv preprint arXiv:2206.03544, 2022. 3
2022 arXiv
-
[34]
Modeling short visual events through the bold moments video fmri dataset and metadata
Benjamin Lahner, Kshitij Dwivedi, Polina Iamshchinina, Monika Graumann, Alex Lascelles, Gemma Roig, Alessan- dro Thomas Gifford, Bowen Pan, SouYoung Jin, N Apurva Ratan Murty, et al. Modeling short visual events through the bold moments video fmri dataset and metadata. Nature ...
2024
-
[35]
EEG- Net: a compact convolutional neural network for EEG-based brain–computer interfaces
Vernon J Lawhern, Amelia J Solon, Nicholas R Waytowich, Stephen M Gordon, Chou P Hung, and Brent J Lance. EEG- Net: a compact convolutional neural network for EEG-based brain–computer interfaces. Journal of Neural Engineering , 15(5):056013, 2018. 6, 7
2018
-
[36]
Visual decoding and reconstruction via EEG embeddings with guided diffusion
Dongyang Li, Chen Wei, Shiying Li, Jiachen Zou, and Quanying Liu. Visual decoding and reconstruction via EEG embeddings with guided diffusion. In Advances in Neural Information Processing Systems, 2024. 2, 6, 8, 1
2024
-
[37]
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In In- ternational Conference on Machine Learning, pages 19730– 19742. PMLR, 2023. 2
2023
-
[38]
Mind reader: Reconstructing complex images from brain activi- ties
Sikun Lin, Thomas Sprague, and Ambuj K Singh. Mind reader: Reconstructing complex images from brain activi- ties. Advances in Neural Information Processing Systems , 35:29624–29636, 2022. 2
2022
-
[39]
Zero-1-to-3: Zero-shot one image to 3D object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok- makov, Sergey Zakharov, and Carl V ondrick. Zero-1-to-3: Zero-shot one image to 3D object. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 9298–9309, 2023. 3
2023
-
[40]
Point- voxel cnn for efficient 3d deep learning
Zhijian Liu, Haotian Tang, Yujun Lin, and Song Han. Point- voxel cnn for efficient 3d deep learning. Advances in Neural Information Processing Systems, 32, 2019. 6
2019
-
[41]
Wonder3D: Sin- gle image to 3D using cross-domain diffusion
Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, et al. Wonder3D: Sin- gle image to 3D using cross-domain diffusion. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pa...
2024
-
[42]
Brain diffusion for visual exploration: Cortical discov- ery using large scale generative models
Andrew Luo, Maggie Henderson, Leila Wehbe, and Michael Tarr. Brain diffusion for visual exploration: Cortical discov- ery using large scale generative models. Advances in Neural Information Processing Systems, 36, 2024. 1
2024
-
[43]
Diffusion probabilistic models for 3D point cloud generation
Shitong Luo and Wei Hu. Diffusion probabilistic models for 3D point cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2837–2845, 2021. 3
2021
-
[44]
Pc2: Projection-conditioned point cloud diffu- sion for single-image 3D reconstruction
Luke Melas-Kyriazi, Christian Rupprecht, and Andrea Vedaldi. Pc2: Projection-conditioned point cloud diffu- sion for single-image 3D reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12923–12932, 2023. 3, 6
2023
-
[45]
A high-performance neuroprosthesis for speech decod- ing and avatar control
Sean L Metzger, Kaylo T Littlejohn, Alexander B Silva, David A Moses, Margaret P Seaton, Ran Wang, Maximil- ian E Dougherty, Jessie R Liu, Peter Wu, Michael A Berger, et al. A high-performance neuroprosthesis for speech decod- ing and avatar control. Nature, 620(7976):1037–104...
2023
-
[46]
Neuroprosthesis for decoding speech in a paralyzed person with anarthria
David A Moses, Sean L Metzger, Jessie R Liu, Gopala K Anumanchipalli, Joseph G Makin, Pengfei F Sun, Josh Chartier, Maximilian E Dougherty, Patricia M Liu, Gary M Abrams, et al. Neuroprosthesis for decoding speech in a paralyzed person with anarthria. New England Journal of Me...
2021
-
[47]
Point-e: A system for generat- ing 3D point clouds from complex prompts
Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. Point-e: A system for generat- ing 3D point clouds from complex prompts. arXiv preprint arXiv:2212.08751, 2022. 3
2022 arXiv
-
[48]
Improved denoising diffusion probabilistic models
Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, pages 8162–8171. PMLR,
-
[49]
Reconstruction of perceived images from fMRI patterns and semantic brain exploration using instance-conditioned GANs
Furkan Ozcelik, Bhavin Choksi, Milad Mozafari, Leila Reddy, and Rufin VanRullen. Reconstruction of perceived images from fMRI patterns and semantic brain exploration using instance-conditioned GANs. In 2022 International Joint Conference on Neural Networks , pages 1–8. IEEE,
2022
-
[50]
Dreamfusion: Text-to-3D using 2D diffusion
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. Dreamfusion: Text-to-3D using 2D diffusion. arXiv preprint arXiv:2209.14988, 2022. 3
2022 arXiv
-
[51]
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in Neural Information Processing Systems, 30, 2017. 1
2017
-
[52]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In International Conference on Machine Learning...
2021
-
[53]
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Confer- 10 ence on Machine Learning, pages 8821–8831. PMLR, 2021. 3
2021
-
[54]
TIGER: Time-varying denoising model for 3D point cloud generation with diffusion process
Zhiyuan Ren, Minchul Kim, Feng Liu, and Xiaoming Liu. TIGER: Time-varying denoising model for 3D point cloud generation with diffusion process. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9462–9471, 2024. 3
2024
-
[55]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022. 3
2022
-
[56]
Intra-and inter-subject variability in eeg-based sensorimotor brain computer inter- face: a review
Simanto Saha and Mathias Baumert. Intra-and inter-subject variability in eeg-based sensorimotor brain computer inter- face: a review. Frontiers in Computational Neuroscience , 13:87, 2020. 1
2020
-
[57]
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information...
2022
-
[58]
Microstimulation in visual area mt: effects on direction discrimination performance
C Daniel Salzman, Chieko M Murasugi, Kenneth H Britten, and William T Newsome. Microstimulation in visual area mt: effects on direction discrimination performance. Journal of Neuroscience, 12(6):2331–2355, 1992. 8
1992
-
[59]
Zeronvs: Zero-shot 360- degree view synthesis from a single real image
Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Herrmann, Hong-Xing Yu, Yunzhi Zhang, Eric Ryan Chan, Dmitry La- gun, Li Fei-Fei, Deqing Sun, et al. Zeronvs: Zero-shot 360- degree view synthesis from a single real image. In Proceed- ings of the IEEE/CVF Conference on Computer V...
2024
-
[60]
Deep learning with convolutional neural networks for eeg decoding and visualization
Robin Tibor Schirrmeister, Jost Tobias Springenberg, Lukas Dominique Josef Fiederer, Martin Glasstetter, Katharina Eggensperger, Michael Tangermann, Frank Hutter, Wolfram Burgard, and Tonio Ball. Deep learning with convolutional neural networks for eeg decoding and visualizati...
2017
-
[61]
Re- constructing the mind’s eye: fMRI-to-image with contrastive learning and diffusion priors
Paul Scotti, Atmadeep Banerjee, Jimmie Goode, Stepan Sha- balin, Alex Nguyen, Aidan Dempster, Nathalie Verlinde, Elad Yundler, David Weisberg, Kenneth Norman, et al. Re- constructing the mind’s eye: fMRI-to-image with contrastive learning and diffusion priors. Advances in Neur...
2024
-
[62]
Mindeye2: Shared-subject models enable fMRI-to-image with 1 hour of data
Paul S Scotti, Mihir Tripathy, Cesar Kadir Torrico Vil- lanueva, Reese Kneeland, Tong Chen, Ashutosh Narang, Charan Santhirasegaran, Jonathan Xu, Thomas Naselaris, Kenneth A Norman, et al. Mindeye2: Shared-subject models enable fMRI-to-image with 1 hour of data. In Internation...
2024
-
[63]
Deep image reconstruction from hu- man brain activity
Guohua Shen, Tomoyasu Horikawa, Kei Majima, and Yukiyasu Kamitani. Deep image reconstruction from hu- man brain activity. PLoS Computational Biology , 15(1): e1006633, 2019. 2
2019
-
[64]
EEG2IMAGE: image reconstruc- tion from EEG brain signals
Prajwal Singh, Pankaj Pandey, Krishna Miyapuram, and Shanmuganathan Raman. EEG2IMAGE: image reconstruc- tion from EEG brain signals. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing, pages 1–5. IEEE, 2023. 2
2023
-
[65]
Eeg conformer: Convolutional transformer for eeg decoding and visualization
Yonghao Song, Qingqing Zheng, Bingchuan Liu, and Xi- aorong Gao. Eeg conformer: Convolutional transformer for eeg decoding and visualization. IEEE Transactions on Neu- ral Systems and Rehabilitation Engineering , 31:710–719,
-
[66]
Decoding natural images from EEG for object recognition
Yonghao Song, Bingchuan Liu, Xiang Li, Nanlin Shi, Yijun Wang, and Xiaorong Gao. Decoding natural images from EEG for object recognition. In The Twelfth International Conference on Learning Representations, 2024. 2, 6, 7, 8
2024
-
[67]
Neurocine: Decoding vivid video se- quences from human brain activties
Jingyuan Sun, Mingxiao Li, Zijiao Chen, and Marie- Francine Moens. Neurocine: Decoding vivid video se- quences from human brain activties. arXiv preprint arXiv:2402.01590, 2024. 3
2024 arXiv
-
[68]
Contrast, attend and diffuse to decode high-resolution images from brain ac- tivities
Jingyuan Sun, Mingxiao Li, Zijiao Chen, Yunhao Zhang, Shaonan Wang, and Marie-Francine Moens. Contrast, attend and diffuse to decode high-resolution images from brain ac- tivities. Advances in Neural Information Processing Systems, 36, 2024. 2
2024
-
[69]
Event-related poten- tial: An overview
Shravani Sur and Vinod Kumar Sinha. Event-related poten- tial: An overview. Industrial Psychiatry Journal, 18(1):70– 73, 2009. 1
2009
-
[70]
High-resolution image re- construction with latent diffusion models from human brain activity
Yu Takagi and Shinji Nishimoto. High-resolution image re- construction with latent diffusion models from human brain activity. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 14453– 14463, 2023. 2
2023
-
[71]
Lion: Latent point dif- fusion models for 3D shape generation
Arash Vahdat, Francis Williams, Zan Gojcic, Or Litany, Sanja Fidler, Karsten Kreis, et al. Lion: Latent point dif- fusion models for 3D shape generation. Advances in Neural Information Processing Systems, 35:10021–10039, 2022. 3
2022
-
[72]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in Neural Information Processing Systems, 2017. 5
2017
-
[73]
Reconstructing rapid natural vision with fMRI- conditional video generative adversarial network
Chong Wang, Hongmei Yan, Wei Huang, Jiyi Li, Yuting Wang, Yun-Shuang Fan, Wei Sheng, Tao Liu, Rong Li, and Huafu Chen. Reconstructing rapid natural vision with fMRI- conditional video generative adversarial network. Cerebral Cortex, 32(20):4502–4511, 2022. 3
2022
-
[74]
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
Tongzhou Wang and Phillip Isola. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In International Conference on Machine Learning, pages 9929–9939. PMLR, 2020. 2
2020
-
[75]
Neural encoding and decod- ing with deep learning for dynamic natural vision
Haiguang Wen, Junxing Shi, Yizhen Zhang, Kun-Han Lu, Ji- ayue Cao, and Zhongming Liu. Neural encoding and decod- ing with deep learning for dynamic natural vision. Cerebral Cortex, 28(12):4136–4160, 2018. 2, 3, 4
2018
-
[76]
Sketch and text guided diffusion model for colored point cloud generation
Zijie Wu, Yaonan Wang, Mingtao Feng, He Xie, and Ajmal Mian. Sketch and text guided diffusion model for colored point cloud generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8929– 8939, 2023. 3, 6
2023
-
[77]
Pointllm: Empowering large language models to understand point clouds
Runsen Xu, Xiaolong Wang, Tai Wang, Yilun Chen, Jiang- miao Pang, and Dahua Lin. Pointllm: Empowering large language models to understand point clouds. European Con- ference on Computer Vision, 2024. 2, 3, 6, 1 11
2024
-
[78]
Learning topology-agnostic eeg representations with geometry-aware modeling
Ke Yi, Yansen Wang, Kan Ren, and Dongsheng Li. Learning topology-agnostic eeg representations with geometry-aware modeling. Advances in Neural Information Processing Sys- tems, 36, 2024. 1, 2
2024
-
[79]
Accelerating text-to-image editing via cache- enabled sparse diffusion inference
Zihao Yu, Haoyang Li, Fangcheng Fu, Xupeng Miao, and Bin Cui. Accelerating text-to-image editing via cache- enabled sparse diffusion inference. In Proceedings of the AAAI Conference on Artificial Intelligence , pages 16605– 16613, 2024. 3
2024
-
[80]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023. 3
2023
-
[81]
Neural decoding of visual information across different neural recording modalities and approaches
Yi-Jun Zhang, Zhao-Fei Yu, Jian K Liu, and Tie-Jun Huang. Neural decoding of visual information across different neural recording modalities and approaches. Machine Intelligence Research, 19(5):350–365, 2022. 2
2022
-
[82]
3D shape generation and completion through point-voxel diffusion
Linqi Zhou, Yilun Du, and Jiajun Wu. 3D shape generation and completion through point-voxel diffusion. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 5826–5835, 2021. 3
2021
-
[83]
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mo- hamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023. 2 12 Neuro-3D: Towards 3D Visual Decoding from EEG Signals Supplementary Material
2023 arXiv
-
[84]
During data acquisition, static 3D im- age and dynamic 3D video stimuli were preceded by a marker to streamline subsequent data processing
EEG Data Preprocessing In this section, we introduce the details of EEG prepro- cessing pipeline. During data acquisition, static 3D im- age and dynamic 3D video stimuli were preceded by a marker to streamline subsequent data processing. The con- tinuous EEG recordings were su...
-
[85]
For 2D image evaluation, a pre-trained ImageNet1K classifier is used to classify both the gener- ated images and their corresponding ground truth images
Evaluation Metrics for Reconstruction Benchmark To assess the quality of the generated outputs, we adopt the N-way, top-K metric, a standard approach in 2D image de- coding [8, 9, 36]. For 2D image evaluation, a pre-trained ImageNet1K classifier is used to classify both the ge...
-
[86]
Analysis of Individual Difference We present the performance variability across individuals on two classification tasks, as illustrated in Fig. 6. On both tasks, individual performance consistently exceeds chance level, demonstrating that EEG signals encode visual per- ception...
-
[87]
More Reconstructed Samples Additional reconstructed results alongside their correspond- ing ground truth point clouds are presented in Fig. 7. The proposed Neuro-3D framework exhibits robust perfor- mance, effectively capturing semantic categories, shape de- tails, and the ove...
-
[88]
8 illustrates representative failure cases, categorized into two principal types: inaccuracies in detailed shape pre- diction and semantic reconstruction errors
Analysis of Failure Cases Fig. 8 illustrates representative failure cases, categorized into two principal types: inaccuracies in detailed shape pre- diction and semantic reconstruction errors. Despite these limitations, certain features of the stimulus objects, includ- ing sha...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.