REVIEW 12 cited by
DreamDiffusion: Generating High-Quality Images from Brain EEG Signals
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper introduces DreamDiffusion, a novel method for generating high-quality images directly from brain electroencephalogram (EEG) signals, without the need to translate thoughts into text. DreamDiffusion leverages pre-trained text-to-image models and employs temporal masked signal modeling to pre-train the EEG encoder for effective and robust EEG representations. Additionally, the method further leverages the CLIP image encoder to provide extra supervision to better align EEG, text, and image embeddings with limited EEG-image pairs. Overall, the proposed method overcomes the challenges of using EEG signals for image generation, such as noise, limited information, and individual differences, and achieves promising results. Quantitative and qualitative results demonstrate the effectiveness of the proposed method as a significant step towards portable and low-cost ``thoughts-to-image'', with potential applications in neuroscience and computer vision. The code is available here \url{https://github.com/bbaaii/DreamDiffusion}.
Forward citations
Cited by 12 Pith papers
-
3D-Telepathy: Reconstructing 3D Objects from EEG Signals
3D-Telepathy reconstructs 3D objects from EEG signals by combining a dual self-attention EEG encoder with stable diffusion and variational score distillation into a NeRF, and reports best 2D-frame metrics among compar...
-
Category-aware EEG image generation based on wavelet transform and contrast semantic loss
A DWT-gated transformer EEG encoder with CLIP alignment and category-aware clustering loss generates semantic images via a pre-trained diffusion model, achieving 43% max single-subject top-1 classification accuracy an...
-
MoTime: A Dataset Suite for Multimodal Time Series Forecasting
MoTime provides a large multimodal forecasting benchmark and shows that external text or images can improve forecasts in some datasets, especially cold-start and sparse settings, though gains are inconsistent.
-
Neuro-3D: Towards 3D Visual Decoding from EEG Signals
A first benchmark for decoding 3D object shape and dominant color from EEG, with above-chance but limited reconstruction performance.
-
Towards Neural Foundation Models for Vision: Aligning EEG, MEG, and fMRI Representations for Decoding, Encoding, and Modality Conversion
A contrastive model aligns EEG, MEG, and fMRI activity to CLIP image embeddings, enabling image retrieval from brain signals, neural retrieval from images, and cross-modal neural retrieval.
-
What Does the Brain See? Multiview Neural Representations to Demystify the Brain-Visual Alignment
Joint temporal-spectral-spatial EEG encoding with contrastive CLIP alignment sets new SOTA on THINGS-EEG zero-shot decoding, including the first systematic cross-session results.
-
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
WorldWeaver reduces temporal drift in long-horizon video generation by jointly modeling RGB and depth perceptual conditions with segmented noise scheduling.
-
Mind2Matter: Creating 3D Models from EEG Signals
Mind2Matter decodes EEG signals into text descriptions and then uses those descriptions to generate 3D scenes with Gaussian splatting, demonstrating an EEG-to-3D pipeline.
-
Decoding Visual Neural Representations by Multimodal with Dynamic Balancing
A multimodal EEG-image-text contrastive framework with dynamic gradient balancing and stochastic noise improves zero-shot object recognition from EEG on ThingsEEG, raising top-1 accuracy from 13.8% to 15.8%.
-
Foundation Models for Cross-Domain EEG Analysis Application: A Survey
A survey that organizes EEG foundation-model research into five output-modality categories: native EEG, text, vision, audio, and multimodal fusion, with a claim to be the first such comprehensive taxonomy.
-
CATVis: Context-Aware Thought Visualization
CATVis combines a Conformer EEG classifier, CLIP-based caption retrieval and re-ranking, and Stable Diffusion to generate images from EEG, reporting large gains over prior work.
-
Dynamic Neural Communication: Convergence of Computer Vision and Brain-Computer Interface
A diffusion-based model classifies short EEG/EMG segments into 15 viseme classes, and an LSTM trained on the test sentences themselves reconstructs the spoken sentences.
Discussion (0). Continue with ORCID to comment.