Pith. sign in

REVIEW 9 cited by

Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion Priors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.18274 v2 pith:6AE7YTPA submitted 2023-05-29 cs.CV cs.AIq-bio.NC

classification cs.CVcs.AIq-bio.NC
keywords mindeyeimagereconstructionbrainretrievalretrievespaceactivity
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present MindEye, a novel fMRI-to-image approach to retrieve and reconstruct viewed images from brain activity. Our model comprises two parallel submodules that are specialized for retrieval (using contrastive learning) and reconstruction (using a diffusion prior). MindEye can map fMRI brain activity to any high dimensional multimodal latent space, like CLIP image space, enabling image reconstruction using generative models that accept embeddings from this latent space. We comprehensively compare our approach with other existing methods, using both qualitative side-by-side comparisons and quantitative evaluations, and show that MindEye achieves state-of-the-art performance in both reconstruction and retrieval tasks. In particular, MindEye can retrieve the exact original image even among highly similar candidates indicating that its brain embeddings retain fine-grained image-specific information. This allows us to accurately retrieve images even from large-scale databases like LAION-5B. We demonstrate through ablations that MindEye's performance improvements over previous methods result from specialized submodules for retrieval and reconstruction, improved training techniques, and training models with orders of magnitude more parameters. Furthermore, we show that MindEye can better preserve low-level image features in the reconstructions by using img2img, with outputs from a separate autoencoder. All code is available on GitHub.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 40 citations worldwide. Full citation record

  1. Real-time Reconstruction of Human Visual Perception from fMRI

    cs.CV 2026-07 conditional novelty 6.0 of 10

    First demonstration that single-trial visual images can be decoded from fMRI in near-real-time (about 10-15 seconds) with roughly one hour of fine-tuning data.

  2. Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Fine-grained text decoded from fMRI with three reward signals improves brain-to-image reconstruction when fused into existing diffusion pipelines.

  3. Optimizing fMRI Data Acquisition for Decoding Natural Speech with Limited Participants

    q-bio.NC 2025-05 conditional novelty 6.0 of 10

    In a small cohort, fMRI decoders improve with more data per participant, and multi-subject training or shared stimuli add no benefit, so deep phenotyping is the recommended acquisition strategy.

  4. Dynadiff: Single-stage Decoding of Images from Continuously Evolving fMRI

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A single-stage, LoRA-finetuned diffusion model decodes seen images directly from continuous BOLD fMRI time series and beats previous pipelines on semantic metrics.

  5. SIM: Surface-based fMRI Analysis for Inter-Subject Multimodal Decoding from Movie-Watching Experiments

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A surface-transformer and tri-modal CLIP model decodes which 3-second movie clip a person watched from 3 seconds of fMRI, generalizing to new people and new clips.

  6. Scaling laws for decoding images from brain activity

    eess.IV 2025-01 conditional novelty 6.0 of 10

    Across four non-invasive brain-imaging devices, image-decoding accuracy grows log-linearly with recorded data per subject and shows little gain from adding subjects.

  7. Towards Neural Foundation Models for Vision: Aligning EEG, MEG, and fMRI Representations for Decoding, Encoding, and Modality Conversion

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A contrastive model aligns EEG, MEG, and fMRI activity to CLIP image embeddings, enabling image retrieval from brain signals, neural retrieval from images, and cross-modal neural retrieval.

  8. VoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI

    cs.CV 2025-09 reject novelty 5.0 of 10

    A 39M-parameter transformer with token merging and query compression gives around 74% top-1 fMRI-to-image retrieval across six training subjects, far below the ~98% of larger state-of-the-art decoders.

  9. Making Your Dreams A Reality: Decoding the Dreams into a Coherent Video Story from fMRI Signals

    cs.CV 2025-01 reject novelty 5.0 of 10

    A zero-shot pipeline claims to generate dream videos from sleep fMRI by transferring a model trained on awake visual perception, with only three dream segments validated.

Pith tools