Pith. sign in

REVIEW 3 cited by

Natural scene reconstruction from fMRI signals using generative latent diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.05334 v2 pith:ZQNT4BQH submitted 2023-03-09 cs.CV cs.AIq-bio.NC

classification cs.CVcs.AIq-bio.NC
keywords imagesdiffusionfmrilatentmodelnaturalpropertiesreconstruction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In neural decoding research, one of the most intriguing topics is the reconstruction of perceived natural images based on fMRI signals. Previous studies have succeeded in re-creating different aspects of the visuals, such as low-level properties (shape, texture, layout) or high-level features (category of objects, descriptive semantics of scenes) but have typically failed to reconstruct these properties together for complex scene images. Generative AI has recently made a leap forward with latent diffusion models capable of generating high-complexity images. Here, we investigate how to take advantage of this innovative technology for brain decoding. We present a two-stage scene reconstruction framework called ``Brain-Diffuser''. In the first stage, starting from fMRI signals, we reconstruct images that capture low-level properties and overall layout using a VDVAE (Very Deep Variational Autoencoder) model. In the second stage, we use the image-to-image framework of a latent diffusion model (Versatile Diffusion) conditioned on predicted multimodal (text and visual) features, to generate final reconstructed images. On the publicly available Natural Scenes Dataset benchmark, our method outperforms previous models both qualitatively and quantitatively. When applied to synthetic fMRI patterns generated from individual ROI (region-of-interest) masks, our trained model creates compelling ``ROI-optimal'' scenes consistent with neuroscientific knowledge. Thus, the proposed methodology can have an impact on both applied (e.g. brain-computer interface) and fundamental neuroscience.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimizing fMRI Data Acquisition for Decoding Natural Speech with Limited Participants

    q-bio.NC 2025-05 conditional novelty 6.0 of 10

    In a small cohort, fMRI decoders improve with more data per participant, and multi-subject training or shared stimuli add no benefit, so deep phenotyping is the recommended acquisition strategy.

  2. VoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI

    cs.CV 2025-09 reject novelty 5.0 of 10

    A 39M-parameter transformer with token merging and query compression gives around 74% top-1 fMRI-to-image retrieval across six training subjects, far below the ~98% of larger state-of-the-art decoders.

  3. Perception Activator: An intuitive and portable framework for brain cognitive exploration

    cs.CV 2025-07 reject novelty 4.0 of 10

    Injecting fMRI vectors into Mask R-CNN via cross-attention produces a small detection AP gain and a slight segmentation AP drop on NSD, contradicting the abstract's claim of improved segmentation accuracy.

Pith tools