Pith. sign in

REVIEW 28 cited by

MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.11207 v2 pith:YERLA64Q submitted 2024-03-17 cs.CV cs.AIq-bio.NC

MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of Data

classification cs.CV cs.AIq-bio.NC
keywords dataspaceclipreconstructionssubjecttrainingbrainfmri
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Reconstructions of visual perception from brain activity have improved tremendously, but the practical utility of such methods has been limited. This is because such models are trained independently per subject where each subject requires dozens of hours of expensive fMRI training data to attain high-quality results. The present work showcases high-quality reconstructions using only 1 hour of fMRI training data. We pretrain our model across 7 subjects and then fine-tune on minimal data from a new subject. Our novel functional alignment procedure linearly maps all brain data to a shared-subject latent space, followed by a shared non-linear mapping to CLIP image space. We then map from CLIP space to pixel space by fine-tuning Stable Diffusion XL to accept CLIP latents as inputs instead of text. This approach improves out-of-subject generalization with limited training data and also attains state-of-the-art image retrieval and reconstruction metrics compared to single-subject approaches. MindEye2 demonstrates how accurate reconstructions of perception are possible from a single visit to the MRI facility. All code is available on GitHub.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 28 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Can neurons speak? Semantic narration of vision at single-cell resolution

    q-bio.NC 2026-06 unverdicted novelty 7.0

    NEURRATOR bridges neural spike trains to frozen CLIP patch embeddings via a learned encoder, then uses a multimodal LM and sparse autoencoder to produce validated natural-language narrations of viewed scenes from Neur...

  2. Brain-IT-VQA: From Brain Signals to Answers

    cs.CV 2026-05 unverdicted novelty 7.0

    Brain-IT-VQA decodes visual question answers from fMRI using a transformer to extract language tokens and introduces the NSD-VQA benchmark with 20 controlled questions per image across 20 categories.

  3. From Activation to Causality: Discovery of Causal Visual Representations in the Human Brain

    cs.CV 2026-05 unverdicted novelty 7.0

    BrainCause recovers known visual localizations and finds new candidate representations by validating causal specificity via counterfactual stimuli and encoding models, showing activation alone produces many false positives.

  4. A foundation model of vision, audition, and language for in-silico neuroscience

    q-bio.NC 2026-05 unverdicted novelty 7.0

    TRIBE v2 is a multimodal AI model that predicts human brain activity more accurately than linear encoding models and recovers established neuroscientific findings through in-silico testing.

  5. NeuroFlow: Toward Unified Visual Encoding and Decoding from Neural Activity

    cs.LG 2026-04 unverdicted novelty 7.0

    NeuroFlow is the first unified flow model for bidirectional visual encoding and decoding from neural activity using NeuroVAE and cross-modal flow matching.

  6. Neuroprobe: Evaluating Intracranial Brain Responses to Naturalistic Stimuli

    cs.LG 2025-09 accept novelty 7.0

    Neuroprobe is a new suite of decoding tasks on the BrainTreebank iEEG dataset for evaluating multi-modal language processing in the brain during naturalistic movie viewing.

  7. Fast Whole-Brain, Geometry-Aware Functional Alignment for Cross-Subject Decoding

    q-bio.NC 2026-07 conditional novelty 6.5

    SpectralOT regularizes entropic optimal transport with the first three Laplace-Beltrami eigenmodes of cortical geometry to produce fast, parsimonious whole-brain functional alignments that improve cross-subject decoding.

  8. Real-time Reconstruction of Human Visual Perception from fMRI

    cs.CV 2026-07 conditional novelty 6.0

    First demonstration that single-trial visual images can be decoded from fMRI in near-real-time (about 10-15 seconds) with roughly one hour of fine-tuning data.

  9. Predicted Cortex Is Not a Domain-General Prior: A Matched-Control Audit of Brain-Encoding Features for Video Memorability

    cs.CV 2026-07 conditional novelty 6.0

    Predicted-brain features beat the visual backbone on VideoMem but lose on Memento10k for video memorability, so the benefit is dataset-specific.

  10. Predicted Cortex Is Not a Domain-General Prior: A Matched-Control Audit of Brain-Encoding Features for Video Memorability

    cs.CV 2026-07 conditional novelty 6.0

    Predicted cortical responses from a brain-encoding model beat their own visual backbone on VideoMem but lose on Memento10k, so they are a dataset-specific memorability representation, not a domain-general prior.

  11. Channel-Oriented Design for EEG-to-Music Reconstruction

    cs.SD 2026-06 unverdicted novelty 6.0

    Introduces a channel-oriented design using per-electrode tokenization, multi-view self-distillation, and structured channel dropout within an encoding-alignment-decoding pipeline to improve EEG-to-music reconstruction...

  12. MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding

    cs.LG 2026-05 unverdicted novelty 6.0

    A tri-modal contrastive learning method for EEG-based zero-shot visual decoding reports 54.1% top-1 accuracy on the Things-EEG2 200-way benchmark, outperforming prior baselines of 32.4%.

  13. Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction

    cs.CV 2026-05 unverdicted novelty 6.0

    CineNeuron improves fMRI-to-video reconstruction by combining bottom-up semantic enrichment with top-down Mixture-of-Memories integration and outperforms prior methods on benchmarks.

  14. Interpreting V1 Population Activity via Image-Neural Latent Representation Alignment

    cs.NE 2026-05 unverdicted novelty 6.0

    DINA is a dual-tower contrastive model that aligns images with mouse V1 neural activity to enable decoding and shows that low-level visual structure, not semantics or fine details, primarily supports the alignment.

  15. Meta-learning In-Context Enables Training-Free Cross Subject Brain Decoding

    cs.LG 2026-04 unverdicted novelty 6.0

    A meta-optimized in-context learning approach enables training-free cross-subject semantic visual decoding from fMRI by inferring individual neural encoding patterns via hierarchical inference on a few examples.

  16. BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain

    cs.CV 2025-12 conditional novelty 6.0

    A new automated pipeline decomposes fMRI activity into components and labels them with visual concepts, claiming thousands of interpretable patterns across the human visual cortex.

  17. A cross-species neural foundation model for end-to-end speech decoding

    cs.CL 2025-11 unverdicted novelty 6.0

    A cross-species pretrained neural encoder combined with end-to-end training and audio LLMs reduces word error rate in neural speech decoding from 24.69% to 10.22% while aligning attempted and imagined speech.

  18. Hi-DREAM: Brain-Inspired Hierarchical Diffusion for fMRI-to-Image Reconstruction via ROI Encoder and VisuAl Mapping

    cs.CV 2025-11 conditional novelty 6.0

    Hi-DREAM conditions a latent diffusion U-Net on early/mid/late visual-ROI streams via a multi-scale cortical pyramid and depth-matched ControlNet, reporting state-of-the-art semantic metrics on NSD fMRI-to-image recon...

  19. BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

    cs.CV 2026-06 unverdicted novelty 5.0

    BrainJanus presents a unified autoregressive model with a brain tokenizer that maps between neural activity, vision, and language for encoding and decoding tasks.

  20. Retrieval-Based Brain Decoding by Alignment, not Complexity

    q-bio.NC 2026-06 unverdicted novelty 5.0

    Linear contrastive decoders outperform ridge regression and non-linear alternatives when mapping fMRI activity to foundation model embeddings in vision, text, and audio.

  21. FlexiBrain: Resolution-Agnostic Voxel-Level Encoding for Native fMRI

    eess.IV 2026-06 unverdicted novelty 5.0

    FlexiBrain is a resolution-agnostic voxel-level encoding framework using Mamba-JEPA that ingests native fMRI data directly and outperforms prior methods by up to 12 points on five neuroscience tasks.

  22. FPED: A Functional-Network Prior-Guided Mixture-of-Experts Framework for Interpretable Brain Decoding

    cs.CV 2026-05 unverdicted novelty 5.0

    FPED is a functional-network prior-guided MoE framework for fMRI visual reconstruction that claims competitive performance at 0.68B parameters and biologically meaningful routing interpretability.

  23. Multi-Level Bidirectional Biomimetic Learning for EEG-Based Visual Decoding

    cs.CV 2026-05 unverdicted novelty 5.0

    MB2L achieves 80.5% top-1 and 97.6% top-5 accuracy on zero-shot EEG-to-image retrieval by using biomimetic modules and bidirectional contrastive learning to align neural and visual features.

  24. Brain-Grasp: Graph-based Saliency Priors for Improved fMRI-based Visual Brain Decoding

    eess.IV 2026-04 unverdicted novelty 5.0

    Graph-informed saliency masks derived from fMRI signals are used to condition a single diffusion model, improving object structure and semantic fidelity in visual brain decoding.

  25. VoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI

    cs.CV 2025-09 reject novelty 5.0

    A 39M-parameter transformer with token merging and query compression gives around 74% top-1 fMRI-to-image retrieval across six training subjects, far below the ~98% of larger state-of-the-art decoders.

  26. BRAIN: Bias-Mitigation Continual Learning Approach to Vision-Brain Understanding

    cs.CV 2025-08 unverdicted novelty 5.0

    BRAIN uses bias-mitigation continual learning with a new de-bias contrastive loss and angular forgetting mitigation to achieve SOTA performance on vision-brain understanding benchmarks despite brain signal inconsisten...

  27. Threat Vectors and the State of the Art in Defense Methods for Security in Neurotechnology

    cs.CR 2026-07 conditional novelty 4.5

    Neurosecurity lags BCI capability; a full-stack attack-surface taxonomy plus transferable defenses from cyber, hardware security, and private ML can close many gaps immediately.

  28. Large-Scale AI and Foundation Models for Neuroscience: A Comprehensive Review

    cs.AI 2025-10 conditional novelty 1.0

    This paper is a survey: it organizes existing foundation-model work in neuroscience into five application domains and lists public datasets, without presenting new experiments.