REVIEW 28 cited by
MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of Data
read the original abstract
Reconstructions of visual perception from brain activity have improved tremendously, but the practical utility of such methods has been limited. This is because such models are trained independently per subject where each subject requires dozens of hours of expensive fMRI training data to attain high-quality results. The present work showcases high-quality reconstructions using only 1 hour of fMRI training data. We pretrain our model across 7 subjects and then fine-tune on minimal data from a new subject. Our novel functional alignment procedure linearly maps all brain data to a shared-subject latent space, followed by a shared non-linear mapping to CLIP image space. We then map from CLIP space to pixel space by fine-tuning Stable Diffusion XL to accept CLIP latents as inputs instead of text. This approach improves out-of-subject generalization with limited training data and also attains state-of-the-art image retrieval and reconstruction metrics compared to single-subject approaches. MindEye2 demonstrates how accurate reconstructions of perception are possible from a single visit to the MRI facility. All code is available on GitHub.
Forward citations
Cited by 28 Pith papers
-
Can neurons speak? Semantic narration of vision at single-cell resolution
NEURRATOR bridges neural spike trains to frozen CLIP patch embeddings via a learned encoder, then uses a multimodal LM and sparse autoencoder to produce validated natural-language narrations of viewed scenes from Neur...
-
Brain-IT-VQA: From Brain Signals to Answers
Brain-IT-VQA decodes visual question answers from fMRI using a transformer to extract language tokens and introduces the NSD-VQA benchmark with 20 controlled questions per image across 20 categories.
-
From Activation to Causality: Discovery of Causal Visual Representations in the Human Brain
BrainCause recovers known visual localizations and finds new candidate representations by validating causal specificity via counterfactual stimuli and encoding models, showing activation alone produces many false positives.
-
A foundation model of vision, audition, and language for in-silico neuroscience
TRIBE v2 is a multimodal AI model that predicts human brain activity more accurately than linear encoding models and recovers established neuroscientific findings through in-silico testing.
-
NeuroFlow: Toward Unified Visual Encoding and Decoding from Neural Activity
NeuroFlow is the first unified flow model for bidirectional visual encoding and decoding from neural activity using NeuroVAE and cross-modal flow matching.
-
Neuroprobe: Evaluating Intracranial Brain Responses to Naturalistic Stimuli
Neuroprobe is a new suite of decoding tasks on the BrainTreebank iEEG dataset for evaluating multi-modal language processing in the brain during naturalistic movie viewing.
-
Fast Whole-Brain, Geometry-Aware Functional Alignment for Cross-Subject Decoding
SpectralOT regularizes entropic optimal transport with the first three Laplace-Beltrami eigenmodes of cortical geometry to produce fast, parsimonious whole-brain functional alignments that improve cross-subject decoding.
-
Real-time Reconstruction of Human Visual Perception from fMRI
First demonstration that single-trial visual images can be decoded from fMRI in near-real-time (about 10-15 seconds) with roughly one hour of fine-tuning data.
-
Predicted Cortex Is Not a Domain-General Prior: A Matched-Control Audit of Brain-Encoding Features for Video Memorability
Predicted-brain features beat the visual backbone on VideoMem but lose on Memento10k for video memorability, so the benefit is dataset-specific.
-
Predicted Cortex Is Not a Domain-General Prior: A Matched-Control Audit of Brain-Encoding Features for Video Memorability
Predicted cortical responses from a brain-encoding model beat their own visual backbone on VideoMem but lose on Memento10k, so they are a dataset-specific memorability representation, not a domain-general prior.
-
Channel-Oriented Design for EEG-to-Music Reconstruction
Introduces a channel-oriented design using per-electrode tokenization, multi-view self-distillation, and structured channel dropout within an encoding-alignment-decoding pipeline to improve EEG-to-music reconstruction...
-
MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding
A tri-modal contrastive learning method for EEG-based zero-shot visual decoding reports 54.1% top-1 accuracy on the Things-EEG2 200-way benchmark, outperforming prior baselines of 32.4%.
-
Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction
CineNeuron improves fMRI-to-video reconstruction by combining bottom-up semantic enrichment with top-down Mixture-of-Memories integration and outperforms prior methods on benchmarks.
-
Interpreting V1 Population Activity via Image-Neural Latent Representation Alignment
DINA is a dual-tower contrastive model that aligns images with mouse V1 neural activity to enable decoding and shows that low-level visual structure, not semantics or fine details, primarily supports the alignment.
-
Meta-learning In-Context Enables Training-Free Cross Subject Brain Decoding
A meta-optimized in-context learning approach enables training-free cross-subject semantic visual decoding from fMRI by inferring individual neural encoding patterns via hierarchical inference on a few examples.
-
BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain
A new automated pipeline decomposes fMRI activity into components and labels them with visual concepts, claiming thousands of interpretable patterns across the human visual cortex.
-
A cross-species neural foundation model for end-to-end speech decoding
A cross-species pretrained neural encoder combined with end-to-end training and audio LLMs reduces word error rate in neural speech decoding from 24.69% to 10.22% while aligning attempted and imagined speech.
-
Hi-DREAM: Brain-Inspired Hierarchical Diffusion for fMRI-to-Image Reconstruction via ROI Encoder and VisuAl Mapping
Hi-DREAM conditions a latent diffusion U-Net on early/mid/late visual-ROI streams via a multi-scale cortical pyramid and depth-matched ControlNet, reporting state-of-the-art semantic metrics on NSD fMRI-to-image recon...
-
BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language
BrainJanus presents a unified autoregressive model with a brain tokenizer that maps between neural activity, vision, and language for encoding and decoding tasks.
-
Retrieval-Based Brain Decoding by Alignment, not Complexity
Linear contrastive decoders outperform ridge regression and non-linear alternatives when mapping fMRI activity to foundation model embeddings in vision, text, and audio.
-
FlexiBrain: Resolution-Agnostic Voxel-Level Encoding for Native fMRI
FlexiBrain is a resolution-agnostic voxel-level encoding framework using Mamba-JEPA that ingests native fMRI data directly and outperforms prior methods by up to 12 points on five neuroscience tasks.
-
FPED: A Functional-Network Prior-Guided Mixture-of-Experts Framework for Interpretable Brain Decoding
FPED is a functional-network prior-guided MoE framework for fMRI visual reconstruction that claims competitive performance at 0.68B parameters and biologically meaningful routing interpretability.
-
Multi-Level Bidirectional Biomimetic Learning for EEG-Based Visual Decoding
MB2L achieves 80.5% top-1 and 97.6% top-5 accuracy on zero-shot EEG-to-image retrieval by using biomimetic modules and bidirectional contrastive learning to align neural and visual features.
-
Brain-Grasp: Graph-based Saliency Priors for Improved fMRI-based Visual Brain Decoding
Graph-informed saliency masks derived from fMRI signals are used to condition a single diffusion model, improving object structure and semantic fidelity in visual brain decoding.
-
VoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI
A 39M-parameter transformer with token merging and query compression gives around 74% top-1 fMRI-to-image retrieval across six training subjects, far below the ~98% of larger state-of-the-art decoders.
-
BRAIN: Bias-Mitigation Continual Learning Approach to Vision-Brain Understanding
BRAIN uses bias-mitigation continual learning with a new de-bias contrastive loss and angular forgetting mitigation to achieve SOTA performance on vision-brain understanding benchmarks despite brain signal inconsisten...
-
Threat Vectors and the State of the Art in Defense Methods for Security in Neurotechnology
Neurosecurity lags BCI capability; a full-stack attack-surface taxonomy plus transferable defenses from cyber, hardware security, and private ML can close many gaps immediately.
-
Large-Scale AI and Foundation Models for Neuroscience: A Comprehensive Review
This paper is a survey: it organizes existing foundation-model work in neuroscience into five application domains and lists public datasets, without presenting new experiments.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.