Pith. sign in

REVIEW 4 cited by

Brain Captioning: Decoding human brain activity into images and text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.11560 v1 pith:IFN5TZIK submitted 2023-05-19 cs.CV cs.AI

classification cs.CVcs.AI
keywords brainimagescaptioningactivitydecodinghumanimagemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Every day, the human brain processes an immense volume of visual information, relying on intricate neural mechanisms to perceive and interpret these stimuli. Recent breakthroughs in functional magnetic resonance imaging (fMRI) have enabled scientists to extract visual information from human brain activity patterns. In this study, we present an innovative method for decoding brain activity into meaningful images and captions, with a specific focus on brain captioning due to its enhanced flexibility as compared to brain decoding into images. Our approach takes advantage of cutting-edge image captioning models and incorporates a unique image reconstruction pipeline that utilizes latent diffusion models and depth estimation. We utilized the Natural Scenes Dataset, a comprehensive fMRI dataset from eight subjects who viewed images from the COCO dataset. We employed the Generative Image-to-text Transformer (GIT) as our backbone for captioning and propose a new image reconstruction pipeline based on latent diffusion models. The method involves training regularized linear regression models between brain activity and extracted features. Additionally, we incorporated depth maps from the ControlNet model to further guide the reconstruction process. We evaluate our methods using quantitative metrics for both generated captions and images. Our brain captioning approach outperforms existing methods, while our image reconstruction pipeline generates plausible images with improved spatial relationships. In conclusion, we demonstrate significant progress in brain decoding, showcasing the enormous potential of integrating vision and language to better understand human cognition. Our approach provides a flexible platform for future research, with potential applications in various fields, including neural art, style transfer, and portable devices.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NSD-Imagery: A benchmark dataset for extending fMRI vision decoding methods to mental imagery

    cs.CV 2025-06 conditional novelty 7.0 of 10

    NSD-Imagery is a released benchmark of fMRI responses to imagined pictures from Natural Scenes Dataset participants, and benchmarks of five decoders show mental imagery performance is largely decoupled from seen-image...

  2. fMRI2Face: A Full-HD fMRI-Video Dataset and Geometry-Guided Neural Decoding Framework for Dynamic Human Face Reconstruction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A 62,856-sample fMRI dataset of full-HD digital-human face videos plus a geometry-guided video-diffusion decoder that reconstructs facial identity and motion from brain signals.

  3. Dynadiff: Single-stage Decoding of Images from Continuously Evolving fMRI

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A single-stage, LoRA-finetuned diffusion model decodes seen images directly from continuous BOLD fMRI time series and beats previous pipelines on semantic metrics.

  4. Exploring The Visual Feature Space for Multimodal Neural Decoding

    cs.CV 2025-05 conditional novelty 5.0 of 10

    VINDEX decodes fMRI into nested 9-token CLIP features that feed a frozen MLLM, improving detailed brain captioning and QA over a regression baseline, with a new benchmark.

Pith tools