Pith. sign in

REVIEW 2 cited by

MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.12630 v2 pith:OZ6CS2VX submitted 2024-04-19 cs.CV cs.MM

MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction

classification cs.CV cs.MM
keywords visualcross-subjectdatadecodingfingerprintfmri-to-textmindtunersubjects
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Decoding natural visual scenes from brain activity has flourished, with extensive research in single-subject tasks and, however, less in cross-subject tasks. Reconstructing high-quality images in cross-subject tasks is a challenging problem due to profound individual differences between subjects and the scarcity of data annotation. In this work, we proposed MindTuner for cross-subject visual decoding, which achieves high-quality and rich semantic reconstructions using only 1 hour of fMRI training data benefiting from the phenomena of visual fingerprint in the human visual system and a novel fMRI-to-text alignment paradigm. Firstly, we pre-train a multi-subject model among 7 subjects and fine-tune it with scarce data on new subjects, where LoRAs with Skip-LoRAs are utilized to learn the visual fingerprint. Then, we take the image modality as the intermediate pivot modality to achieve fMRI-to-text alignment, which achieves impressive fMRI-to-text retrieval performance and corrects fMRI-to-image reconstruction with fine-tuned semantics. The results of both qualitative and quantitative analyses demonstrate that MindTuner surpasses state-of-the-art cross-subject visual decoding models on the Natural Scenes Dataset (NSD), whether using training data of 1 hour or 40 hours.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Real-time Reconstruction of Human Visual Perception from fMRI

    cs.CV 2026-07 conditional novelty 6.0

    First demonstration that single-trial visual images can be decoded from fMRI in near-real-time (about 10-15 seconds) with roughly one hour of fine-tuning data.

  2. MindAdapter: Few-Shot Parameter-Efficient Residual Calibration of Cross-Subject Brain-to-Visual Decoding Models

    cs.CV 2026-05 unverdicted novelty 5.0

    MindAdapter introduces a decoupled linear-residual adapter with topology-anchored dual-stream manifold constraints for few-shot cross-subject calibration of brain-to-visual models, improving reconstruction and retriev...