Instruction-tuned multimodal LLMs predict fMRI responses to natural images better than vision-only models and on par with CLIP, though most explained variance is shared across instructions.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
q-bio.NC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
Instruction-tuned multimodal LLMs predict fMRI responses to natural images better than vision-only models and on par with CLIP, though most explained variance is shared across instructions.