Pith. sign in

REVIEW 2 cited by

Vision-Language Integration in Multimodal Video Transformers (Partially) Aligns with the Brain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.07766 v1 pith:WYWY66P7 submitted 2023-11-13 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords brainmulti-modalinformationmodalitiesvideoalignmentevidencemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Integrating information from multiple modalities is arguably one of the essential prerequisites for grounding artificial intelligence systems with an understanding of the real world. Recent advances in video transformers that jointly learn from vision, text, and sound over time have made some progress toward this goal, but the degree to which these models integrate information from modalities still remains unclear. In this work, we present a promising approach for probing a pre-trained multimodal video transformer model by leveraging neuroscientific evidence of multimodal information processing in the brain. Using brain recordings of participants watching a popular TV show, we analyze the effects of multi-modal connections and interactions in a pre-trained multi-modal video transformer on the alignment with uni- and multi-modal brain regions. We find evidence that vision enhances masked prediction performance during language processing, providing support that cross-modal representations in models can benefit individual modalities. However, we don't find evidence of brain-relevant information captured by the joint multi-modal transformer representations beyond that captured by all of the individual modalities. We finally show that the brain alignment of the pre-trained joint representation can be improved by fine-tuning using a task that requires vision-language inferences. Overall, our results paint an optimistic picture of the ability of multi-modal transformers to integrate vision and language in partially brain-relevant ways but also show that improving the brain alignment of these models may require new approaches.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A transformer-based encoder that combines text, audio, and video embeddings predicts whole-brain fMRI responses to movies across subjects and won the Algonauts 2025 competition.

  2. Multi-modal brain encoding models for multi-modal stimuli

    q-bio.NC 2025-05 conditional novelty 5.0 of 10

    On movie-watching fMRI data, multi-modal vision-audio transformers predict brain activity better than unimodal video or speech models, with video dominating cross-modal alignment and video plus audio jointly contribut...

Pith tools