Pith. sign in

REVIEW 10 cited by

TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.22229 v1 pith:LEGCY23A submitted 2025-07-29 cs.LG

TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction

classification cs.LG
keywords brainmodelapproachcorticalfirstfmrimodalitiesmodels
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Historically, neuroscience has progressed by fragmenting into specialized domains, each focusing on isolated modalities, tasks, or brain regions. While fruitful, this approach hinders the development of a unified model of cognition. Here, we introduce TRIBE, the first deep neural network trained to predict brain responses to stimuli across multiple modalities, cortical areas and individuals. By combining the pretrained representations of text, audio and video foundational models and handling their time-evolving nature with a transformer, our model can precisely model the spatial and temporal fMRI responses to videos, achieving the first place in the Algonauts 2025 brain encoding competition with a significant margin over competitors. Ablations show that while unimodal models can reliably predict their corresponding cortical networks (e.g. visual or auditory networks), they are systematically outperformed by our multimodal model in high-level associative cortices. Currently applied to perception and comprehension, our approach paves the way towards building an integrative model of representations in the human brain. Our code is available at https://github.com/facebookresearch/algonauts-2025.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. NEvo: Neural-Guided Evolutionary Video Synthesis for Dynamic Visual Selectivity

    cs.CV 2026-07 unverdicted novelty 7.0

    NEvo performs evolutionary search guided by a dynamic voxel-level encoding model to synthesize videos that maximize predicted activity in target brain ROIs, recovering known selectivities and revealing temporal dynami...

  2. NeuralBench: A Unifying Framework to Benchmark NeuroAI Models

    cs.LG 2026-05 conditional novelty 7.0

    NeuralBench is a new benchmarking framework for neuroAI models on EEG data that finds foundation models only marginally outperform task-specific ones while many tasks like cognitive decoding stay highly challenging.

  3. It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability

    cs.CV 2026-07 conditional novelty 6.0

    Predicted-brain features beat the visual backbone on VideoMem but lose on Memento10k for video memorability, so the benefit is dataset-specific.

  4. A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps

    cs.SE 2026-07 conditional novelty 6.0

    TRIBE’s predicted cortical drive does not predict YouTube most-replayed heatmaps beyond position and low-level baselines, with the null bounded near r≈0.14.

  5. MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding

    cs.LG 2026-05 unverdicted novelty 6.0

    A tri-modal contrastive learning method for EEG-based zero-shot visual decoding reports 54.1% top-1 accuracy on the Things-EEG2 200-way benchmark, outperforming prior baselines of 32.4%.

  6. EmoMind: Decoding Affective Captions from Human Brain fMRI

    cs.LG 2026-05 unverdicted novelty 6.0

    EmoMind is the first end-to-end pipeline that decodes continuous affective captions from fMRI by combining brain-decoded visual features with a 34D emotion vector and classifier-free guidance to balance semantic fidel...

  7. Feature Visualization Recovers Known Cortical Selectivity from TRIBE v2

    q-bio.NC 2026-05 unverdicted novelty 6.0

    Feature visualization on TRIBE v2 brain encoders recovers the known ventral visual hierarchy from V1 to V4 and produces distinctive patterns for MT, FFA, and PPA, with optimized stimuli driving ~4x higher activation t...

  8. OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens

    q-bio.NC 2026-04 unverdicted novelty 6.0

    OmniMouse demonstrates data-driven scaling in multi-task brain models on a 150B-token neural dataset, achieving SOTA across prediction, decoding, and forecasting while model size gains saturate.

  9. A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps

    cs.SE 2026-07 unverdicted novelty 5.0

    Predicted global field power from TRIBE on 48 YouTube videos shows pooled partial correlation of +0.058 (p=0.23) with replay heatmaps, null across readouts and tests.

  10. Retrieval-Based Brain Decoding by Alignment, not Complexity

    q-bio.NC 2026-06 unverdicted novelty 5.0

    Linear contrastive decoders outperform ridge regression and non-linear alternatives when mapping fMRI activity to foundation model embeddings in vision, text, and audio.