REVIEW 5 cited by
The Algonauts Project 2025 Challenge: How the Human Brain Makes Sense of Multimodal Movies
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
There is growing symbiosis between artificial and biological intelligence sciences: neural principles inspire new intelligent machines, which are in turn used to advance our theoretical understanding of the brain. To promote further collaboration between biological and artificial intelligence researchers, we introduce the 2025 edition of the Algonauts Project challenge: How the Human Brain Makes Sense of Multimodal Movies (https://algonautsproject.com/). In collaboration with the Courtois Project on Neuronal Modelling (CNeuroMod), this edition aims to bring forth a new generation of brain encoding models that are multimodal and that generalize well beyond their training distribution, by training them on the largest dataset of fMRI responses to movie watching available to date. Open to all, the 2025 challenge provides transparent, directly comparable results through a public leaderboard that is updated automatically after each submission to facilitate rapid model assessment and guide development. The challenge will end with a session at the 2025 Cognitive Computational Neuroscience (CCN) conference that will feature winning models. We welcome researchers interested in collaborating with the Algonauts Project by contributing ideas and datasets for future challenges.
Forward citations
Cited by 5 Pith papers
-
NeuroWorld: A Latent Brain World Model for Stimulus-Conditioned Human Brain Dynamics
A latent world model of the brain, trained to predict the next latent fMRI state from past brain states and current movie stimuli, outperforms regression-based encoders in causal multi-step rollout on three naturalist...
-
RABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brain
A 12.9M-parameter audio-to-fMRI encoder achieves zero-shot prediction of speech-evoked brain responses on 324 unseen participants and few-shot adaptation with ~10 minutes of data, outperforming larger baselines and pe...
-
A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps
TRIBE’s predicted cortical drive does not predict YouTube most-replayed heatmaps beyond position and low-level baselines, with the null bounded near r≈0.14.
-
TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction
A transformer-based encoder that combines text, audio, and video embeddings predicts whole-brain fMRI responses to movies across subjects and won the Algonauts 2025 competition.
-
VIBE: Video-Input Brain Encoder for fMRI Response Modeling
A two-stage multimodal Transformer with fixed pretrained feature extractors predicts fMRI activity from movies with Pearson r = 0.32 (in-domain) and 0.21 (out-of-domain), beating the challenge baseline by about 0.12.
Discussion (0). Sign in to comment.