{"id":"d5094b91-7884-4da4-ae6f-46c5c5e8c9b2","arxiv_id":"2501.00504","paper_version":2,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The Algonauts 2025 challenge will benchmark fMRI encoding models on multimodal movies, with winners selected on out-of-distribution generalization.","lead":"This paper announces the 2025 Algonauts Project challenge, in which teams build computer models that predict brain activity from movies using video, audio, and language. It is the first version of this benchmark to use about 80 hours of fMRI data per person and to choose winners based on performance on movies outside the training distribution.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The OOD winner-selection metric may be too noisy to rank models: 2 hours of withheld movies with baseline r=0.09, plus 10 visible submissions, could select on noise rather than on generalization.","rationale":"The reader's weakest_assumption correctly identifies the OOD selection protocol as the load-bearing element: for the challenge's central claim to hold, the 2-hour OOD set must be reliable enough to rank encoding models, and the ten-submission cap must prevent test-set overfitting. My concern is the same one, sharpened with the specific quantitative context from the paper: the OOD baseline is r=0.09, the OOD set is only 2 hours across four subjects, and the leaderboard is updated after each of the ten allowed submissions. The manuscript provides no estimate of the metric's reliability, so the risk that winner selection is noise-driven is real and unaddressed. I agree with the reader that the appropriate verdict is UNVERDICTED: the paper makes no empirical claim that can be accepted or rejected yet, but its design promise is conditional on OOD ranking stability. The concrete split-half test would settle whether this concern lands, and it can be run either on the actual OOD data during the competition or on a pre-challenge proxy using existing CNeuroMod movies. I do not see a more fundamental internal inconsistency; the largest-dataset claim and multimodal framing are plausible, and the rules are generally clear.","tokens_in":820,"tokens_out":1043,"duration_ms":62237,"concrete_test":"Compute the split-half reliability of the OOD leaderboard metric on the actual withheld OOD data once collected: split the 2 hours of OOD stimuli into two halves (or split by movie), score every submitted model on each half, and correlate the two score vectors across models. If the split-half correlation is below about 0.8, the OOD ranking is too unstable to select winners. A pre-challenge proxy: on existing CNeuroMod data, train the baseline on Friends seasons 1-6 plus a subset of Movie10, treat one withheld Movie10 film as a 2-hour OOD proxy, and measure how well rankings from random 2-hour subsets reproduce the ranking from the full withheld film.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central promise that winners will be models that generalize OOD rests entirely on the 2-hour OOD evaluation in the model selection phase. The manuscript reports no reliability estimate for the OOD leaderboard metric. Baseline OOD correlation is r=0.09 (vs r=0.20 ID); at this level, with 2 hours of stimuli across 4 subjects and 1000 parcels, the mean score has substantial sampling error, and differences between models may be dominated by noise. The problem is compounded because the OOD leaderboard is updated after each of the maximum 10 submissions (Figure 2c, Model selection phase), so participants can select among 10 attempts on the same small set; with a noisy metric, this is selection on noise rather than on generalization. Nothing in the manuscript (e.g., split-half reliability, bootstrap confidence intervals, or an error bar on the baseline) establishes that the OOD ranking is stable enough to identify a true winner. Without that, the challenge may crown a model whose OOD score is a chance high draw, and the stated goal of selecting models solely based on OOD performance cannot be validated. This is a testable design risk, not a claim about the authors' conduct.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript introduces the Algonauts Project 2025 challenge, which asks participants to build encoding models that predict fMRI responses to multimodal movie stimuli. The challenge uses the CNeuroMod dataset: four subjects, almost 80 hours of fMRI per subject, with training data from six seasons of Friends and a set of four movies (Movie10), an in-distribution test on season 7 of Friends, and a two-hour out-of-distribution test on withheld movies. Winners are selected solely on OOD performance during a one-week selection phase with at most ten submissions, after a six-month building phase with unlimited ID submissions. The paper describes the data, phases, rules, baseline model (a linearizing encoding model with visual, audio, and language features, reporting r=0.20 ID and r=0.09 OOD), development kit, and the scientific rationale for multimodal naturalistic stimulation and OOD generalization.","tokens_in":12598,"tokens_out":4262,"duration_ms":41258,"significance":"If the challenge runs as designed, it will provide the field with a large, open, multimodal fMRI encoding benchmark with a public leaderboard and an explicit OOD generalization test, which are valuable and complementary to existing initiatives such as Brain-Score and Sensorium. The choice of CNeuroMod data and the release of a development kit and automated scoring infrastructure are concrete strengths. The paper's central scientific promise, however, is that winners selected on the OOD score will be models that genuinely generalize beyond the training distribution; that promise depends on the statistical reliability of the OOD leaderboard, which the manuscript does not yet establish.","major_comments":[{"comment":"The central claim that winners are selected 'solely based on their OOD performance' depends on the OOD score being reliable enough to rank models. The paper reports only r=0.09 for the baseline OOD score, with no error bars, split-half reliability, bootstrap confidence intervals, or subject-wise and parcel-wise variance. With 2 hours of OOD stimuli per subject and averaging over 1,000 parcels, the sampling variance of this mean correlation is non-negligible; the manuscript should quantify the reliability of the OOD leaderboard (e.g., split-half correlation across OOD movies, bootstrap CI on the baseline, or a noise ceiling estimate) and state a criterion for when differences between models are meaningful.","section":"Model selection phase and Baseline model"},{"comment":"Because the OOD leaderboard is updated after each of the up to ten submissions, participants can choose their best of ten scores on the same small OOD set. If the OOD metric is as noisy as the r=0.09 baseline suggests, this protocol selects on noise rather than on generalization. The paper should describe a safeguard, such as a final hold-out split of OOD data used only after the ten submissions are frozen, or a statistical test comparing the top submission against the baseline and against other top submissions.","section":"Rules and Model selection phase"},{"comment":"The baseline description is too underspecified to be reproduced: 'extracts visual, audio, and language features' does not state which features are used, how they are temporally aligned to the fMRI time series, what regression or regularization is applied, or whether the model is fit per subject and parcel. Since the baseline is the reference score against which all entries are judged, this omission weakens the scientific value of the benchmark and should be fixed by providing a precise specification or a link to the baseline code.","section":"Baseline model"},{"comment":"The term 'out-of-distribution' is used without defining the distribution shift. Because the OOD movies are unrevealed until the selection phase, participants cannot know the shift, but the organizers should specify what dimensions of shift are intended (e.g., new narrative content, different genres, new audiovisual statistics) and ideally provide a planned post-hoc measure of distribution shift to verify that the OOD set is actually outside the training distribution.","section":"Model selection phase"}],"minor_comments":[{"comment":"The phrase 'all episodes of seasons 7 of the Friends dataset' should be 'season 7'.","section":"Model building phase"},{"comment":"The word 'premiating' should be replaced with 'rewarding' or 'prizing'.","section":"Discussion"},{"comment":"The two Richards et al. 2019 entries are identical; one duplicate should be removed.","section":"References"},{"comment":"The sentence 'Further information on the challenge stimuli and fMRI data is provided on the challenge data repository and development kit' lacks a URL or specific citation; please add the repository link.","section":"Data"},{"comment":"The caption and text mention the maximum of ten submissions in the model selection phase, but the figure should visually indicate this cap to avoid ambiguity.","section":"Figure 2c"}],"recommendation":"major_revision","confidential_remarks":"This is a challenge description paper without new experimental results; its contribution is organizational and infrastructural. The main risk is that the OOD leaderboard may be too noisy to support winner selection. If the journal regularly publishes such benchmark/challenge papers, this is a reasonable fit; otherwise, the reliability analysis should be a required part of the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this is a benchmark specification, not a research paper with a falsifiable claim. What is actually new: the 2025 Algonauts edition moves to multimodal movie stimuli, provides roughly 80 hours of fMRI per subject from CNeuroMod, and makes out-of-distribution generalization the explicit winner-selection criterion. That combination is a genuine step beyond the 2021 and 2023 editions. The protocol is clearly described, the rules are sensible (no using withheld responses, one account, code release for top-3), and the development kit lowers the entry barrier. The baseline correlations (ID r=0.20, OOD r=0.09) give participants a reference point. Credit where due: the paper does not oversell; it presents the challenge as a platform and acknowledges that prediction and explanation are different things.\n\nThe soft spots are the ones you would expect from a design-only paper. Most importantly, the OOD leaderboard may be too noisy to rank models reliably. The selection phase uses only 2 hours of withheld movies, the baseline OOD correlation is 0.09, and the leaderboard is updated after each of up to ten submissions. With a noisy metric, selecting the best of ten attempts can become selection on noise. The manuscript gives no split-half reliability, bootstrap confidence intervals, or subject-wise breakdowns. Averaging over 1000 parcels and 4 subjects helps, but it does not automatically save the ranking. This is a real, testable risk, and the authors should either report a reliability estimate for the OOD score or reduce the number of selection submissions. It is not fatal to the challenge concept, but it would be embarrassing if the 2025 winner turned out to be a chance high draw.\n\nTwo minor points: the baseline is reported without any measure of variance, which is easy to fix, and the OOD stimuli are described only as \"withheld movies\"—readers cannot judge how far out of distribution they are. The authors probably keep the identity secret to prevent overfitting, but a post-hoc characterization would help.\n\nOverall, this is a useful community resource that deserves a serious referee. The central design is sound; the OOD noise concern is legitimate but addressable. I would send it to review with the request that the authors add a reliability analysis for the OOD metric before publication. For a reading group on benchmarks and encoding models, it is a reasonable \"maybe\"—the paper itself is short and the real content is the challenge infrastructure.","headline":"A well-specified challenge announcement, not a results paper; the OOD leaderboard may be too noisy to crown a reliable winner, but that is a fixable design gap, not a fatal flaw.","tokens_in":13191,"tokens_out":1703,"would_cite":true,"duration_ms":19945,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces a brain-encoding challenge whose central claim is that the best models of the brain are those that generalize beyond their training distribution, and it operationalizes that claim by selecting the winner solely on…","keywords":["brain encoding models","out-of-distribution generalization","multimodal movies","naturalistic stimuli","large-scale fMRI","community benchmark","fMRI movie watching"],"falsifier":"Compute a feature-space distance, using a pretrained audio-visual or language model, between the withheld out-of-distribution movies and the sitcom and film training material; if that distance is no larger than the distance between two seasons of the same sitcom, the out-of-distribution set is not actually out-of-distribution and the challenge's main ranking would not measure generalization.","tokens_in":12213,"feed_emoji":"🧠","tokens_out":8139,"duration_ms":78032,"temperature":0.7,"pith_summary":"This paper introduces a community competition for building models of how the brain responds to movies, and states the central goal: encoding models should predict whole-brain fMRI responses to naturalistic, multimodal stimulation and keep working on stimuli outside their training distribution. To make that goal testable, the challenge supplies almost 80 hours of movie-watching fMRI data per subject, including sitcom episodes and feature films, together with visual, audio, and transcript tracks. Participants train on six seasons of the sitcom and four films, then submit predicted brain responses for a held-out seventh season with unlimited attempts and for two hours of unrevealed out-of-distribution movies with at most ten attempts. The winning models are chosen solely on the out-of-distribution scores. The baseline encoding model scores r=0.20 in-distribution but only r=0.09 out-of-distribution, so the paper's design makes the generalization gap the explicit target.","feed_headline":"Brain-encoding challenge rewards models that generalize to new movies","feed_subtitle":"Almost 80 hours of fMRI per viewer and a winner-selection rule that only counts predictions on unseen films.","key_machinery":"The central mechanism is a two-phase, two-leaderboard evaluation split. In the six-month building phase, models train on 55 hours of the sitcom and 10 hours of four films, and are tested on the held-out subsequent season with unlimited submissions; this gives in-distribution performance. In the one-week selection phase, models predict brain responses to two hours of withheld movies from outside that distribution, with a maximum of ten submissions, and winners are ranked solely on this out-of-distribution score. Scoring is done by averaging Pearson correlation over the 1,000 functionally defined brain parcels, then over out-of-distribution movies or held-out episodes, then over the four subjects. The load-bearing baseline is a linearizing encoding model that maps extracted visual, audio, and language features to fMRI responses, achieving r=0.20 in-distribution and r=0.09 out-of-distribution.","core_discovery":"The paper's claim is that the next generation of brain encoding models should be multimodal and should be selected by their out-of-distribution generalization, and it offers a concrete way to measure that: the largest single-subject fMRI movie-watching dataset assembled so far, split so that training, in-distribution testing, and out-of-distribution testing draw on different content. An encoding model is any algorithm that maps movie stimuli, including visual frames, audio, and time-stamped language transcripts, to predicted fMRI activity in 1,000 cortical parcels for each of four subjects. The quality of a model is a single number: Pearson correlation between predicted and recorded responses, averaged across parcels, then across stimuli, then across subjects. The challenge is organized so that the winner cannot be chosen by repeated probing of the out-of-distribution test set; the out-of-distribution movies are revealed only in a one-week selection phase, and only ten submissions are allowed. The intended payoff is a transparent leaderboard that ranks models both on a familiar test from the same TV series and on a genuinely novel test, making out-of-distribution robustness part of what a good brain model means.","pith_inferences":["A natural extension the paper gestures at but does not implement is disaggregating out-of-distribution scores by modality or by movie genre; such breakdowns could reveal exactly which stimulus properties models fail to capture.","The design implicitly predicts that in-distribution leaderboard ranks will not match out-of-distribution ranks; if they do match closely, selecting winners by out-of-distribution performance would add little information.","After the challenge, a decisive validity check would be to compare the out-of-distribution-selected winners against a fresh, never-revealed movie set; if their advantage evaporates, the two-hour out-of-distribution set was too small or too similar to the training data.","If the benchmark works, the same train-withhold-out-of-distribution structure could be adopted for other brain-encoding problems, for example predicting responses during tasks with active cognition rather than passive movie watching."],"forward_implications":["Multimodal fusion of visual frames, audio, and transcripts becomes a prerequisite for top scores, since the stimuli and scoring combine all three modalities.","Unlimited in-distribution submissions let teams tune their models on a familiar test, while the ten-submission out-of-distribution cap makes the final ranking a test of generalization rather than leaderboard overfitting.","The single averaged correlation score makes models of any architecture directly comparable on identical data and identical preprocessing.","The indefinite post-challenge phase turns the challenge into a permanent public benchmark with separate in-distribution and out-of-distribution leaderboards.","If participants close the gap between the r=0.20 and r=0.09 baselines, that would demonstrate that data-hungry, end-to-end trained encoding models can generalize to new movies."],"supporting_citations":[{"why":"Supplies the source dataset of intensive single-subject fMRI responses to naturalistic movies that the challenge is built on.","marker":"Boyle et al. 2023"},{"why":"Predecessor challenge that establishes the format of predicting brain responses and comparing models on a public leaderboard.","marker":"Cichy et al. 2021"},{"why":"Previous challenge edition that the 2025 edition extends to multimodal, movie-length stimuli.","marker":"Gifford et al. 2023"},{"why":"Defines the linearizing encoding model used to compute the baseline correlation scores.","marker":"Naselaris et al. 2011"},{"why":"Provides the 1,000-parcel cortical atlas used to define the fMRI response units that are scored.","marker":"Schaefer et al. 2018"},{"why":"Shows that out-of-distribution test scores can separate models with similar in-distribution scores, motivating the winner-selection rule.","marker":"Ren and Bashivan 2024"}],"fun_headline_variants":["Brain encoding challenge pushes models to generalize to new movies","Algonauts 2025: Predict brain activity from multimodal movie watching","Multimodal movie fMRI challenge: win by generalizing to unseen films","Large movie fMRI dataset spurs brain-encoding models that generalize"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole design rests on two hours of withheld movies being different enough from the training films, and reliable enough, that an out-of-distribution correlation score is a fair and stable ranking of model quality; if those movies are too similar, too noisy, or too short, the central test of generalization fails.","fun_headline_variants_meta":{"raw":{"variants":["Brain encoding challenge pushes models to generalize to new movies","Algonauts 2025: Predict brain activity from multimodal movie watching","Multimodal movie fMRI challenge: win by generalizing to unseen films","Large movie fMRI dataset spurs brain-encoding models that generalize"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000455,"raw_usage":{"total_tokens":2299,"prompt_tokens":973,"completion_tokens":1326,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":1253}},"tokens_in":589,"tokens_out":1326,"duration_ms":10917,"temperature":1.0,"reasoning_tokens":1253,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:48:53.307015+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute a feature-space distance, using a pretrained audio-visual or language model, between the withheld out-of-distribution movies and the sitcom and film training material; if that distance is no larger than the distance between two seasons of the same sitcom, the out-of-distribution set is not actually out-of-distribution and the challenge's main ranking would not measure generalization.","supporting_citations":[],"review_version":1}