{"id":"c485d9e7-f324-4a30-8f36-4f57ddb70635","arxiv_id":"2607.01400","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"TRIBE’s predicted cortical drive does not predict YouTube most-replayed heatmaps beyond position and low-level baselines, with the null bounded near r≈0.14.","lead":"A top multimodal brain-encoding model’s predicted fMRI drive does not track which moments of YouTube videos people rewatch, even after careful controls. The result is a bounded null that warns against treating brain encoders as off-the-shelf engagement predictors.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the paper's own scope bounds; the bounded null is internally well-supported.","rationale":"The paper's central contribution is a carefully bounded negative result plus a mechanistic account (group-trained encoder collapses idiosyncratic structure) and a methodological caution (spurious supervised r=0.47 under coarse detrend). The statistical package (partial correlation, baselines, Bayes, equivalence, reliability, permutation, matched/mismatched probes, own per-subject ISC) is coherent and the null is not a bare non-significant p. The reader's weakest_assumption correctly identifies the main residual risks (most-replayed bias, missing NAcc, N=48 popular 60s windows), but the paper already states these as scope limits rather than claiming universal transfer failure. Because the strongest claim is already written with those bounds, the concern does not force a verdict change from CONDITIONAL. A concrete re-run on the denser-marker subset under the stricter spline control would still be worth doing for robustness, but is not expected to overturn the reported numbers. Verdict remains CONDITIONAL: accept the bounded null within stated scope; do not over-generalize.","tokens_in":12267,"tokens_out":576,"duration_ms":6452,"concrete_test":"Using the released video-ID manifest and code, recompute the primary pooled position-controlled partial correlation after (i) restricting to the subset of clips with ≥20 in-window native markers and (ii) applying the cubic-spline (7 df) residualization used in §6.5; if the point estimate remains inside [-0.05, 0.12] and BF01 stays >3, the bounded null is stable under the paper's own stricter controls.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest_assumption (biased most-replayed target + cortical mesh without NAcc + popular-only 60s clips) correctly flags the main scope limit, but it is not a load-bearing threat to the paper's actual strongest claim. That claim is already scoped as 'on this target and these readouts... approximately no content-specific re-watch signal, up to r≈0.14' and is backed by position-controlled partial r=+0.058, baselines, BF01=3.2, TOST/equivalence, reliability ceiling, network/ROI/signed variants, permutation, and the supervised-probe artifact autopsy. The anatomy and target biases are disclosed in §8 and §3.3; they limit generalization, not the internal validity of the reported null. No hidden inconsistency or unaddressed confound overturns the bounded claim as written.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript tests whether predicted cortical responses from TRIBE (the 2025 Algonauts winner) forecast moment-level YouTube re-watch behavior. TRIBE is run unmodified on 48 videos; its predicted surface response is reduced to a per-TR global field power (GFP) engagement curve and compared to each video’s most-replayed heatmap. The primary metric is a position-controlled partial correlation (quadratic detrend). The pooled partial r is +0.058 (95% CI [−0.04, 0.15]; t(47)=1.21, p=0.23), not above loudness/motion baselines, and near zero in raw form outside a music-video onset artifact. The null is stable across network and signed value/salience ROI readouts, a circular-shift permutation, and a supervised leave-one-video-out cortical probe that only appears successful (r=0.47) under a coarse detrend and collapses under spline + matched/mismatched controls. A borderline content-specific signal is reported only in the visual input stream; audio, text, and predicted cortex show none. Because the released TRIBE checkpoint is subject-averaged, the authors fit their own per-subject encoders and find that predicted ISC also fails to track re-watch (r=−0.04). They bound rather than merely fail to reject the null (BF01=3.2; effects ≳0.14 excluded; target split-half reliability ≈0.82). Code, video-ID manifest, and a SABR-resilient acquisition method are released.","tokens_in":12457,"tokens_out":1682,"duration_ms":38295,"significance":"If the bounded null holds, the paper is a useful negative result for anyone tempted to treat modern multimodal brain encoders (or their foundation features) as off-the-shelf engagement predictors. Its main contributions are (i) a carefully scoped claim with Bayes factor, equivalence-style bound, and reliability ceiling rather than a bare non-significant p-value; (ii) a concrete methodological caution that most-replayed labels plus a coarse detrend can manufacture large cross-validated correlations that are pure shared temporal shape; and (iii) a mechanistic sketch that group-trained encoding collapses the weak video-specific structure present in visual features. Full release of code, video IDs, and an acquisition pipeline that works under SABR is a genuine reproducibility strength and should be credited. The result is narrow by design (one model, cortical surface, biased target, 60 s window), but within that scope it is informative.","major_comments":[{"comment":"§6.4 vs abstract/conclusion: the pre-specified TOST against δ=0.10 does not establish equivalence (p=0.20). The claim that effects larger than r≈0.14 are excluded comes from the descriptive 90% CI, not from a successful TOST at a pre-registered SEI. Please state explicitly what was pre-specified versus post-hoc, and align the abstract wording (“an equivalence test excludes effects above r≈0.14”) with §6.4 so readers do not read a failed TOST as a passed equivalence test. BF01=3.2 and the reliability ceiling already support a bounded null; the equivalence language just needs to match the procedure.","section":"§6.4, Abstract, Conclusion"},{"comment":"§3.2 and §5 fix the analysis to the first 60 s (≈60 TRs) of every clip. Most-replayed markers span the full video, and intros/onsets are exactly the position artifact the paper works hard to remove. §6.8 shows that marker density and duration do not correlate with the effect, which is helpful, but does not test whether the null is specific to the opening minute. A sensitivity check—e.g., a mid-clip 60 s window, or full-length analysis where TRIBE segment onsets allow—would make the “moment-level” claim less dependent on a single, intro-heavy window. If that is infeasible for encoding cost, state the limitation more prominently in §8 and qualify the claim accordingly.","section":"§3.2, §5, §6.8, §8"},{"comment":"§6.7 is presented as a direct test of the closest prior positive result (ISC). The custom per-subject ridge encoders are validated at only r≈0.15 in-domain and r≈0.10 cross-domain. That is within the published range for feature-to-cortex models, but it leaves open how large a true ISC–rewatch effect could have been and still been missed. Please add a brief power or attenuation note (e.g., expected attenuation of a behavioral correlation given encoder r≈0.10–0.15) so the ISC null is interpretable as evidence against transfer rather than only as a weak encoder. The paper already notes that the released TRIBE checkpoint cannot supply per-subject predictions; that disclosure is good and should stay.","section":"§6.7"}],"minor_comments":[{"comment":"Table 1 leaves raw r blank for the loudness and motion baselines; either fill those cells or note in the caption that only partial r is reported for baselines by design.","section":"Table 1"},{"comment":"Figure 1c category labels include “?” for n=4; replace with the “misc” label used in §5 for consistency.","section":"Figure 1c"},{"comment":"Figure 2a labels the equivalence region ±0.139 while the text uses r≈0.14; pick one rounding and use it throughout abstract, §6.4, and the figure.","section":"Figure 2a, §6.4"},{"comment":"Eq. (1) GFP is clear; Eq. (2) would benefit from an explicit statement that the same B=[1,t,t²] is fit independently to e and g (which is implied but not written).","section":"§3.4, Eq. (2)"},{"comment":"§6.1 video-level ranking (views/likes) is a useful negative control but sits somewhat apart from the moment-level claim; a one-sentence pointer in the introduction or discussion would help readers see why it is included.","section":"§6.1"},{"comment":"Primary arXiv category is listed as cs.SE; for discoverability the authors may want cs.AI / q-bio.NC / cs.CV as cross-lists if the venue allows, since the contribution is brain encoding and engagement prediction rather than software engineering per se.","section":"Metadata / venue fit"},{"comment":"Minor prose: “theneuroforecastingliterature” and similar missing spaces appear in the introduction PDF text; check for tokenization artifacts before camera-ready.","section":"§1–§2"}],"recommendation":"minor_revision","confidential_remarks":"The central bounded null is carefully argued and the supervised-probe artifact autopsy is genuinely useful; I do not see a load-bearing flaw that would justify major revision or reject. The three major points are clarifications and one sensitivity request, all fixable without new data collection beyond re-analysis. Venue fit is the only soft concern: the work is computational neuroscience / multimodal encoding, not software engineering, despite the cs.SE primary category and the SABR acquisition engineering. If the journal’s scope is methods/ML-for-neuro, this is appropriate; if it is classical SE, it is a mismatch. No citation or novelty concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple. On 48 YouTube clips, TRIBE’s predicted whole-cortex drive (GFP) does not forecast most-replayed after a quadratic position control: pooled partial r = +0.058, CI crosses zero, not above loudness or motion. They bound it rather than just fail to reject—BF01 = 3.2, equivalence excluding effects above ~0.14, target split-half reliability ~0.82—and they show a supervised LOVO probe that looks like r = 0.47 is pure shared temporal shape under a spline + matched/mismatched check.\n\nWhat is new is the transfer test itself. Prior neuroforecasting used measured signals, mostly cross-item or ISC. Here they take a current Algonauts winner, reduce it to a moment-level engagement curve, and ask whether the prediction inherits re-watch signal. The input-vs-cortex probe is useful: a small borderline video-specific effect sits in the visual stream and disappears in the predicted cortex, which supports their “regress to the mean” story. Rebuilding per-subject encoders to test predicted ISC (null at r ≈ −0.04) is the right move given the released checkpoint averages subjects. Code, video IDs, and the SABR acquisition path are released; that is real work.\n\nSoft spots are mostly scope, and the paper already names them. N = 48 popular videos, first 60 s only, most-replayed is a biased target (onsets, chapters, seek-back), and fsaverage5 has no nucleus accumbens. Those limit how far you can generalize, not the internal claim as written. Free choices (window length, basis order, ridge probe, equivalence delta) are ordinary analysis knobs, not circularity. Math and stats look solid; citations track the right measured-neuroforecasting and encoding lines.\n\nThis is for people who might treat brain-encoding models or their foundation features as off-the-shelf engagement predictors, and for anyone using most-replayed as a label. It is a methodological caution with a bounded empirical result, not a theory rewrite. I would send it to peer review. A serious referee can push on sample balance and subcortical follow-up without the paper collapsing. Worth engaging if you work in this area; cite the null and the probe autopsy if you need a careful negative.","headline":"A carefully bounded null: TRIBE predicted cortical drive does not track YouTube most-replayed beyond position and low-level baselines, with real artifact controls and released code.","tokens_in":13094,"tokens_out":578,"would_cite":true,"duration_ms":5064,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Predicted cortical drive from a top brain-encoding model does not track which moments of YouTube videos people re-watch, up to a bound of r≈0.14.","keywords":["brain encoding","neuroforecasting","fMRI prediction","YouTube most-replayed","TRIBE","global field power","inter-subject correlation","engagement prediction"],"falsifier":"A subcortex-inclusive encoder that includes nucleus accumbens, or a cleaner creator-side audience-retention target on a larger balanced sample, yielding a position-controlled partial correlation clearly above the loudness baseline and outside the current r≈0.14 equivalence bound.","tokens_in":13108,"feed_emoji":"🧠","tokens_out":736,"duration_ms":6039,"temperature":0.7,"pith_summary":"Deep multimodal models can now predict fMRI responses to naturalistic video with high accuracy. This paper asks whether those predicted neural signals also forecast real-world engagement, using YouTube “most replayed” heatmaps as a passive, moment-level proxy for re-watch. The authors run TRIBE, a trimodal encoder, on 48 videos, collapse its predicted cortex into a per-second global-field-power curve, and find no content-specific association after position control: the pooled partial correlation is near zero, no better than loudness or motion baselines, and robust across network and ROI readouts. A supervised cortical probe that first looks successful collapses into a shared temporal-shape artifact once position is controlled properly. A small, borderline video-specific signal appears only in the visual input stream and is lost by the encoding step; predicted inter-subject correlation, the closest prior positive route, also fails. The authors bound the null with a Bayes factor, an equivalence test, and a high target reliability ceiling, and they release the pipeline so others can test subcortical or cleaner-target variants.","feed_headline":"Predicted brain signals fail to forecast YouTube re-watches","feed_subtitle":"Top multimodal encoder’s cortical drive is near zero after position control, bounded at r≈0.14","key_machinery":"Position-controlled partial correlation of TRIBE global field power (root-mean-square over cortical vertices) against YouTube most-replayed heatmaps, with matched/mismatched leave-one-video-out probes and an equivalence/Bayes bound on the null.","core_discovery":"On this target and these readouts, a predicted-fMRI drive signal from TRIBE carries approximately no content-specific re-watch signal, up to a bound of r≈0.14. The pooled position-controlled partial correlation is +0.058 (95% CI [−0.04, 0.15]), indistinguishable from zero and from low-level baselines; the null holds for network GFP, signed value ROIs, permutation tests, and predicted ISC, while an apparent supervised r=0.47 is a shared temporal-shape artifact.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["TRIBE's predicted fMRI drive fails to predict YouTube rewatches","Global predicted-fMRI signal shows no link to replay heatmaps","Predicted cortical drive unbound from YouTube most-replayed maps","TRIBE encoder responses carry no content-specific re-watch signal","Null: multimodal brain prediction does not forecast rewatch heat"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That most-replayed heatmaps on already-popular clips, over a 60-second window and compared mainly via cortical-surface summaries, are a fair enough behavioral target to detect neuroforecasting transfer if it existed.","fun_headline_variants_meta":{"raw":{"variants":["TRIBE's predicted fMRI drive fails to predict YouTube rewatches","Global predicted-fMRI signal shows no link to replay heatmaps","Predicted cortical drive unbound from YouTube most-replayed maps","TRIBE encoder responses carry no content-specific re-watch signal","Null: multimodal brain prediction does not forecast rewatch heat"]},"model":"grok-4.5","effort":"low","cost_usd":0.005554,"raw_usage":{"total_tokens":1687,"prompt_tokens":1046,"num_sources_used":0,"completion_tokens":92,"cost_in_usd_ticks":55540000,"prompt_tokens_details":{"text_tokens":1046,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":549,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":1046,"tokens_out":92,"duration_ms":4389,"temperature":1.0,"reasoning_tokens":549,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T08:51:32.577008+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A subcortex-inclusive encoder that includes nucleus accumbens, or a cleaner creator-side audience-retention target on a larger balanced sample, yielding a position-controlled partial correlation clearly above the loudness baseline and outside the current r≈0.14 equivalence bound.","supporting_citations":[],"review_version":2}