{"id":"06a9d2db-47b0-45ad-8f72-6f00e967e903","arxiv_id":"2505.21304","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"In a small cohort, fMRI decoders improve with more data per participant, and multi-subject training or shared stimuli add no benefit, so deep phenotyping is the recommended acquisition strategy.","lead":"The authors trained deep neural networks to decode what people listened to from their fMRI brain scans, reaching 27% average top-10 retrieval accuracy on three deeply scanned participants. Their results suggest that, when only a few participants can be scanned, collecting more hours per person helps more than scanning more people, which matters for designing fMRI speech-decoding studies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'deep phenotyping over more subjects' recommendation rests on a multi-subject negative result obtained with one specific architecture (subject-specific first layer, no functional alignment) and with the extra data for subjects 1–3 excluded; the claim is therefore narrower than the…","rationale":"The central positive claim—that retrieval accuracy scales with per-participant data—is well supported by Figure 2's within-subject curves and by the incremental setup analysis in Figure 3; I do not dispute it. The load-bearing step is the inference from multi-subject training does not improve accuracy to the recommendation that multi-subject natural-speech decoding will require deeper phenotyping or substantially larger cohorts. That inference needs the tested subject-specific-layer architecture to be representative of multi-subject learning. It is not: the architecture has a single subject-specific input layer and a shared backbone, with no functional alignment of the input spaces, and Figure 4 explicitly excludes the extra data for subjects 1–3, so the comparison runs at roughly 4 hours per subject rather than the 13.5-hour regime highlighted in the abstract. The authors themselves list functional alignment and hierarchical Bayesian models as untested alternatives in the Discussion, which is honest but also concedes that the central recommendation is narrower than the abstract implies. The reader's conditional verdict already captures this concern; no change is needed. Secondary issues (the 0.05 percent chance level should be about 0.5 percent for a 2000-candidate retrieval set, and Figure 4 plots the best over all subject combinations) should be corrected but do not affect the main argument. The concrete test—a functional-aligned multi-subject decoder at full per-subject data—would settle whether the negative result generalizes.","tokens_in":15793,"tokens_out":11499,"duration_ms":122379,"concrete_test":"On the LeBel dataset, train the same contrastive decoder in a multi-subject setup after first mapping all subjects into a common space with a functional alignment method (e.g., shared response model or fused unbalanced Gromov Wasserstein, Thual et al. 2022), using the full available data per subject (13.5 hours for subjects 1–3, about 4–6 hours for the rest); evaluate per-subject top-10 retrieval on the same test splits as Figure 2 and compare with the single-subject curves at matched per-subject data. If aligned multi-subject accuracy is not significantly above single-subject accuracy for any of the 8 subjects, the recommendation survives; if it is, the multi-subject negative result is method-specific and the central recommendation must be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing condition for the abstract claim that multi-subject speech decoding likely requires deeper phenotyping or a substantially larger cohort is that the tested multi-subject implementation is a fair test of what multi-subject training can do. It is not. Section 2.2 describes a single architecture: a subject-specific first layer followed by a shared backbone, with voxels selected independently per subject. This is a weak form of inter-subject alignment; it cannot adjudicate between single-subject and multi-subject learning generally. The authors explicitly concede in the Discussion that advanced functional alignment techniques or hierarchical Bayesian models could change the result. Moreover, Figure 4, the main evidence for the negative multi-subject result, was run without using the extra data available for the 3 first subjects, so it tests roughly 4 hours per subject, not the 13.5-hour deep-phenotyping regime in which the headline 27 percent accuracy is achieved. The contrastive decoder itself is credible and the per-subject scaling in Figure 2 is a real finding; the weakness is specifically the generalization from this SSL decoder did not benefit to multi-subject approaches likely require deeper phenotyping or larger cohorts.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates data acquisition strategies for fMRI-based decoding of natural speech, using the LeBel et al. (2023) dataset of 8 participants. The authors train deep neural networks with a contrastive objective to predict LLM text embeddings from fMRI activity and evaluate performance with top-10 retrieval accuracy. They report that single-subject decoders scale with the amount of training data per participant, achieving 27% average top-10 accuracy for the three deeply phenotyped subjects (max 36% for subject 3), which far exceeds their stated chance level and the BrainLLM baseline. They also report that multi-subject training does not improve decoding accuracy relative to single-subject models, that the degree of stimulus overlap across subjects has negligible effect, and that decoders appear to capture syntactic structure better than semantic content. Based on these findings, the authors recommend deep phenotyping over collecting data from more participants and conclude that leveraging multi-subject data for natural speech decoding likely requires deeper phenotyping or a substantially larger cohort.","tokens_in":15993,"tokens_out":4596,"duration_ms":46723,"significance":"If the results hold, the paper provides a useful empirical benchmark for fMRI speech decoding, demonstrating that contrastive learning with LLM embeddings can yield high retrieval accuracy from relatively few participants. The scaling result (Figure 2) is a concrete, falsifiable finding that supports the deep-phenotyping strategy in limited-cohort regimes. The comparisons between single-subject and multi-subject training and the stimulus-overlap experiment address practically important design questions. The paper uses a public dataset, applies proper session- and story-holdout cross-validation, and reports confidence intervals, which strengthens the credibility of the per-subject scaling claims. However, the central negative result on multi-subject training is currently narrower than the abstract claims, and the reported chance level contains a quantitative error that affects several interpretations.","major_comments":[{"comment":"The reported chance level of 0.05% is incorrect for the evaluation protocol. With a retrieval set of approximately 2000 chunks and top-10 accuracy, random chance is 10/2000 = 0.5%. The value 0.05% corresponds to a top-1 chance level. This error appears in the Figure 2 caption, the BrainLLM baseline discussion, and the Results section, where the baseline is described as 'far above the chance-level of 0.05%'. The authors should correct this value and recompute the corresponding statements about how much above chance the decoders perform, since the correct chance level is an order of magnitude higher.","section":"§2.1, Figure 2, §3.1"},{"comment":"The central negative result on multi-subject training does not support the abstract conclusion that 'leveraging multi-subject for natural speech decoding likely requires deeper phenotyping or a substantially larger cohort.' Figure 4 explicitly excludes the extra data available for subjects 1–3, so the multi-subject comparison is performed at roughly 4 hours of training data per subject, not in the 13.5-hour deep-phenotyping regime in which the headline 27% accuracy is achieved. The claim that multi-subject training does not improve decoding in the deep-phenotyping regime is therefore untested. The authors should either run the multi-subject experiment using the full deep-phenotyping data for subjects 1–3 or clearly restrict the conclusion to the lower-data regime.","section":"§3.2, Figure 4, Abstract"},{"comment":"The multi-subject architecture tested here (a subject-specific first layer followed by a shared backbone, with no explicit inter-subject alignment) is a limited instantiation of multi-subject learning. The Discussion and Limitations already concede that advanced functional alignment or hierarchical Bayesian models could change the result. Since the recommendation to favor deep phenotyping over larger cohorts rests on this negative evidence, the paper should either provide additional multi-subject evidence (e.g., with a functional-alignment method) or temper the conclusion to the specific architecture and preprocessing pipeline evaluated.","section":"§2.2, §3.2, §5"},{"comment":"In Figure 4, for each subject and each number of training subjects, the authors report the best accuracy obtained across all 255 combinations of subjects. Taking the maximum over combinations introduces an optimistic bias and makes the displayed 'best accuracy' curve difficult to compare with the single-subject baseline. The paper should report the distribution (e.g., mean and range across combinations) or justify why the best-case selection is the appropriate comparison, and discuss how this selection affects the conclusion that multi-subject training is not beneficial.","section":"§3.2, Figure 4"},{"comment":"The conclusion that the decoder 'better differentiates between syntactically dissimilar chunks than semantically dissimilar ones' is based on comparing the normalized slopes of two different similarity metrics (Levenshtein distance on POS-tagged chunks vs. GloVe bag-of-words cosine). These metrics have different scales, distributions, and noise properties, so the observed difference in profile steepness could be an artifact of the metrics rather than a property of the decoder. The authors should provide a more direct and controlled comparison, for example by using a single representational space or by calibrating the two metrics on matched random baselines, before drawing the syntactic-vs-semantic conclusion.","section":"§3.3.2, Figure 6"}],"minor_comments":[{"comment":"In the sentence 'We do not to explicitly model inter-subject variability', there is a typo: 'do not to' should be 'do not'.","section":"§2.1"},{"comment":"The x-axis label '0 103 30 min' is unclear; it should be formatted as '10^3' and the corresponding time annotation should be explicit (e.g., '~30 min').","section":"Figure 2"},{"comment":"The statement 'we achieve an average top-10 accuracy of 27%' should explicitly specify that this average is over subjects 1, 2, and 3, and it would be helpful to report the per-subject values for all eight participants in a table rather than only in figures.","section":"§3.1"},{"comment":"The claim 'To our knowledge, these results represent the first successful decoding of natural speech from fMRI data using a contrastive objective' is strong. The authors should verify this claim against prior fMRI-based contrastive decoding work and, if it stands, provide a more precise statement of what specifically is novel.","section":"§3.1"},{"comment":"The row 'Total number of layers 3 Subject specific layer + Linear + Residual' is ambiguous; rephrase to make the layer structure clear, e.g., '3 layers: one subject-specific layer, one linear layer, one residual layer'.","section":"Table 2"},{"comment":"The Limitations section correctly acknowledges that the findings may not hold for larger cohorts, but this acknowledgement is not carried into the abstract or conclusion. The abstract should be aligned with the scope of the evidence presented in the paper.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid empirical core: the contrastive decoding pipeline and the per-participant scaling result are valuable, and the cross-validation design (held-out sessions and stories) is appropriate. The main problems are fixable: the incorrect chance level is a quantitative error, and the multi-subject negative result is presented as stronger than the experiments warrant. The authors should either add the deep-phenotyping multi-subject experiment or substantially narrow the conclusions. I also recommend that the editor ask for the code and trained model predictions to be made available in the revision, since the paper promises code 'upon publication' but the exact retrieval sets and preprocessing would be useful for reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: the per-subject scaling result is real and useful, but the headline 'deep phenotyping over more subjects' is a reasonable guess rather than a proven conclusion. The multi-subject null was obtained with a specific architecture (subject-specific first layer, shared backbone), without functional alignment, and Figure 4 does not use the extra data available for subjects 1–3. So the claim is narrower than the abstract suggests.\n\nWhat's new and good: first contrastive decoding of natural speech from fMRI on the LeBel dataset, with a systematic comparison of single- vs multi-subject training, stimulus overlap, and data quantity. The evaluation is careful: held-out runs, cross-validation, retrieval sets roughly matched. The BrainLLM baseline comparison is honest and useful.\n\nSoft spots: the reported 0.05% chance level is wrong by an order of magnitude—with a 2000-chunk retrieval set it should be about 0.5%. That error appears several times and makes the gains look larger than they are. Figure 4 plots the best accuracy over all 255 subject combinations, which is selection-biased; the average or a pre-specified combination would be the fair test. More substantively, the multi-subject conclusion rests on one architecture and on data subsampled to ~4 hours per subject, so it does not test whether multi-subject learning helps in the deep-phenotyping regime that produces the headline 27%. The authors do concede in the Discussion that functional alignment or hierarchical Bayesian models could change things, but the abstract's phrasing is too strong. The stimulus-overlap experiment is also limited to subjects 1–3 with reduced per-subject data, so 'negligible effect' is provisional.\n\nWho this is for: researchers planning fMRI speech-decoding studies or budgeting acquisition time. It deserves a serious referee—the core results are interesting and the issues are fixable. I'd ask for a corrected chance level, a less cherry-picked multi-subject figure, and a more cautious abstract. Code is promised but not yet out, so independent verification is limited, but the methods are described in enough detail to reproduce.","headline":"Solid contrastive-decoding results for natural speech fMRI, but the deep-phenotyping recommendation gets ahead of the evidence: the multi-subject null is one architecture, one data subset.","tokens_in":16585,"tokens_out":2490,"would_cite":true,"duration_ms":27525,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that for decoding natural speech from fMRI in small cohorts, deep phenotyping—more recorded hours per participant—outperforms adding more participants, and that multi-subject training with a shared backbone yields no…","keywords":["fMRI decoding","natural speech","deep phenotyping","contrastive learning","LLM embeddings","multi-subject decoding","retrieval accuracy","syntax and semantics"],"falsifier":"Re-run the multi-subject comparison on the same eight participants with functional alignment or a hierarchical Bayesian model instead of subject-specific linear layers, using the same held-out test stories: if per-subject top-10 accuracy beats the single-subject decoders for most participants, the paper's central negative result is an artefact of its architecture. Independently, extend the three deeply phenotyped participants past 13.5 recorded hours and check whether top-10 accuracy keeps rising; a plateau would falsify the claim that per-participant data volume still has headroom.","tokens_in":15548,"feed_emoji":"🧠","tokens_out":10497,"duration_ms":106911,"temperature":0.7,"pith_summary":"Researchers planning fMRI studies of language face a practical tradeoff: scan a few people for many hours or many people for a few hours each. This paper argues that when the goal is decoding the natural speech a person hears, the first option wins in the small-cohort regime. Training deep networks to predict language-model text embeddings from brain activity with a contrastive retrieval objective yields 27% average top-10 accuracy for the three participants with the most data (36% for the best), far above a 1.6% prior baseline and near-zero chance. Accuracy rises with recorded hours per participant and shows no plateau at 13.5 hours, whereas adding other participants to training gives no improvement and the amount of story overlap across participants does not matter. The authors conclude that inter-subject variability is the binding constraint, so acquisition should favour deep phenotyping unless cohorts become substantially larger.","feed_headline":"For fMRI speech decoding, hours per person beat participant count","feed_subtitle":"Single-subject decoders hit 27% top-10 accuracy after 13.5 hours; adding subjects gave no gain.","key_machinery":"The load-bearing mechanism is contrastive retrieval training of a brain decoder on LLM2Vec chunk embeddings, evaluated as top-10 accuracy against a retrieval set of about two thousand chunks per test fold. The contrastive loss pulls the predicted embedding toward the true chunk and pushes it away from other chunks in the same batch, and the pipeline's gains come from four tuned components: a 6-second lag for the haemodynamic delay, an 8-second context of preceding text, averaging 4 brain volumes for temporal smoothing, and pre-selecting 4096 voxels by their Ridge encoding $R^2$. For the multi-subject comparison, the central object is a decoder with a subject-specific first linear layer followed by a shared backbone, tested at hidden widths 64 and 4096; this architecture is what the no-multi-subject-gain claim rests on.","core_discovery":"The central discovery is a practical scaling result for neural speech decoding: with eight participants and up to 13.5 hours of story-listening fMRI per person, per-participant data volume drives retrieval accuracy, whereas adding subjects does not. A decoder trained per subject with a contrastive objective to predict LLM2Vec chunk embeddings from 4096 encoding-selected voxels achieves 27% average top-10 accuracy for the three most-recorded participants and 36% for subject 3, against a 1.6% BrainLLM baseline and near-zero chance. In the multi-subject setup, a shared backbone with subject-specific first layers fails to beat single-subject decoders across all 255 subsets of the eight participants, and the overlap of stories heard during training has negligible effect. The authors further report that decoders discriminate syntactic structure more sharply than semantic content, and that stories with conversational, simple syntax are decoded far better than reflective, complex ones. From this they conclude that the limiting factor is inter-subject variability, so acquisition should favour deep phenotyping unless cohorts become substantially larger.","pith_inferences":["A cost-budget reading not stated in the paper: if fixed total acquisition hours are the constraint, the data favour concentrating them into as few participants as possible, at least until a per-person saturation point is measured; this is testable by simulating scan-budget allocation with this dataset's scaling curve.","The multi-subject null result is likely tied to the subject-specific-layer architecture; functional alignment or hierarchical Bayesian sharing remains a plausible route to cross-subject gains, so the strongest general claim is about today's architectures, not about multi-subject learning in principle.","The negligible overlap effect implies that multi-site and cross-language pooling could be a practical path out of the limited-participant regime, turning the bottleneck from participant recruitment into embedding alignment across languages and recording sites.","The syntax-over-semantics asymmetry could be an artefact of the LLM2Vec representation rather than a property of neural speech signals; comparing decoders trained on syntax-probing versus semantics-probing targets would separate these explanations."],"forward_implications":["Per-participant recorded hours are the main driver of decoding accuracy in small cohorts: 13.5 hours of training data give 27% average top-10 accuracy, versus 6% for roughly 4 hours.","Adding more participants to a shared-backbone decoder will not by itself improve individual decoding accuracy in this regime, so acquisition plans should not trade per-person hours for cohort size.","Because story overlap across participants barely matters, experimenters can use diverse or even different stimuli per participant, and datasets recorded in different languages could be pooled if the text embeddings are language-agnostic.","Since decoders separate syntactic structure better than semantic content, improving decoding of reflective, semantically dense stories likely needs new objectives or representations rather than simply more of the same data.","Accuracy has not plateaued at 13.5 hours, so extending the same participants with more listening sessions is expected to yield further gains before any multi-subject advantage appears."],"supporting_citations":[{"why":"This citation supplies the eight-participant natural speech fMRI dataset, with about 6 hours of stories per person and 16.5 hours for three subjects; every decoding result rests on this data.","marker":"LeBel et al. (2023)"},{"why":"This citation provides the BrainLLM baseline and the closest prior decoder on the same dataset; its predicted embeddings yield the 1.6% average baseline used for comparison.","marker":"Ye et al. (2025)"},{"why":"This citation provides LLM2Vec, the 4096-dimensional sentence embeddings that are both the prediction target and the retrieval ground truth.","marker":"BehnamGhader et al. (2024)"},{"why":"This citation supplies the contrastive loss formulation that the decoders optimize for retrieval.","marker":"Radford et al. (2021)"},{"why":"This citation supplies the decoder architecture inspiration (layer normalization and skip connections) and the shared-model-with-per-subject-layers design.","marker":"Scotti et al. (2024)"},{"why":"This citation is the earlier language-decoding approach on this dataset using beam search with an encoding model, which the paper contrasts with its decoding-first method.","marker":"Tang et al. (2023)"},{"why":"This citation is cited as evidence that functional alignment boosts multi-subject decoding elsewhere, framing why the paper's null multi-subject result is architecture-specific.","marker":"Thual et al. (2023)"}],"fun_headline_variants":["Deep phenotyping beats cohort size for fMRI speech decoding","For fMRI speech decoding, hours per head beat head count","Single-subject fMRI decoders outperform multi-subject for speech","Per-brain data volume, not participant count, drives fMRI speech decoding"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion that multi-subject training does not help assumes that the subject-specific-layer-plus-shared-backbone decoder fairly represents multi-subject learning; if functional alignment or hierarchical Bayesian sharing works better, the recommendation to prefer deep phenotyping over larger cohorts loses its main support.","fun_headline_variants_meta":{"raw":{"variants":["Deep phenotyping beats cohort size for fMRI speech decoding","For fMRI speech decoding, hours per head beat head count","Single-subject fMRI decoders outperform multi-subject for speech","Per-brain data volume, not participant count, drives fMRI speech decoding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000469,"raw_usage":{"total_tokens":2320,"prompt_tokens":912,"completion_tokens":1408,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":1339}},"tokens_in":528,"tokens_out":1408,"duration_ms":11187,"temperature":1.0,"reasoning_tokens":1339,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:30:50.350765+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the multi-subject comparison on the same eight participants with functional alignment or a hierarchical Bayesian model instead of subject-specific linear layers, using the same held-out test stories: if per-subject top-10 accuracy beats the single-subject decoders for most participants, the paper's central negative result is an artefact of its architecture. Independently, extend the three deeply phenotyped participants past 13.5 recorded hours and check whether top-10 accuracy keeps rising; a plateau would falsify the claim that per-participant data volume still has headroom.","supporting_citations":[],"review_version":1}