{"id":"89af49bb-b724-44e7-a9fc-87e565948583","arxiv_id":"1908.08807","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"An fMRI decoding method that corrects standard encoding predictions using a PCA-based latent 'inner state' estimated from the measured brain pattern, reporting higher image identification accuracy.","lead":"This paper adds a 'brain inner state' correction to standard fMRI encoding models, using leftover prediction errors from other brain voxels to improve how well a scanned image can be identified from brain activity. The idea is practical for brain decoding, but the way the authors measure improvement may be overly generous because it uses the brain response being predicted as part of the input.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validation-time inner state is estimated from true measured responses of other voxels (Sec. II-A), so reported encoding R2 gains—and possibly the decoding gains—reflect target leakage rather than predictive encoding.","rationale":"The reader's weakest assumption matches my reading: the phrase 'one needs to know other voxels' true activities' in Section II-A, combined with the validation R2 protocol in Section III-B.1, is the exact point where the argument breaks. The zero-model control is not merely a curiosity: R2=0.35 with zero stimulus features is the cleanest demonstration that the framework's encoding gain is achieved by feeding measured responses back into the prediction. I agree with REJECT. I would not soften the verdict based on the decoding results because the test-time procedure is not specified; however, the paper's core concept of modeling residual correlations is not inherently invalid, so a revised version with proper cross-validation and explicit test-time equations could be conditionally acceptable. No change to the reader's verdict is needed.","tokens_in":15874,"tokens_out":6831,"duration_ms":71456,"concrete_test":"Recompute the encoding evaluation with a no-leak inner state: fit the connectivity graph, PCA weights α_i, and λ_i on training residuals exactly as in Section II-A, but for each of the 120 validation images construct s_i for voxel i from the forward model's predicted activities of the connected voxels for that image (never from measured validation responses), then recompute mean R2 for GLO, GWP, RO, and ZM. If the ISF-minus-forward increments (0.28→0.42 for GLO; 0→0.35 for ZM) collapse, the original encoding claim is driven by target leakage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that ISF combines stimulus features with a brain inner state to predict fMRI responses and improve image identification. The load-bearing step is the inner-state estimate in Section II-A: s_i is the first principal component of residual vectors of 'connected' voxels, and the paper states that 'to estimate the inner state of one voxel, one needs to know other voxels' true activities.' In Section III-B.1, encoding R2 is computed on the 120 validation images. If the validation image's true activities of connected voxels are used to form s_i, then the ISF prediction v_i = f_i(X)β_i + s_iλ_i is not a stimulus-driven prediction: it uses the measured response of the same image through other voxels. The ISF+ZM result confirms this: the zero model has no image features, yet ISF+ZM reaches R2=0.35, which can only arise from correlations among measured validation activities. Therefore the reported encoding improvements (e.g., GLO R2 from 0.28 to 0.42) are not evidence for the framework's predictive content. The decoding section (II-C) is also underspecified: it says the inner-state model is 'applied to the predicted pattern and the measured pattern' but gives no equations. If the measured pattern is used to build the inner state for each candidate image, the identification accuracy can be inflated by self-similarity; the paper does not rule this out.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an encoding framework (ISF) that augments a conventional forward stimulus-response encoding model with a data-driven \"brain inner state\" term. The inner state is estimated as the first principal component of the forward-model residuals of voxels that are highly correlated with the target voxel. The authors evaluate the framework on the Kay et al. natural-image fMRI dataset, reporting that ISF improves mean encoding R2 (e.g., GLO from 0.28 to 0.42) and image identification accuracy (e.g., ISF+GLO 100% for subject 1) relative to forward-only models, and that even a zero forward model combined with ISF reaches R2=0.35. The paper also studies identification accuracy as a function of image-set size and voxel count. The central claim is that adding the inner-state term to any forward encoding model yields substantial, robust gains in both encoding and decoding.","tokens_in":16178,"tokens_out":2663,"duration_ms":28409,"significance":"If the reported results were valid, the paper would present a conceptually interesting extension of voxel-wise encoding models, with a plausible neuroscientific motivation from intrinsic connectivity and a practical gain in decoding performance. The use of a public dataset, four forward models, and a zero-model control is a commendable design. However, the encoding evaluation is fundamentally compromised by target leakage: the inner state for a validation image is constructed from the true measured activities of other voxels for that same image, so the reported R2 gains do not reflect stimulus-driven prediction. The zero-model result of R2=0.35 without any image features is a direct demonstration of this leakage. The decoding procedure is also underspecified, leaving open the same leakage concern. As a result, the paper's main claims are not supported, and the contribution cannot be accepted in its current form.","major_comments":[{"comment":"The encoding evaluation is circular. In Eqs. (5)-(6), the inner state s_i for a voxel is estimated as the first principal component of the residual vectors of its connected voxels, and the introduction explicitly states that 'to estimate the inner state of one voxel, one needs to know other voxels' true activities.' For the validation images, those true activities are the measured responses to the same image being predicted. Substituting this s_i into Eq. (2) means that each voxel's prediction uses the measured activity of other voxels from the same validation image. The reported encoding R2 values (e.g., ISF+GLO R2=0.42 vs. GLO R2=0.28, Fig. 3) therefore do not measure stimulus-driven predictive accuracy; they are inflated by information from the target image itself. This is a load-bearing flaw for the claim that ISF 'achieves much better performance' in encoding.","section":"Section II-A, Eqs. (5)-(6); Section III-B.1"},{"comment":"The zero-model control result is decisive evidence of target leakage. The zero model has no image features (Eq. 14), so ISF+ZM receives no stimulus information. Its reported mean R2 of 0.35 can only arise from correlations among the measured validation activities of connected voxels via Eqs. (5)-(6). This shows that the inner-state term, as implemented, is not extracting stimulus-independent brain dynamics; it is exploiting the measured response pattern of the same image. The paper itself notes that ISF+ZM 'still achieved a better predictive power of R2 = 0.35,' but the authors do not recognize that this invalidates rather than supports the framework's encoding claims.","section":"Fig. 3 and Section III-B.1"},{"comment":"The decoding procedure is underspecified in a way that matters for the identification results. The text says that the inner state model is 'applied to the predicted pattern and the measured pattern' to yield an updated prediction, but no equations are given for how the test-time inner state is computed. If the measured pattern (the one whose image identity is being sought) is used to estimate the inner state for each candidate image, then the updated prediction can be made artificially similar to the measured pattern for the correct image, inflating identification accuracy. The paper does not rule out this possibility, and hence the reported 100% and 95.83% identification accuracies for ISF+GLO are not interpretable without a precise statement of the test-time estimation procedure.","section":"Section II-C"}],"minor_comments":[{"comment":"The choices of the correlation threshold for defining connected voxels and the number of principal components retained are not reported; these free parameters are central to the method and should be stated explicitly or provided in a supplementary table.","section":"Section II-A"},{"comment":"The notation is inconsistent: Eq. (6) defines \\tilde{s}_i as the PCA-based estimate, Eq. (7) uses \\tilde{s}_i to compute \\tilde{\\lambda}_i, but Eq. (2) uses s_i without tildes. Please clarify the relationship between the training-time estimate and the validation-time estimate.","section":"Eq. (7)"},{"comment":"The scatter plots in the second column of Fig. 4 are very small and the axis labels overlap; larger panels or a density-color representation would make the voxel-level comparison readable.","section":"Fig. 4"},{"comment":"The phrase 'extern stimuli' should read 'external stimuli'.","section":"Section IV-D"},{"comment":"Reference [15] and [16] lack complete bibliographic information (volume, pages, and publication year are missing); please complete these entries.","section":"References"},{"comment":"The paper claims state-of-the-art performance but does not compare with previously published identification accuracies on the same Kay et al. dataset; a comparison table would strengthen the decoding claim.","section":"Section IV-B and Fig. 7"}],"recommendation":"reject","confidential_remarks":"The zero-model result (ISF+ZM R2=0.35) is a clear red flag that the encoding improvements are driven by using the true validation responses rather than by stimulus-based prediction. In my assessment, the central encoding claim is not fixable by local revisions; a proper evaluation would require a fundamentally different inner-state estimation procedure that does not use the target image's measured responses, and even then the decoding section would need to be rewritten with explicit equations. The idea of using residual structure may be salvageable in future work, but the current manuscript does not establish its validity. The paper's fit to the journal is acceptable in topic, but the technical flaw is decisive."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about this paper is that its central encoding claim does not survive a close look at the evaluation. The inner state for a validation image is estimated from the true measured activities of other voxels in that same image (Section II-A: 'To estimate the inner state of one voxel, one needs to know other voxels' true activities'). So the reported R2 gains — GLO from 0.28 to 0.42, and the zero model from 0 to 0.35 — are not evidence of stimulus-driven prediction. The zero-model result is a smoking gun: with no image features at all, ISF+ZM 'predicts' voxel responses with R2=0.35 purely by exploiting inter-voxel correlations in the validation data. That is target leakage, not encoding.\n\nWhat is genuinely new is the specific combination: taking PCA of forward-model residuals and using it as a correction term for natural-image identification. The authors also deserve credit for including the zero-model control at all — it is an honest thing to do, even if it undercuts their own claim. The paper is clearly written, uses the standard Kay et al. dataset, and the set-size scaling analysis is a useful addition.\n\nThe soft spots are not minor. First, the encoding evaluation is circular: each voxel's prediction uses other voxels' measured responses to the same image, so the R2 values do not measure how well the model generalizes from stimulus to activity. Second, the decoding procedure is underspecified — 'the inner state model is applied to the predicted pattern and the measured pattern' is not enough to know whether test-time leakage inflates the identification accuracies. Third, the method is not compared to existing residual-based or noise-normalization baselines, so even the decoding gains are not established as special. Fourth, the free parameters (connectivity threshold, number of PCs, voxel count) appear to be chosen on the validation set without nested cross-validation. Finally, the Discussion's claim that the inner state is 'linearly independent to image features' rests on orthogonality of training residuals, which does not hold for new images.\n\nThis paper is for someone working on fMRI decoding who might be interested in residual-based corrections, but only as a starting point — the current evaluation would need to be redone with proper stimulus-only encoding cross-validation and explicit test-time equations before the claims can be trusted. It deserves a serious referee rather than a desk reject, but as it stands the encoding result should not be accepted.","headline":"The ISF idea is a reasonable denoising heuristic, but the encoding evaluation leaks the target through other voxels' measured activity, and the zero-model control makes that plain.","tokens_in":16706,"tokens_out":3138,"would_cite":false,"duration_ms":34936,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding a brain 'inner state' term to fMRI encoding models lifts natural image identification to 100 percent.","keywords":["fMRI encoding","brain decoding","natural image identification","inner state model","voxel-wise encoding","functional connectivity","prediction residuals","PCA"],"falsifier":"Re-run ISF with the inner state for validation images estimated only from training-data residuals and the forward model's predictions (never from the validation image's measured pattern); if the encoding $R^2$ and identification gains over forward-only models largely vanish, the reported improvements depend on access to the measured response being predicted.","tokens_in":15624,"feed_emoji":"🧠","tokens_out":9517,"duration_ms":84421,"temperature":0.7,"pith_summary":"The paper sets out to show that fMRI voxel responses to natural images are better explained and decoded when a standard forward encoding model is augmented with an estimate of the brain's 'inner state.' It proposes a framework in which any forward model's per-voxel prediction is corrected by a linear term built from the residuals of other voxels, on the grounds that visual cortex is a state machine with intrinsic connections, not a passive sensor. On a two-subject dataset, the correction raises mean encoding $R^2$ from 0.28 to 0.42 for the best forward model and lifts natural-image identification accuracy from 95.83% to 100% on subject 1 and from 88.33% to 95.83% on subject 2. The paper also reports that the zero model, which receives no image features at all, reaches $R^2 = 0.35$ under the framework, which it reads as evidence that the residual structure carries real intrinsic information. If true, the framework would give a general way to improve any fMRI encoding model without redesigning its stimulus features, with potential value for brain-computer interfaces and neural image reconstruction.","feed_headline":"Inner-state model hits 100% on fMRI image identification","feed_subtitle":"Residual-based correction raises encoding R2 from 0.28 to 0.42 and beats forward-only decoding on natural images.","key_machinery":"The load-bearing object is the inner-state model built from prediction residuals. After fitting a forward model on training data, the paper forms a residual matrix $E = [\\epsilon_1, \\dots, \\epsilon_p]$ with $\\epsilon_i = v_i - f_i(X)\\tilde{\\beta}_i$; for each voxel, it selects 'connected' voxels whose residual vectors have Pearson correlation above a threshold, and sets the inner state $\\tilde{s}_i = E_i \\alpha_i$, the first principal component of those connected voxels' residuals. A scalar weight $\\tilde{\\lambda}_i$ is then fitted by least squares, and the final ISF prediction is $v_i = f_i(X)\\beta_i + \\tilde{s}_i \\tilde{\\lambda}_i$. In decoding, the inner-state correction is applied to both the forward model's predicted pattern and the measured pattern, and identification picks the image whose corrected predicted pattern has the highest Pearson correlation with the measured pattern. This machinery is what carries the paper's claim that residual structure—rather than noise—contains recoverable information about intrinsic brain state.","core_discovery":"Traditionally, a voxel-wise encoding model predicts each voxel's response as $v_i = f_i(X)\\beta_i + e_i$, treating the brain as a stimulus-response mapper. The paper's central discovery is that adding an inner-state term $s_i\\lambda_i$ to this equation—where $s_i$ is estimated as the first principal component of the prediction residuals of voxels whose residual time courses are highly correlated with voxel $i$—produces substantially higher encoding $R^2$ and identification accuracy than the forward model alone. Because the correction is built from residuals after the stimulus features have explained their portion, the paper argues it captures influence from intrinsic connections rather than from the image. The framework is deliberately agnostic to the forward model: Gabor wavelet pyramid, gross local orientation, retinotopy-only, and even a zero model can be plugged in. With effective forward models the corrected predictions also degrade less as the candidate image set grows from 100 to 1,000 images and remain accurate when pattern similarity is measured by Euclidean distance instead of Pearson correlation.","pith_inferences":["Editorial inference: because the inner state for a validation image is computed from that image's measured voxel activities, part of the reported encoding $R^2$ gain likely reflects the model seeing the target pattern through other voxels; a test that estimates inner state from training data alone would reveal how much.","Editorial inference: the zero-model result suggests the residual-based correction may also be absorbing shared noise or non-stimulus response components; applying ISF to image labels shuffled across the validation set would measure how much of the identification gain is specific to image content.","Editorial inference: the framework's flexibility implies it could be combined with deep-network features or applied to EEG/MEG single-trial decoding, where a measured spatiotemporal pattern is available to seed the inner-state estimate."],"forward_implications":["Plugging any of the four tested forward models into ISF yields a statistically significant encoding $R^2$ increase over the forward-only version (t test, p < 0.01).","ISF identification accuracy stays high as the candidate image set grows from 100 to 1,000 images, with a mean decline of 10.25%, 5.23%, and 17.44% for GWP, GLO, and RO versus 35.52%, 18.92%, and 51.51% for the forward-only versions.","ISF predictions remain close to measured patterns under Euclidean distance as well as Pearson correlation; for subject 1, ISF+GLO identification stays at 98.33% with Euclidean distance while forward-only GLO drops from 95.87% to 75%.","The zero-model ISF reaches $R^2 = 0.35$ without any image features, but cannot identify images above chance, indicating the inner-state term alone carries non-stimulus structure but not image identity."],"supporting_citations":[{"why":"Supplies the fMRI dataset, the Gabor wavelet pyramid and retinotopy-only forward models, and the image identification procedure that ISF extends.","marker":"[4]"},{"why":"Establishes the voxel-wise encoding/decoding formulation that the inner-state framework modifies.","marker":"[1]"},{"why":"Reports a mean encoding $R^2$ of 0.28 on natural images, used as a motivating baseline for low explained variance.","marker":"[10]"},{"why":"Reports a mean $R^2$ of 0.25 for deep neural network features, another baseline in the same motivation.","marker":"[11]"},{"why":"Provides resting-state functional connectivity evidence that intrinsic networks carry signal, used to justify the inner-state term.","marker":"[23]"},{"why":"Provides lasso regression used to fit the high-dimensional Gabor and retinotopy forward models.","marker":"[34]"}],"fun_headline_variants":["Inner-state encoding beats forward-only on fMRI images","Brain state model lifts image ID accuracy and R2","Residual-aware encoding sharpens natural image recall","Intrinsic brain signals improve fMRI image identification","Flexible encoding plus inner state boosts brain decoding"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework estimates a voxel's inner state in a validation image from the true measured activities of other voxels in that same image, so the prediction is not made from stimulus information alone.","fun_headline_variants_meta":{"raw":{"variants":["Inner-state encoding beats forward-only on fMRI images","Brain state model lifts image ID accuracy and R2","Residual-aware encoding sharpens natural image recall","Intrinsic brain signals improve fMRI image identification","Flexible encoding plus inner state boosts brain decoding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000211,"raw_usage":{"total_tokens":1416,"prompt_tokens":949,"completion_tokens":467,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":565,"completion_tokens_details":{"reasoning_tokens":396}},"tokens_in":565,"tokens_out":467,"duration_ms":5075,"temperature":1.0,"reasoning_tokens":396,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:41:23.436778+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run ISF with the inner state for validation images estimated only from training-data residuals and the forward model's predictions (never from the validation image's measured pattern); if the encoding $R^2$ and identification gains over forward-only models largely vanish, the reported improvements depend on access to the measured response being predicted.","supporting_citations":[{"cited_title":"Encoding and decoding in fMRI,","cited_arxiv_id":null,"evidence_quote":"Establishes the voxel-wise encoding/decoding formulation that the inner-state framework modifies."},{"cited_title":"Unsupervised Feature Learning Improves Prediction of Human Brain Activity in Response to Natural Images,","cited_arxiv_id":null,"evidence_quote":"Reports a mean encoding $R^2$ of 0.28 on natural images, used as a motivating baseline for low explained variance."},{"cited_title":"Deep neural networks reveal a gradient in the complexity of neural representations across the ventral stream,","cited_arxiv_id":null,"evidence_quote":"Reports a mean $R^2$ of 0.25 for deep neural network features, another baseline in the same motivation."},{"cited_title":"Functional connectivity in the resting brain: A network analysis of the default mode hypothesis,","cited_arxiv_id":null,"evidence_quote":"Provides resting-state functional connectivity evidence that intrinsic networks carry signal, used to justify the inner-state term."}],"review_version":1}