{"id":"2a515582-10c6-4d9b-b739-b0c437e1331b","arxiv_id":"2607.04107","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Attention-aware contextual semantic relevance predicts N400 and especially P600 EEG voltages during naturalistic reading beyond GPT-2 surprisal and lexical controls.","lead":"This EEG study finds that an attention-weighted measure of how well a word fits its recent context predicts N400 and P600 voltages during naturalistic reading, beyond GPT-2 surprisal and lexical controls. It matters because it offers a simple, interpretable computational link between discourse semantic fit and classic language ERP components.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Uncorrected window means plus fixed 3-word weights leave the incremental P600 claim vulnerable to residual overlap and arbitrary local-fit parameterization.","rationale":"The reader correctly isolates the fixed 3-word hand-weighted metric and the uncorrected window means as the weakest assumptions supporting the central incremental-explanatory-value claim. Those choices are not peripheral: they define both the predictor that is said to index “discourse semantic fit” and the ERP quantities said to index N400/P600 dynamics. The dual rERP + GAMM pipeline, lexical controls, participant/token random effects, and weak surprisal correlation are genuine strengths and make a pure null result unlikely, so the paper remains publishable-shaped under a CONDITIONAL verdict. No stronger internal inconsistency (e.g., contradictory model comparisons or circular use of EEG to set weights) is present. The concrete test above directly probes whether the reported P600 advantage for semantic relevance is an artifact of residual overlap or of the particular local kernel; if it survives, the claim is substantially more secure; if not, the interpretation must be narrowed to local semantic similarity under RSVP. This does not require rejecting the paper, only retaining the reader’s caution about how strongly one can conclude about naturalistic discourse integration.","tokens_in":30069,"tokens_out":730,"duration_ms":6727,"concrete_test":"Re-run the primary channel-wise P600 GAMMs (full vs reduced excluding semantic relevance) after (a) applying a conventional –200–0 ms baseline correction to the released epochs and (b) recomputing semantic relevance under at least two alternative schemes (uniform weights; 5-word window with the same graded decay). If the number of FDR-significant channels for semantic relevance and the median ΔAIC drop substantially (e.g., below the corresponding surprisal values or lose the reported P600 advantage), the incremental discourse-integration claim is not robust to these modeling choices.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim is that contextual semantic relevance adds explanatory value for N400/P600-window voltages beyond GPT-2 surprisal and lexical controls, especially in the P600 window (Results 4.2–4.3; Discussion 5.1–5.2). That claim rests on two jointly load-bearing choices in Methods 3.1–3.4 and Appendix A: (1) dependent variables are uncorrected mean voltages from released DERCo epochs (baseline=None; no additional prestimulus correction), justified by possible residual activity from the preceding RSVP word in the –200–0 ms interval; and (2) semantic relevance is a fixed a-priori local score over only the three preceding words with hand-set weights (0.3/0.6/0.9 target–context; 0.2 context–context), not estimated from EEG and not varied. Because successive 200 ms + 300 ms RSVP epochs heavily overlap the 300–500 and 500–800 ms analysis windows, residual carry-over can inflate late-window associations. Because the metric is a short, fixed-weight local kernel rather than a discourse-level representation, significant ΔAIC and channel counts may reflect local lexical-semantic similarity (or residual overlap correlated with it) rather than discourse semantic fit. The paper reports only weak linear correlation with surprisal (r = –0.10) and channel-wise/ROI significance, but does not show that the incremental P600 support survives alternative windows, baselines, or weightings.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"This paper tests whether an attention-aware contextual semantic relevance metric, defined over a three-word local context with fixed a priori weights, predicts word-locked EEG voltages in N400 (300–500 ms) and P600 (~500–800 ms) windows during naturalistic RSVP reading in the public DERCo corpus, beyond GPT-2 surprisal and lexical controls. Across 22 participants and 32 channels, the authors combine time-resolved rERP regressions with channel-wise GAMMs that include participant and token random-effect smooths, FDR correction, and ΔAIC model comparisons. Both predictors are associated with EEG activity and show partly distinct temporal and scalp patterns; semantic relevance is reported as especially robust for P600-window voltages and as adding explanatory value beyond surprisal. The central claim is that naturalistic reading depends on both lexical expectation and local semantic integration, and that the proposed metric provides an interpretable computational link between discourse semantic fit and ERP dynamics.","tokens_in":30451,"tokens_out":1222,"duration_ms":9648,"significance":"If the incremental P600 (and N400) effects survive stronger controls for residual RSVP overlap and metric parameterization, the paper would be a useful contribution to computational neurolinguistics: it jointly models surprisal and a transparent local semantic-fit measure on public naturalistic EEG, uses complementary rERP and GAMM pipelines with item-level random effects, and reports weak predictor correlation (r = −0.10) plus ΔAIC evidence. Strengths include open DERCo data, dual analysis strategies, lexical controls, FDR correction, and an explicitly fixed rather than EEG-fitted weighting scheme. The work is relevant to ongoing debates about N400/P600 functional interpretation and sequential vs. parallel processing, but its theoretical reach depends on whether the metric indexes discourse-level fit rather than short-range lexical similarity or presentation-mode carry-over.","major_comments":[{"comment":"Methods 3.1–3.2: Dependent variables are uncorrected window mean voltages from released DERCo epochs (baseline=None; no additional prestimulus correction), justified by possible residual activity from the preceding RSVP word. With 200 ms word + 300 ms blank presentation, successive epochs heavily overlap the 300–500 and 500–800 ms analysis windows, so residual carry-over can inflate late-window associations. The incremental P600 claim is load-bearing for the paper’s strongest conclusion; without a sensitivity analysis (e.g., prestimulus baseline, alternative late windows, or overlap-aware deconvolution), it remains unclear whether semantic relevance predicts true late integration or residual overlap correlated with local semantic structure.","section":null},{"comment":"Methods 3.4 and Appendix A / Eq. for Sem_Rel: Contextual semantic relevance is a fixed a priori local kernel over only three preceding words with hand-set weights (0.3/0.6/0.9 target–context; 0.2 context–context), not estimated from EEG and not varied. The manuscript repeatedly frames this as discourse semantic fit, yet the operationalization is short-range lexical-semantic similarity plus local coherence. Because the central claim is that this metric contributes beyond surprisal especially in the P600 window (Results 4.2–4.3; Discussion 5.1–5.2), the paper needs either (i) robustness checks over window size/weights or (ii) a clearer, more modest interpretation as local semantic fit rather than discourse-level integration.","section":null},{"comment":"Results 4.2–4.3 and Discussion 5.3–5.4: Channel-wise significance counts and ROI partial-effect plots are used to argue partly distinct temporal/scalp roles for surprisal vs. semantic relevance and to support graded sequential/parallel processing claims. These analyses are descriptive scalp-level models under volume conduction; they do not establish functional or source dissociation. The interpretive leap from broader P600 channel counts / larger ΔAIC for semantic relevance to “later integration and discourse updating” should be tempered, or supported by stronger spatiotemporal modeling and explicit alternative-window tests.","section":null}],"minor_comments":[{"comment":"Abstract/Introduction vs. Methods 3.4: “discourse context” / “discourse semantic fit” language is stronger than the three-word local metric actually computed; align terminology throughout.","section":null},{"comment":"Section 3.2 vs. 3.5: P600 window is stated as 500–800 ms in quantification but 500–699 ms in the GAMM methods paragraph; reconcile the exact sample range.","section":null},{"comment":"Table 4 / Appendix D: Some ΔAIC values are near zero or negative; the text should more carefully distinguish relative dominance from positive model-improvement evidence.","section":null},{"comment":"Figure captions and Appendix C: Several figure references and channel-count statements for surprisal N400 effects are slightly inconsistent across main text and appendix; clean for consistency.","section":null},{"comment":"Related Work / self-citation: Prior semantic-relevance papers are appropriately cited for metric development, but the novelty claim relative to Frank & Willems (2017) and Michaelov et al. (2024) cosine-context baselines could be stated more sharply.","section":null}],"recommendation":"major_revision","confidential_remarks":"The core design is publishable after revision, but the manuscript currently overclaims “discourse” integration relative to a fixed 3-word local kernel and underplays RSVP residual-overlap risk for the P600 result. I would not reject on novelty grounds—the dual rERP+GAMM pipeline on public data is useful—but I would require the sensitivity analyses or a clearly narrowed interpretation before acceptance at a serious journal."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: on public DERCo RSVP EEG, their attention-aware local semantic-relevance score (fastText cosines over the three preceding words with fixed 0.3/0.6/0.9 and 0.2 weights) predicts N400- and especially P600-window mean voltages beyond GPT-2 surprisal, length, and frequency, with only weak correlation to surprisal (r = –0.10). Dual analysis (rERP + channel-wise GAMMs with participant and token random effects, FDR, ΔAIC) is the right toolkit for this claim.\n\nWhat is actually new is the joint test on naturalistic discourse EEG with that particular metric, previously used mainly for eye movements. Surprisal–ERP and semantic-similarity–ERP links already exist (Frank, Broderick, Michaelov, etc.). The paper does the controls carefully enough that the incremental channel counts and ΔAIC, especially in the P600 window, look real rather than collinear noise. Self-citation of their prior semantic-relevance work is proportionate; the formula is a priori and not fit to these ERPs, so circularity is low.\n\nSoft spots are real but bounded. Dependent variables are uncorrected window means from released epochs (baseline=None), justified by residual prior-word activity in the –200–0 ms interval under 200+300 ms RSVP. That choice, plus heavy epoch overlap into 300–500 and 500–800 ms, leaves late-window associations open to residual carry-over. The metric is a short fixed-weight local kernel, not a discourse model, so “discourse semantic fit” is stronger language than the computation supports. Free parameters (window size, weights, k=5, windows) are not varied. No shipped analysis code. These limit how far one can push a clean prediction-vs-integration story; they do not erase the incremental predictive result on this dataset.\n\nThis is for people who already work on computational predictors of ERPs or naturalistic reading. It is not a methods breakthrough and not a theory closer. I would send it to peer review: the data, dual stats, and incremental claim are sharp enough to deserve referee time, with the baseline/window and local-vs-discourse caveats front and center. Worth engaging if you care about semantic predictors beyond surprisal; not required reading otherwise.","headline":"Solid DERCo reanalysis showing a fixed local semantic-fit score adds incremental N400/P600 variance beyond GPT-2 surprisal; useful, not a clean prediction-vs-integration dissociation.","tokens_in":31072,"tokens_out":572,"would_cite":true,"duration_ms":6286,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Local semantic fit predicts N400 and P600 EEG responses beyond word surprisal in naturalistic reading.","keywords":["naturalistic reading","word surprisal","contextual semantic relevance","N400","P600","rERP","GAMM","EEG"],"falsifier":"Recompute the same channel-wise GAMMs and ΔAIC comparisons after replacing the fixed 0.3/0.6/0.9 + 0.2 weights and three-word window with data-driven or longer-context alternatives (or after proper prestimulus baselining); if semantic relevance no longer improves fit beyond surprisal in the P600 window across most channels, the central claim fails.","tokens_in":30910,"feed_emoji":"🧠","tokens_out":616,"duration_ms":5289,"temperature":0.7,"pith_summary":"This paper asks whether the brain tracks not only how unexpected a word is, but also how well that word fits the recent discourse meaning. Using word-locked EEG from naturalistic RSVP reading, the authors compare GPT-based word surprisal with a simple attention-aware score of contextual semantic relevance built from fastText similarities over the three preceding words. Across 22 readers and 32 channels, both predictors relate to voltages in the classic N400 and later P600 windows after lexical controls, but their time courses and scalp patterns partly diverge. Semantic relevance remains informative even when surprisal is already in the model, and model comparisons give it especially strong support in the P600 window. The result matters because it supplies an interpretable computational link between discourse-level semantic fit and continuous ERP dynamics, supporting the view that naturalistic reading depends on both lexical expectation and local semantic integration rather than expectation alone.","feed_headline":"Semantic fit beats pure surprise in reading EEG signals","feed_subtitle":"Local discourse relevance explains N400 and P600 voltages beyond GPT surprisal","key_machinery":"Attention-aware contextual semantic relevance: a fixed weighted sum of cosine similarities between a target word embedding and its three preceding context words (weights 0.3, 0.6, 0.9) plus pairwise context-context similarities (weight 0.2), treated as a continuous predictor alongside surprisal in rERP and GAMM analyses.","core_discovery":"Contextual semantic relevance contributes unique explanatory value for word-locked N400- and especially P600-window EEG voltages during naturalistic reading, beyond GPT-based surprisal and standard lexical controls, and it shows partly distinct temporal and scalp patterns from surprisal.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Semantic relevance predicts N400 and P600 beyond GPT surprisal","Contextual semantic fit explains ERP voltages past lexical surprise","Discourse relevance adds unique value to N400-P600 reading signals","Local semantic fit shows distinct patterns from word surprisal in EEG","Attention-aware relevance captures P600 dynamics beyond surprisal"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The claim rests on treating a hand-weighted three-word local similarity score, and uncorrected window mean voltages from released RSVP epochs, as valid indexes of discourse semantic fit and of N400/P600 dynamics.","fun_headline_variants_meta":{"raw":{"variants":["Semantic relevance predicts N400 and P600 beyond GPT surprisal","Contextual semantic fit explains ERP voltages past lexical surprise","Discourse relevance adds unique value to N400-P600 reading signals","Local semantic fit shows distinct patterns from word surprisal in EEG","Attention-aware relevance captures P600 dynamics beyond surprisal"]},"model":"grok-4.5","effort":"low","cost_usd":0.004078,"raw_usage":{"total_tokens":1245,"prompt_tokens":752,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":40780000,"prompt_tokens_details":{"text_tokens":752,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":427,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":752,"tokens_out":66,"duration_ms":3613,"temperature":1.0,"reasoning_tokens":427,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T21:37:54.393851+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Recompute the same channel-wise GAMMs and ΔAIC comparisons after replacing the fixed 0.3/0.6/0.9 + 0.2 weights and three-word window with data-driven or longer-context alternatives (or after proper prestimulus baselining); if semantic relevance no longer improves fit beyond surprisal in the P600 window across most channels, the central claim fails.","supporting_citations":[],"review_version":1}