REVIEW 3 major objections 5 minor 15 references
Contextual Semantic Relevance and Word Surprisal Predict N400 and P600 Dynamics During Naturalistic Reading
T0 review · 3 major / 5 minor · reviewed 2026-07-11 · grok-4.5
Pith's one-line read Local semantic fit predicts N400 and P600 EEG responses beyond word surprisal in naturalistic reading.
desk verdict Solid DERCo reanalysis showing a fixed local semantic-fit score adds incremental N400/P600 variance beyond GPT-2 surprisal; useful, not a clean prediction-vs-integration dissociation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Attention-aware contextual semantic relevance: a fixed weighted sum of cosine similarities between a target word embedding and its three preceding context words (weights 0.3, 0.6, 0.9) plus pairwise context-context similarities (weight 0.2), treated as a continuous predictor alongside surprisal in rERP and GAMM analyses.
What would settle it
Recompute the same channel-wise GAMMs and ΔAIC comparisons after replacing the fixed 0.3/0.6/0.9 + 0.2 weights and three-word window with data-driven or longer-context alternatives (or after proper prestimulus baselining); if semantic relevance no longer improves fit beyond surprisal in the P600 window across most channels, the central claim fails.
Extended reading notes
Core claim
Contextual semantic relevance contributes unique explanatory value for word-locked N400- and especially P600-window EEG voltages during naturalistic reading, beyond GPT-based surprisal and standard lexical controls, and it shows partly distinct temporal and scalp patterns from surprisal.
Load-bearing premise
The claim rests on treating a hand-weighted three-word local similarity score, and uncorrected window mean voltages from released RSVP epochs, as valid indexes of discourse semantic fit and of N400/P600 dynamics.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper tests whether an attention-aware contextual semantic relevance metric, defined over a three-word local context with fixed a priori weights, predicts word-locked EEG voltages in N400 (300–500 ms) and P600 (~500–800 ms) windows during naturalistic RSVP reading in the public DERCo corpus, beyond GPT-2 surprisal and lexical controls. Across 22 participants and 32 channels, the authors combine time-resolved rERP regressions with channel-wise GAMMs that include participant and token random-effect smooths, FDR correction, and ΔAIC model comparisons. Both predictors are associated with EEG activity and show partly distinct temporal and scalp patterns; semantic relevance is reported as especially robust for P600-window voltages and as adding explanatory value beyond surprisal. The central claim is that naturalistic reading depends on both lexical expectation and local semantic integration, and that the proposed metric provides an interpretable computational link between discourse semantic fit and ERP dynamics.
Significance. If the incremental P600 (and N400) effects survive stronger controls for residual RSVP overlap and metric parameterization, the paper would be a useful contribution to computational neurolinguistics: it jointly models surprisal and a transparent local semantic-fit measure on public naturalistic EEG, uses complementary rERP and GAMM pipelines with item-level random effects, and reports weak predictor correlation (r = −0.10) plus ΔAIC evidence. Strengths include open DERCo data, dual analysis strategies, lexical controls, FDR correction, and an explicitly fixed rather than EEG-fitted weighting scheme. The work is relevant to ongoing debates about N400/P600 functional interpretation and sequential vs. parallel processing, but its theoretical reach depends on whether the metric indexes discourse-level fit rather than short-range lexical similarity or presentation-mode carry-over.
major comments (3)
- Methods 3.1–3.2: Dependent variables are uncorrected window mean voltages from released DERCo epochs (baseline=None; no additional prestimulus correction), justified by possible residual activity from the preceding RSVP word. With 200 ms word + 300 ms blank presentation, successive epochs heavily overlap the 300–500 and 500–800 ms analysis windows, so residual carry-over can inflate late-window associations. The incremental P600 claim is load-bearing for the paper’s strongest conclusion; without a sensitivity analysis (e.g., prestimulus baseline, alternative late windows, or overlap-aware deconvolution), it remains unclear whether semantic relevance predicts true late integration or residual overlap correlated with local semantic structure.
- Methods 3.4 and Appendix A / Eq. for Sem_Rel: Contextual semantic relevance is a fixed a priori local kernel over only three preceding words with hand-set weights (0.3/0.6/0.9 target–context; 0.2 context–context), not estimated from EEG and not varied. The manuscript repeatedly frames this as discourse semantic fit, yet the operationalization is short-range lexical-semantic similarity plus local coherence. Because the central claim is that this metric contributes beyond surprisal especially in the P600 window (Results 4.2–4.3; Discussion 5.1–5.2), the paper needs either (i) robustness checks over window size/weights or (ii) a clearer, more modest interpretation as local semantic fit rather than discourse-level integration.
- Results 4.2–4.3 and Discussion 5.3–5.4: Channel-wise significance counts and ROI partial-effect plots are used to argue partly distinct temporal/scalp roles for surprisal vs. semantic relevance and to support graded sequential/parallel processing claims. These analyses are descriptive scalp-level models under volume conduction; they do not establish functional or source dissociation. The interpretive leap from broader P600 channel counts / larger ΔAIC for semantic relevance to “later integration and discourse updating” should be tempered, or supported by stronger spatiotemporal modeling and explicit alternative-window tests.
minor comments (5)
- Abstract/Introduction vs. Methods 3.4: “discourse context” / “discourse semantic fit” language is stronger than the three-word local metric actually computed; align terminology throughout.
- Section 3.2 vs. 3.5: P600 window is stated as 500–800 ms in quantification but 500–699 ms in the GAMM methods paragraph; reconcile the exact sample range.
- Table 4 / Appendix D: Some ΔAIC values are near zero or negative; the text should more carefully distinguish relative dominance from positive model-improvement evidence.
- Figure captions and Appendix C: Several figure references and channel-count statements for surprisal N400 effects are slightly inconsistent across main text and appendix; clean for consistency.
- Related Work / self-citation: Prior semantic-relevance papers are appropriately cited for metric development, but the novelty claim relative to Frank & Willems (2017) and Michaelov et al. (2024) cosine-context baselines could be stated more sharply.
Circularity Check
No load-bearing circularity: the semantic-relevance score is defined independently of EEG with fixed a priori weights; self-citations only supply the metric recipe, not the N400/P600 result.
-
ansatz smuggled in via citation
[§2.1 Related Work; §3.4 Computing semantic relevance; Appendix A]
"Because this “attention-aware” method has shown better performance than earlier cosine and dynamic approaches, the present study adopts it to compute contextual semantic relevance. ... The weights used in the metric were fixed a priori rather than estimated from the EEG data. The same graded-weighting approach has been used in previous work on eye movements and reading times (Sun et al., 2023; Sun and Liu, 2025; Sun et al., 2026)."
The 3-word window and fixed weights (0.3/0.6/0.9 target–context; 0.2 context–context) are not derived from EEG or from an independent uniqueness result; they are imported from the authors’ prior behavioral ansatz (forgetting-curve / attention-aware weighting). This is mild ansatz inheritance via self-citation. It does not make the EEG associations tautological: weights were not fit to DERCo voltages, and the N400/P600 claims are tested as out-of-sample empirical associations rather than forced by the definition of the score.
full rationale
The paper’s central claim is empirical: an independently computed local semantic-fit score and GPT-2 surprisal jointly explain word-locked N400- and P600-window voltages on DERCo, with semantic relevance adding ΔAIC beyond surprisal and lexical controls. The semantic-relevance formula (weighted cosines over three preceding words plus pairwise context terms) does not contain EEG voltages, N400/P600 labels, or any parameter estimated from the present ERP data; the paper states explicitly that weights were fixed a priori. Surprisal is taken from an external language model. Dependent measures are window means from released epochs; model comparisons are standard nested rERP/GAMM tests. Those steps are not self-definitional and are not fitted-then-predicted on the same target. The only mild circularity-adjacent pattern is methodological self-citation: the attention-aware window and hand-set weights are adopted because the same authors’ prior eye-movement/reading-time work used them. That is normal inheritance of an ansatz, not a uniqueness theorem or a derivation that forces the EEG outcome. The ERP associations remain externally falsifiable on DERCo and would change if the metric failed. Score 2 reflects that single non-load-bearing self-citation chain; the main prediction is not circular by construction.
Assumptions & free parameters
free parameters (5)
- target-context distance weights (0.3, 0.6, 0.9)
- context-context pairwise weight (0.2)
- local context window size (3 preceding words)
- GAMM smooth basis dimension k=5
- N400/P600 analysis windows (300-500 ms; ~500-800 ms)
assumptions (5)
- domain assumption Word-locked mean voltages in fixed post-onset windows index N400- and P600-related processing during RSVP reading without additional prestimulus baseline correction.
- domain assumption GPT-2 token-summed surprisal is an adequate computational proxy for lexical expectation.
- ad hoc to paper fastText cosine similarities with distance-sensitive local weighting measure contextual semantic fit relevant to comprehension.
- standard math Participant and token-level random-effect smooths adequately control repeated measures and item dependence.
- domain assumption Scalp-channel and ROI patterns can be interpreted as descriptive distributions without source localization.
invented entities (1)
-
attention-aware contextual semantic relevance (Sem_Rel) score
independent evidence
Cite this review
Pith. "Pith review of Contextual Semantic Relevance and Word Surprisal Predict N400 and P600 Dynamics During Naturalistic Reading." pith.science (2026). https://pith.science/paper/BYOXQSXL
@misc{pith2026260704107,
author = {Pith},
title = {Pith review of: Contextual Semantic Relevance and Word Surprisal Predict N400 and P600 Dynamics During Naturalistic Reading},
year = {2026},
howpublished = {\url{https://pith.science/paper/BYOXQSXL}},
note = {Machine review of arXiv:2607.04107}
}
read the original abstract
Word surprisal is a well-established computational predictor of human neural responses during language comprehension, but it remains less clear whether local semantic fit explains neural response variation beyond lexical expectation during naturalistic reading. Using the Dublin EEG-based Reading Experiment Corpus (DERCo), this study examined whether contextual semantic relevance predicts word-locked EEG activity in the N400 and P600 windows. Contextual semantic relevance was computed as an attention-aware measure of how strongly a target word is semantically connected to its recent discourse context, and it was compared with GPT-based word surprisal. Across 22 participants and 32 EEG channels, we tested both predictors using regression-based ERP analyses and generalized additive mixed models while controlling for lexical variables and repeated observations. Both predictors were reliably associated with EEG responses, but they showed partly different temporal and scalp-level patterns. Surprisal captured expectancy-related variation, whereas contextual semantic relevance showed robust effects across N400- and P600-window mean voltages, with particularly strong explanatory support in the P600 window. Model comparisons indicated that contextual semantic relevance contributed explanatory value beyond lexical controls and surprisal. These findings suggest that naturalistic reading depends on both lexical expectation and local semantic integration, and that contextual semantic relevance offers an interpretable computational link between discourse semantic fit and ERP dynamics.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Hale, John. (2001). A probabilistic Earley parser as a psych olinguistic model. The Second Meeting of the North American Chapter of the Assoc iation for Computational Linguistics. Levy, Roger. (2008). Expectation-based syntactic compreh ension. Cognition, 106(3), 1126–1177. Anderson, John R, Matessa, Michael, Lebiere, Christian. (1 997). ACT-R: A theory o...
2001
-
[2]
Boston, Marisa Ferrara, Hale, John, Kliegl, Reinhold, Pati l, Umesh, Va- sishth, Shravan. (2008). Parsing costs as predictors of rea ding difficulty: An evaluation using the Potsdam Sentence Corpus. Journal of Eye Move- ment Research, 2(1). Gibson, Edward. (1998). Linguistic complexity: Locality o f syntactic depen- dencies. Cognition, 68(1), 1–76. Roland, ...
2008
-
[3]
Goldstein, Ariel, Zada, Zaid, Buchnik, Eliav, Schain, Mari ano, Price, Amy, Aubrey, Bobbi, Nastase, Samuel A, Feder, Amir, Emanuel, Dot an, Cohen, Alon, others. (2022). Shared computational principles for language pro- cessing in humans and deep language models. Nature neurosci ence, 25(3), 369–380. Frank, Stefan L, Willems, Roel M. (2017). Word predictab...
arXiv 2022
-
[4]
DeLong, Katherine A, Quante, Laura, Kutas, Marta. (2014). P redictability, plausibility, and two late ERP positivities during written sentence com- prehension. Neuropsychologia, 61, 150–162. Veldre, Aaron, Andrews, Sally. (2016). Is semantic preview benefit due to relatedness or plausibility?. Journal of Experimental Psy chology: Human Perception and Perfo...
2014
-
[5]
Jabeen, Shahida, Gao, Xiaoying, Andreae, Peter. (2020). Se mantic associa- tion computation: a comprehensive survey. Artificial Intel ligence Review, 53(6), 3849–3899. 40 Harispe, Sébastien, Ranwez, Sylvie, Montmain, Jacky, othe rs. (2022). Se- mantic similarity from natural language and ontology analy sis. Springer Nature. Pereira, Francisco, Lou, Bin, Pr...
2020
-
[6]
Broderick, Michael P, Anderson, Andrew J, Di Liberto, Giova nni M, Crosse, Michael J, Lalor, Edmund C. (2018). Electrophysiological c orrelates of se- mantic dissimilarity reflect the comprehension of natural, narrative speech. Current Biology, 28(5), 803–809. Sun, Kun, Liu, Haitao. (2025). Attention-aware semantic relevance predicting Chinese sentence rea...
-
[7]
Van Petten, Cyma, Kutas, Marta. (1990). Interactions betwe en sentence con- text and word frequencyinevent-related brainpotentials. Memory & cogni- tion, 18, 380–393. Federmeier, Kara D, Kutas, Marta. (1999). A rose by any other name: Long- term memory structure and sentence processing. Journal of M emory and Language, 41(4), 469–495. Hagoort, Peter, Hald...
1990
-
[8]
Regel, Stefanie, Meyer, Lars, Gunter, Thomas C. (2014). Dis tinguishing neu- rocognitive processes reflected by P600 effects: Evidence fr om ERPs and neural oscillations. PloS One, 9(5), e96840. Schuster, Sarah, Hawelka, Stefan, Hutzler, Florian, Kronb ichler, Martin, Richlan, Fabio. (2016). Words in context: The effects of leng th, frequency, and predictabi...
2014
Show all 15 references
-
[9]
42 Gramfort, Alexandre, Luessi, Martin, Larson, Eric, Engema nn, Denis A, Strohmeier, Daniel, Brodbeck, Christian, Parkkonen, Laur i, Hämäläinen, Matti S. (2014). MNE software for processing MEG and EEG data . neu- roimage, 86, 446–460. Alday, Phillip M. (2019). How much basel...
2014
-
[10]
Lewis, Richard L, Vasishth, Shravan. (2005). An activation -based model of sentence processing as skilled memory retrieval. Cognitiv e science, 29(3), 375–419. Christiansen, Morten H, Chater, Nick. (2016). The now-or-n ever bottleneck: A fundamental constraint on language. Beh...
2005 arXiv
-
[11]
43 Wood, Simon N. (2017). Generalized Additive Models: An Intr oduction with R. Chapman and Hall/CRC. Wood, Simon N, Pya, Natalya, Säfken, Benjamin. (2016). Smoo thing pa- rameter and model selection for general smooth models. Jour nal of the American Statistical Association, ...
2017
-
[12]
Herbert H
doi: 10.1016/j.rmal.2023.100079. Herbert H. Clark. (1973). The language-as-fixed-effect fall acy: A critique of language statistics in psychological research. Journal of Verbal Learning and Verbal Behavior, 12(4), 335-359. R.H. Baayen, D.J. Davidson, D.M. Bates. (2008). Mixed-eff...
2023 doi
-
[13]
44 Kutas, Marta, Federmeier, Kara D. (2000). Electrophysiolo gy reveals seman- tic memory use in language comprehension. Trends in cogniti ve sciences, 4(12), 463–470. DeLong, Katherine A, Urbach, Thomas P, Kutas, Marta. (2005) . Proba- bilistic word pre-activation during lang...
2000
- [14]
-
[15]
attention-aware
5 Sum the weighted target-context similarities and the weig hted context- context similarities to obtain the contextual semantic rel evance score. Appendix Appendix A: Attention-aware approach and its memory capability This section details the rationale behind the attention-aw...
1985
Reviewed July 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.