Pith. sign in

REVIEW 3 major objections 5 minor 15 references

Contextual Semantic Relevance and Word Surprisal Predict N400 and P600 Dynamics During Naturalistic Reading

T0 review · 3 major / 5 minor · reviewed 2026-07-11 · grok-4.5

Pith's one-line read Local semantic fit predicts N400 and P600 EEG responses beyond word surprisal in naturalistic reading.

desk verdict Solid DERCo reanalysis showing a fixed local semantic-fit score adds incremental N400/P600 variance beyond GPT-2 surprisal; useful, not a clean prediction-vs-integration dissociation. read the letter →

arxiv 2607.04107 v2 pith:BYOXQSXL submitted 2026-07-05 cs.CL

classification cs.CL
keywords naturalisticreadingwordsurprisalcontextualsemanticrelevanceN400P600rERPGAMMEEG
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether the brain tracks not only how unexpected a word is, but also how well that word fits the recent discourse meaning. Using word-locked EEG from naturalistic RSVP reading, the authors compare GPT-based word surprisal with a simple attention-aware score of contextual semantic relevance built from fastText similarities over the three preceding words. Across 22 readers and 32 channels, both predictors relate to voltages in the classic N400 and later P600 windows after lexical controls, but their time courses and scalp patterns partly diverge. Semantic relevance remains informative even when surprisal is already in the model, and model comparisons give it especially strong support in the P600 window. The result matters because it supplies an interpretable computational link between discourse-level semantic fit and continuous ERP dynamics, supporting the view that naturalistic reading depends on both lexical expectation and local semantic integration rather than expectation alone.

What carries the argument

Attention-aware contextual semantic relevance: a fixed weighted sum of cosine similarities between a target word embedding and its three preceding context words (weights 0.3, 0.6, 0.9) plus pairwise context-context similarities (weight 0.2), treated as a continuous predictor alongside surprisal in rERP and GAMM analyses.

What would settle it

Recompute the same channel-wise GAMMs and ΔAIC comparisons after replacing the fixed 0.3/0.6/0.9 + 0.2 weights and three-word window with data-driven or longer-context alternatives (or after proper prestimulus baselining); if semantic relevance no longer improves fit beyond surprisal in the P600 window across most channels, the central claim fails.

Watch

Extended reading notes

Core claim

Contextual semantic relevance contributes unique explanatory value for word-locked N400- and especially P600-window EEG voltages during naturalistic reading, beyond GPT-based surprisal and standard lexical controls, and it shows partly distinct temporal and scalp patterns from surprisal.

Load-bearing premise

The claim rests on treating a hand-weighted three-word local similarity score, and uncorrected window mean voltages from released RSVP epochs, as valid indexes of discourse semantic fit and of N400/P600 dynamics.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper tests whether an attention-aware contextual semantic relevance metric, defined over a three-word local context with fixed a priori weights, predicts word-locked EEG voltages in N400 (300–500 ms) and P600 (~500–800 ms) windows during naturalistic RSVP reading in the public DERCo corpus, beyond GPT-2 surprisal and lexical controls. Across 22 participants and 32 channels, the authors combine time-resolved rERP regressions with channel-wise GAMMs that include participant and token random-effect smooths, FDR correction, and ΔAIC model comparisons. Both predictors are associated with EEG activity and show partly distinct temporal and scalp patterns; semantic relevance is reported as especially robust for P600-window voltages and as adding explanatory value beyond surprisal. The central claim is that naturalistic reading depends on both lexical expectation and local semantic integration, and that the proposed metric provides an interpretable computational link between discourse semantic fit and ERP dynamics.

Significance. If the incremental P600 (and N400) effects survive stronger controls for residual RSVP overlap and metric parameterization, the paper would be a useful contribution to computational neurolinguistics: it jointly models surprisal and a transparent local semantic-fit measure on public naturalistic EEG, uses complementary rERP and GAMM pipelines with item-level random effects, and reports weak predictor correlation (r = −0.10) plus ΔAIC evidence. Strengths include open DERCo data, dual analysis strategies, lexical controls, FDR correction, and an explicitly fixed rather than EEG-fitted weighting scheme. The work is relevant to ongoing debates about N400/P600 functional interpretation and sequential vs. parallel processing, but its theoretical reach depends on whether the metric indexes discourse-level fit rather than short-range lexical similarity or presentation-mode carry-over.

major comments (3)
  1. Methods 3.1–3.2: Dependent variables are uncorrected window mean voltages from released DERCo epochs (baseline=None; no additional prestimulus correction), justified by possible residual activity from the preceding RSVP word. With 200 ms word + 300 ms blank presentation, successive epochs heavily overlap the 300–500 and 500–800 ms analysis windows, so residual carry-over can inflate late-window associations. The incremental P600 claim is load-bearing for the paper’s strongest conclusion; without a sensitivity analysis (e.g., prestimulus baseline, alternative late windows, or overlap-aware deconvolution), it remains unclear whether semantic relevance predicts true late integration or residual overlap correlated with local semantic structure.
  2. Methods 3.4 and Appendix A / Eq. for Sem_Rel: Contextual semantic relevance is a fixed a priori local kernel over only three preceding words with hand-set weights (0.3/0.6/0.9 target–context; 0.2 context–context), not estimated from EEG and not varied. The manuscript repeatedly frames this as discourse semantic fit, yet the operationalization is short-range lexical-semantic similarity plus local coherence. Because the central claim is that this metric contributes beyond surprisal especially in the P600 window (Results 4.2–4.3; Discussion 5.1–5.2), the paper needs either (i) robustness checks over window size/weights or (ii) a clearer, more modest interpretation as local semantic fit rather than discourse-level integration.
  3. Results 4.2–4.3 and Discussion 5.3–5.4: Channel-wise significance counts and ROI partial-effect plots are used to argue partly distinct temporal/scalp roles for surprisal vs. semantic relevance and to support graded sequential/parallel processing claims. These analyses are descriptive scalp-level models under volume conduction; they do not establish functional or source dissociation. The interpretive leap from broader P600 channel counts / larger ΔAIC for semantic relevance to “later integration and discourse updating” should be tempered, or supported by stronger spatiotemporal modeling and explicit alternative-window tests.
minor comments (5)
  1. Abstract/Introduction vs. Methods 3.4: “discourse context” / “discourse semantic fit” language is stronger than the three-word local metric actually computed; align terminology throughout.
  2. Section 3.2 vs. 3.5: P600 window is stated as 500–800 ms in quantification but 500–699 ms in the GAMM methods paragraph; reconcile the exact sample range.
  3. Table 4 / Appendix D: Some ΔAIC values are near zero or negative; the text should more carefully distinguish relative dominance from positive model-improvement evidence.
  4. Figure captions and Appendix C: Several figure references and channel-count statements for surprisal N400 effects are slightly inconsistent across main text and appendix; clean for consistency.
  5. Related Work / self-citation: Prior semantic-relevance papers are appropriately cited for metric development, but the novelty claim relative to Frank & Willems (2017) and Michaelov et al. (2024) cosine-context baselines could be stated more sharply.

Circularity Check

1 steps flagged · score 2.0 of 10

No load-bearing circularity: the semantic-relevance score is defined independently of EEG with fixed a priori weights; self-citations only supply the metric recipe, not the N400/P600 result.

  1. ansatz smuggled in via citation [§2.1 Related Work; §3.4 Computing semantic relevance; Appendix A]
    "Because this “attention-aware” method has shown better performance than earlier cosine and dynamic approaches, the present study adopts it to compute contextual semantic relevance. ... The weights used in the metric were fixed a priori rather than estimated from the EEG data. The same graded-weighting approach has been used in previous work on eye movements and reading times (Sun et al., 2023; Sun and Liu, 2025; Sun et al., 2026)."

    The 3-word window and fixed weights (0.3/0.6/0.9 target–context; 0.2 context–context) are not derived from EEG or from an independent uniqueness result; they are imported from the authors’ prior behavioral ansatz (forgetting-curve / attention-aware weighting). This is mild ansatz inheritance via self-citation. It does not make the EEG associations tautological: weights were not fit to DERCo voltages, and the N400/P600 claims are tested as out-of-sample empirical associations rather than forced by the definition of the score.

full rationale

The paper’s central claim is empirical: an independently computed local semantic-fit score and GPT-2 surprisal jointly explain word-locked N400- and P600-window voltages on DERCo, with semantic relevance adding ΔAIC beyond surprisal and lexical controls. The semantic-relevance formula (weighted cosines over three preceding words plus pairwise context terms) does not contain EEG voltages, N400/P600 labels, or any parameter estimated from the present ERP data; the paper states explicitly that weights were fixed a priori. Surprisal is taken from an external language model. Dependent measures are window means from released epochs; model comparisons are standard nested rERP/GAMM tests. Those steps are not self-definitional and are not fitted-then-predicted on the same target. The only mild circularity-adjacent pattern is methodological self-citation: the attention-aware window and hand-set weights are adopted because the same authors’ prior eye-movement/reading-time work used them. That is normal inheritance of an ansatz, not a uniqueness theorem or a derivation that forces the EEG outcome. The ERP associations remain externally falsifiable on DERCo and would change if the metric failed. Score 2 reflects that single non-load-bearing self-citation chain; the main prediction is not circular by construction.

Assumptions & free parameters 5 free parameters · 5 assumptions · 1 invented entities

The central claim rests on standard psycholinguistic ERP assumptions plus several paper-specific modeling choices: a fixed local three-word context, hand-set distance weights, embedding cosine as semantic fit, GPT-2 surprisal as expectancy, uncorrected window means as N400/P600 size, and additive mixed models with chosen smooth basis dimension. No new physical entity is postulated; the invented construct is the specific attention-aware relevance score as an ERP predictor.

free parameters (5)
  • target-context distance weights (0.3, 0.6, 0.9)
    Fixed a priori graded weights for wt-3, wt-2, wt-1; not estimated from the EEG data but load-bearing for the relevance score.
  • context-context pairwise weight (0.2)
    Fixed weight on all three preceding-word pairwise similarities; chosen by authors rather than fit or derived.
  • local context window size (3 preceding words)
    Hard-coded short window motivated by memory/RSVP considerations; alternative windows would change the predictor.
  • GAMM smooth basis dimension k=5
    Chosen flexibility for continuous predictors; affects estimated nonlinear partial effects.
  • N400/P600 analysis windows (300-500 ms; ~500-800 ms)
    Predefined component windows used as dependent measures; endpoints vary slightly across sections.
assumptions (5)
  • domain assumption Word-locked mean voltages in fixed post-onset windows index N400- and P600-related processing during RSVP reading without additional prestimulus baseline correction.
    Methods 3.1-3.2 explicitly analyze released uncorrected epochs as window means.
  • domain assumption GPT-2 token-summed surprisal is an adequate computational proxy for lexical expectation.
    Standard in the field; used as the main expectancy predictor in 3.3.
  • ad hoc to paper fastText cosine similarities with distance-sensitive local weighting measure contextual semantic fit relevant to comprehension.
    Core definition of Sem_Rel in 3.4/Appendix A; weights and 3-word window are paper-chosen.
  • standard math Participant and token-level random-effect smooths adequately control repeated measures and item dependence.
    Standard mixed-model assumption invoked for primary GAMMs.
  • domain assumption Scalp-channel and ROI patterns can be interpreted as descriptive distributions without source localization.
    Stated repeatedly due to volume conduction.
invented entities (1)
  • attention-aware contextual semantic relevance (Sem_Rel) score independent evidence
    purpose: Provide an interpretable local semantic-fit predictor of N400/P600 dynamics complementary to surprisal.
    Defined by a specific weighted combination of target-context and context-context cosines; not a new brain mechanism, but a paper-central computational construct.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Contextual Semantic Relevance and Word Surprisal Predict N400 and P600 Dynamics During Naturalistic Reading." pith.science (2026). https://pith.science/paper/BYOXQSXL

@misc{pith2026260704107,
  author       = {Pith},
  title        = {Pith review of: Contextual Semantic Relevance and Word Surprisal Predict N400 and P600 Dynamics During Naturalistic Reading},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BYOXQSXL}},
  note         = {Machine review of arXiv:2607.04107}
}
read the original abstract

Word surprisal is a well-established computational predictor of human neural responses during language comprehension, but it remains less clear whether local semantic fit explains neural response variation beyond lexical expectation during naturalistic reading. Using the Dublin EEG-based Reading Experiment Corpus (DERCo), this study examined whether contextual semantic relevance predicts word-locked EEG activity in the N400 and P600 windows. Contextual semantic relevance was computed as an attention-aware measure of how strongly a target word is semantically connected to its recent discourse context, and it was compared with GPT-based word surprisal. Across 22 participants and 32 EEG channels, we tested both predictors using regression-based ERP analyses and generalized additive mixed models while controlling for lexical variables and repeated observations. Both predictors were reliably associated with EEG responses, but they showed partly different temporal and scalp-level patterns. Surprisal captured expectancy-related variation, whereas contextual semantic relevance showed robust effects across N400- and P600-window mean voltages, with particularly strong explanatory support in the P600 window. Model comparisons indicated that contextual semantic relevance contributed explanatory value beyond lexical controls and surprisal. These findings suggest that naturalistic reading depends on both lexical expectation and local semantic integration, and that contextual semantic relevance offers an interpretable computational link between discourse semantic fit and ERP dynamics.

Figures

Figures reproduced from arXiv: 2607.04107 by the authors.

Figure 1
Figure 1. The EEG Topomap with regions was measured in the 500–800 ms interval and is referred to here as the P600- window amplitude. This label denotes a late post-onset positivity window rather than a condition-defined P600 effect. As described in Section 3.1, the released DERCo epochs were used without additional prestimulus baseline correction. The dependent variables in this analysis are therefore uncorrected mean voltag… view at source ↗
Figure 2
Figure 2. Computation of contextual semantic relevance. The figure illustrates the computation of semantic relevance for the target word “water” in the sentence “you should give me some water”. The three preceding words, “give” (vt−3), “me” (vt−2), and ‘some” (vt−1), form the local context. Seman￾tic relevance is computed as a weighted sum of cosine similarities between the target word and these context words, with closer wor… view at source ↗
Figure 3
Figure 3. Raw rERP window topographies for semantic relevan [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: Representative ROI-level rERP beta waveforms wit [PITH_FULL_IMAGE:figures/full_fig_p021_4.png]
Figure 5
Figure 5. Figure 5: Channel-wise GAMM partial effects of contextual se [PITH_FULL_IMAGE:figures/full_fig_p024_5.png]
Figure 6
Figure 6. Figure 6: ROI-level GAMM partial effects for N400 and P600 siz [PITH_FULL_IMAGE:figures/full_fig_p026_6.png]
Figure 7
Figure 7. Figure 7: GAMM partial effects of contextual semantic releva [PITH_FULL_IMAGE:figures/full_fig_p027_7.png]
Figure 8
Figure 8. Figure 8: Computation of semantic relevance and its memory c [PITH_FULL_IMAGE:figures/full_fig_p049_8.png]
Figure 9
Figure 9. Figure 9: Correlation heatmap for the predictors in EEG data [PITH_FULL_IMAGE:figures/full_fig_p050_9.png]
Figure 10
Figure 10. Figure 10: ROI-level rERP waveforms with 95% confidence inte [PITH_FULL_IMAGE:figures/full_fig_p052_10.png]
Figure 11
Figure 11. Figure 11: Sliding-time analysis showing the number of sign [PITH_FULL_IMAGE:figures/full_fig_p053_11.png]
Figure 12
Figure 12. Figure 12: Overlap and dominance maps for contextual semant [PITH_FULL_IMAGE:figures/full_fig_p054_12.png]
Figure 13
Figure 13. Figure 13: Comparison between full and reduced rERP models. [PITH_FULL_IMAGE:figures/full_fig_p055_13.png]
Figure 14
Figure 14. Figure 14: GAMM partial effects of surprisal on N400 size acro [PITH_FULL_IMAGE:figures/full_fig_p056_14.png]
Figure 15
Figure 15. Figure 15: Additional channel-wise GAMM partial effects of s [PITH_FULL_IMAGE:figures/full_fig_p057_15.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

15 extracted references · 2 canonical work pages

  1. [1]

    Hale, John. (2001). A probabilistic Earley parser as a psych olinguistic model. The Second Meeting of the North American Chapter of the Assoc iation for Computational Linguistics. Levy, Roger. (2008). Expectation-based syntactic compreh ension. Cognition, 106(3), 1126–1177. Anderson, John R, Matessa, Michael, Lebiere, Christian. (1 997). ACT-R: A theory o...

  2. [2]

    Boston, Marisa Ferrara, Hale, John, Kliegl, Reinhold, Pati l, Umesh, Va- sishth, Shravan. (2008). Parsing costs as predictors of rea ding difficulty: An evaluation using the Potsdam Sentence Corpus. Journal of Eye Move- ment Research, 2(1). Gibson, Edward. (1998). Linguistic complexity: Locality o f syntactic depen- dencies. Cognition, 68(1), 1–76. Roland, ...

  3. [3]

    Goldstein, Ariel, Zada, Zaid, Buchnik, Eliav, Schain, Mari ano, Price, Amy, Aubrey, Bobbi, Nastase, Samuel A, Feder, Amir, Emanuel, Dot an, Cohen, Alon, others. (2022). Shared computational principles for language pro- cessing in humans and deep language models. Nature neurosci ence, 25(3), 369–380. Frank, Stefan L, Willems, Roel M. (2017). Word predictab...

  4. [4]

    DeLong, Katherine A, Quante, Laura, Kutas, Marta. (2014). P redictability, plausibility, and two late ERP positivities during written sentence com- prehension. Neuropsychologia, 61, 150–162. Veldre, Aaron, Andrews, Sally. (2016). Is semantic preview benefit due to relatedness or plausibility?. Journal of Experimental Psy chology: Human Perception and Perfo...

  5. [5]

    Jabeen, Shahida, Gao, Xiaoying, Andreae, Peter. (2020). Se mantic associa- tion computation: a comprehensive survey. Artificial Intel ligence Review, 53(6), 3849–3899. 40 Harispe, Sébastien, Ranwez, Sylvie, Montmain, Jacky, othe rs. (2022). Se- mantic similarity from natural language and ontology analy sis. Springer Nature. Pereira, Francisco, Lou, Bin, Pr...

  6. [6]

    Broderick, Michael P, Anderson, Andrew J, Di Liberto, Giova nni M, Crosse, Michael J, Lalor, Edmund C. (2018). Electrophysiological c orrelates of se- mantic dissimilarity reflect the comprehension of natural, narrative speech. Current Biology, 28(5), 803–809. Sun, Kun, Liu, Haitao. (2025). Attention-aware semantic relevance predicting Chinese sentence rea...

  7. [7]

    Van Petten, Cyma, Kutas, Marta. (1990). Interactions betwe en sentence con- text and word frequencyinevent-related brainpotentials. Memory & cogni- tion, 18, 380–393. Federmeier, Kara D, Kutas, Marta. (1999). A rose by any other name: Long- term memory structure and sentence processing. Journal of M emory and Language, 41(4), 469–495. Hagoort, Peter, Hald...

  8. [8]

    Regel, Stefanie, Meyer, Lars, Gunter, Thomas C. (2014). Dis tinguishing neu- rocognitive processes reflected by P600 effects: Evidence fr om ERPs and neural oscillations. PloS One, 9(5), e96840. Schuster, Sarah, Hawelka, Stefan, Hutzler, Florian, Kronb ichler, Martin, Richlan, Fabio. (2016). Words in context: The effects of leng th, frequency, and predictabi...

Show all 15 references
  1. [9]

    42 Gramfort, Alexandre, Luessi, Martin, Larson, Eric, Engema nn, Denis A, Strohmeier, Daniel, Brodbeck, Christian, Parkkonen, Laur i, Hämäläinen, Matti S. (2014). MNE software for processing MEG and EEG data . neu- roimage, 86, 446–460. Alday, Phillip M. (2019). How much basel...

  2. [10]

    Lewis, Richard L, Vasishth, Shravan. (2005). An activation -based model of sentence processing as skilled memory retrieval. Cognitiv e science, 29(3), 375–419. Christiansen, Morten H, Chater, Nick. (2016). The now-or-n ever bottleneck: A fundamental constraint on language. Beh...

  3. [11]

    43 Wood, Simon N. (2017). Generalized Additive Models: An Intr oduction with R. Chapman and Hall/CRC. Wood, Simon N, Pya, Natalya, Säfken, Benjamin. (2016). Smoo thing pa- rameter and model selection for general smooth models. Jour nal of the American Statistical Association, ...

  4. [12]

    Herbert H

    doi: 10.1016/j.rmal.2023.100079. Herbert H. Clark. (1973). The language-as-fixed-effect fall acy: A critique of language statistics in psychological research. Journal of Verbal Learning and Verbal Behavior, 12(4), 335-359. R.H. Baayen, D.J. Davidson, D.M. Bates. (2008). Mixed-eff...

  5. [13]

    44 Kutas, Marta, Federmeier, Kara D. (2000). Electrophysiolo gy reveals seman- tic memory use in language comprehension. Trends in cogniti ve sciences, 4(12), 463–470. DeLong, Katherine A, Urbach, Thomas P, Kutas, Marta. (2005) . Proba- bilistic word pre-activation during lang...

  6. [14]

    Van Herten, Marieke, Kolk, Herman HJ, Chwilla, Dorothee J. ( 2005). An ERP study of P600 effects elicited by semantic anomalies. Cog nitive Brain Research, 22(2), 241–255. Tiedt, Hannes O, Ehlen, Felicitas, Klostermann, Fabian. (2 020). Age-related dissociation of N400 effect an...

  7. [15]

    attention-aware

    5 Sum the weighted target-context similarities and the weig hted context- context similarities to obtain the contextual semantic rel evance score. Appendix Appendix A: Attention-aware approach and its memory capability This section details the rationale behind the attention-aw...

Pith tools

Reviewed July 11, 2026 · model on record in the stance chip above.