REVIEW 6 major objections 6 minor 14 references
The Predictive Brain: Neural Correlates of Word Expectancy Align with Large Language Model Prediction Probabilities
T0 review · 6 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper claims that during natural audiobook listening, the brain's word-expectancy signals—reflected in N400 amplitude and pre-onset activity—track BERT's continuous predictability scores rather than only a binary high/low distinction.
desk verdict A useful but statistically fragile replication of known N400-surprisal effects; the 'first naturalistic' claim is wrong, and the headline r² values are inflated by data-driven selection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a per-noun predictability score produced by masking each noun in the audiobook text and reading BERT's softmax probability for the original word (averaged over sub-word tokens), which is then compared with word-locked neural responses. The neural side is the N400 component—the well-known negative deflection around 300–500 ms after a word that indexes semantic processing effort—plus pre-onset anticipatory activity. Forced alignment supplies precise word-onset times, and the analysis correlates mean activity in ten equal-sized predictability bins with the log of predictability scores, so the test is whether brain responses scale gradually with probability rather than only differing between high and low bins.
What would settle it
A reader could re-run the ten-bin correlation on the same audiobook data with the parietal EEG electrodes and left-frontal MEG sensors fixed in advance, adding word frequency, word length, and acoustic prominence as covariates; if the EEG correlation at 300–450 ms and the MEG correlation at 500–650 ms no longer reach significance, the claim that BERT predictability specifically drives the effect would be refuted.
Extended reading notes
Core claim
The central discovery is a continuous inverse correlation between BERT predictability and N400 amplitude in naturalistic speech, together with a positive correlation between predictability and pre-word-onset neural activity. In EEG, the post-onset effect appears at 300–450 ms over parietal electrodes ($p = 0.0006$, $r^2 = 0.79$ over ten predictability bins), and in MEG at 500–650 ms over left frontal sensors ($p = 0.0013$, $r^2 = 0.75$). Pre-onset activity also scales with predictability (EEG at $-100$ to $0$ ms: $p = 0.028$, $r^2 = 0.47$; MEG at $-350$ to $-250$ ms: $p = 0.0315$, $r^2 = 0.46$), and this pre-activation is negatively related to the later N400 amplitude, linking anticipation to reduced processing cost. Source reconstruction places the post-onset difference in parietal and sensorimotor cortex and the EEG pre-onset effect in left fronto-temporal regions, consistent with a left-lateralized language network. The authors further claim this is the first demonstration of these effects with naturalistic speech stimuli.
Load-bearing premise
The correlations rest on the assumption that the selected electrodes, MEG sensors, and time windows were not chosen from the same data used to test them, and that BERT's probability captures human predictability without needing to separate word frequency, word length, or acoustic prominence.
Editorial extensions
If this is right
- If the claim holds, N400 amplitude in natural listening can be read as a graded neural surprisal signal, and language-model probabilities become a quantitative predictor of neural processing effort.
- The effect generalizing from controlled visual sentence reading to continuous audiobook speech means predictive-coding accounts of language are not limited to artificial stimulus protocols.
- The negative relation between pre-onset activity and N400 amplitude supports a direct anticipatory mechanism: stronger top-down pre-activation reduces bottom-up integration cost.
- Convergent EEG and MEG correlations imply the predictability effect is robust across measurement modalities and is not a single-sensor artifact.
- Because the correlation is computed over ten bins rather than a high/low split, the brain's response appears to track probability continuously, not just categorical expectancy violations.
Reading between the lines
- A natural next test, not run in the paper, is a mixed-effects model on single trials that includes word frequency, word length, and acoustic prominence as covariates; if the BERT-probability slope survives, the case for context-specific prediction is much stronger.
- The MEG tendency toward sensorimotor pre-onset engagement for low-predictability nouns suggests a testable motoric component of anticipation: compare nouns that differ in articulatory complexity while matching predictability.
- Because BERT's masked probabilities are bidirectional, they include future context; comparing with a causal left-to-right language model would clarify whether the brain's anticipatory signal respects temporal order.
- The pre-onset EEG window ($-100$ to $0$ ms) was inspected after the binary comparison showed no effect there, so a pre-registered replication with fixed windows is the cleanest way to confirm the anticipatory result.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an EEG/MEG study in which 29 participants listened to a German audiobook, and BERT-based predictability scores for nouns were correlated with neural responses. The central claims are that higher predictability is associated with reduced N400-like amplitudes (EEG: p=0.0006, r2=0.79 at 300–450 ms; MEG: p=0.0013, r2=0.75 at 500–650 ms) and with increased pre-onset activity (EEG: p=0.028, r2=0.47 at -100–0 ms; MEG: p=0.0315, r2=0.46 at -350 to -250 ms). These results are interpreted as evidence for continuous alignment between LLM word-expectancy probabilities and predictive brain processing in naturalistic speech, and the paper claims to be the first demonstration of such effects with naturalistic stimuli.
Significance. If robust, the central finding would extend previous N400-surprisal work to continuous, naturalistic listening with simultaneous EEG and MEG, and would strengthen the case that transformer-based language model probabilities provide a useful quantitative account of neural word-expectancy signals. The study has genuine strengths: a naturalistic stimulus, a plausible use of BERT-derived predictability, simultaneous EEG/MEG acquisition, and explicit attempts to check reproducibility with additional low-predictability subsets. However, the headline quantitative claims rest on analytic choices that were made after inspecting the same dataset that is then used for the reported p-values and r2 values. Because sensor selection, time-window selection, and the EEG pre-onset window were data-driven, the reported effect sizes are likely inflated, and the central claim that the brain's expectancy signals are 'continuously aligned' with BERT probabilities is not yet supported at the stated level of confidence. The qualitative direction of the N400 effect is consistent with prior literature, but the paper's specific quantitative evidence requires substantially stronger statistical validation.
major comments (6)
- [§2.3 and §3.3] The sensor and time-window selections used for the headline correlations are data-driven. In §2.3, the authors state that they 'identified peak responses in the time window from -1.0s until 2.0s and extracted the topographic distribution at the time point of the strongest negative responses,' and then selected parietal EEG electrodes and left frontal MEG sensors. In §3.3, the correlation analysis uses mean activity over these selected sensors and over time windows that were identified as significant in the same dataset. The reported p=0.0006/r2=0.79 and p=0.0013/r2=0.75 are therefore not out-of-sample estimates; selecting the spatial and temporal region that looks most favorable can inflate correlations substantially. The authors should either pre-specify sensors and time windows, or report results across the full sensor array and a broad fixed window, or use cross-validation/permutation procedures that account for the selection step.
- [§3.3] The EEG pre-onset window (-100 to 0 ms) is explicitly post hoc. The text states: 'in the EEG data, where no significant pre-onset effects were observed, we examined a pre-word interval at -100ms to assess early anticipatory processes.' This window was therefore tested after an initial null result, yet the reported p=0.028 is presented without correction for this additional test. The pre-onset EEG finding cannot be treated as confirmatory evidence; it should be labeled exploratory and the p-value should be adjusted, or an independent dataset/pre-registered analysis should be provided.
- [§3.3] The correlation analysis does not include any confounding regressors. In natural speech, BERT predictability is correlated with word frequency, word length, acoustic duration, speech rate, and possibly lexical neighborhood properties. The regression in §3.3 models only mean predictability per bin against mean neural activity, so the observed correlations could be driven by these confounds rather than by predictability per se. The authors should add control analyses that include word frequency, word length, and acoustic duration as covariates or as binning variables, or otherwise demonstrate that the predictability effect survives such controls.
- [§2.2 and §3.1] The handling of the class imbalance is not statistically transparent. The authors select a random subset of 194 low-predictability trials to match the 194 high-predictability trials, and then report that two additional low-predictability subsets show 'consistent' effects. However, no seed is given for the random selection, no distribution across many possible subsets is reported, and one of the additional EEG subsets shows only a trend (p=0.11), which is nonetheless interpreted as supporting reproducibility. The authors should report the full set of subset analyses, state whether they were preplanned, and quantify variability across random draws rather than selecting favorable subsets.
- [§2.3] The baseline correction used for the epoch analysis may materially affect the pre-onset findings. The epochs are baseline-corrected from -1.0 to 0.0 s, which includes the exact pre-onset intervals later analyzed (-350 to -250 ms and -100 to 0 ms). This means that the reported pre-onset 'activity' is measured relative to a baseline that already contains the anticipatory effect, potentially removing or distorting the very signal of interest. The authors should report whether the pre-onset correlations survive when using a baseline outside the analyzed window (e.g., -1000 to -600 ms) or when no baseline correction is applied.
- [§3.2] The source-space results are reported with strong language ('significant activity was observed') but no statistical test is described for the source reconstructions. It is unclear whether the source-space differences between low- and high-predictability nouns were tested with cluster-based permutation statistics, corrected for multiple comparisons across sources, or are purely descriptive. The authors need to specify the statistical procedure used in source space, or explicitly label these results as descriptive and not claim significance.
minor comments (6)
- [§3.3 and Figure 4] There is an inconsistency in the reported EEG post-stimulus p-value: the text states p=0.0006, while the Figure 4 caption states p=0.00006. The authors should correct this discrepancy and ensure all p-values/r2 values match between text and figures.
- [Introduction] Typo: 'many studies investigation predictability' should be 'many studies investigating predictability.'
- [§5 Acknowledgments] Typo: 'programmme' should be 'programme' (or 'program').
- [§2.2] The description of the semi-logarithmic scaling as 'reverses the softmax operation' is imprecise; taking the logarithm of a softmax probability does not invert the softmax mapping. The sentence should be rephrased as applying a log transform to approximate surprisal.
- [Discussion / Abstract] The claim that this is 'the first study to demonstrate these effects using naturalistic speech stimuli' is difficult to reconcile with the authors' own citation of Goldstein et al., who used naturalistic story listening and reported similar predictability-related neural responses around 400 ms. The novelty claim should be narrowed or supported by an explicit comparison of stimuli and analyses.
- [General] No data or code availability statement is provided, and the analysis pipeline is described only briefly. Given the data-driven nature of several analytic choices, sharing the analysis code and processed trial-level data would be important for assessing reproducibility.
Circularity Check
No definitional circularity: the BERT-neural correlations are empirical fits, not derivations from their inputs.
full rationale
I walked the claimed derivation chain and found no step in which a predicted quantity is defined in terms of the data used to test it in the sense of an equation or fitted parameter being renamed as a prediction. The BERT predictability scores are computed from the audio book text independently of the EEG/MEG recordings, and the neural responses are measured directly; the reported correlations are ordinary regression fits to binned mean amplitudes. The methodology does contain post hoc selections: parietal electrodes and left frontal sensors were chosen from the topographic peak of the same data, and the EEG pre-onset window (-100 to 0 ms) was analyzed only after the binary comparison showed no significant pre-onset effect. These choices can inflate p-values and r-squared values by double dipping, and they are a genuine statistical validity concern, but they do not make the central correlation equivalent to its inputs by construction: the monotonic relation across ten predictability bins is not forced by the window and channel selection. The self-citations (e.g., [Koe+24; Kra+24b] for protocol and [KSK23] for artifact removal) are methodological context, not load-bearing justifications of the central result. Therefore, under the definitional circularity standard, the paper is not circular.
Assumptions & free parameters
free parameters (7)
- low/high predictability threshold =
0.5
- EEG ROI channels =
CP2, CPz, CP1, P2, Pz, P1
- MEG ROI sensors =
A229, A212, A178, A154, A126, A230, A213, A179, A155, A127, A177, A153, A125 (left frontal)
- post-stimulus analysis windows =
EEG 300-450 ms; MEG 500-650 ms
- pre-stimulus analysis windows =
EEG -100 to 0 ms; MEG -350 to -250 ms
- number of correlation bins =
10
- baseline window =
-1.0 to 0.0 s relative to word onset
assumptions (4)
- domain assumption BERT masked-token probability is a valid operationalization of human word predictability for the audiobook text.
- domain assumption Forced-alignment word onset timestamps from WebMAUS are accurate to within the temporal precision needed for ERP/ERF epoching.
- domain assumption The 1-20 Hz bandpass, baseline correction from -1.0 to 0.0 s, and ICA artifact removal preserve the anticipatory pre-onset signals under study.
- ad hoc to paper The selected parietal EEG channels, left frontal MEG sensors, and specific time windows are representative without selection bias.
Cite this review
Pith. "Pith review of The Predictive Brain: Neural Correlates of Word Expectancy Align with Large Language Model Prediction Probabilities." pith.science (2026). https://pith.science/paper/3KK3FS5Q
@misc{pith2026250608511,
author = {Pith},
title = {Pith review of: The Predictive Brain: Neural Correlates of Word Expectancy Align with Large Language Model Prediction Probabilities},
year = {2026},
howpublished = {\url{https://pith.science/paper/3KK3FS5Q}},
note = {Machine review of arXiv:2506.08511}
}
read the original abstract
Predictive coding theory suggests that the brain continuously anticipates upcoming words to optimize language processing, but the neural mechanisms remain unclear, particularly in naturalistic speech. Here, we simultaneously recorded EEG and MEG data from 29 participants while they listened to an audio book and assigned predictability scores to nouns using the BERT language model. Our results show that higher predictability is associated with reduced neural responses during word recognition, as reflected in lower N400 amplitudes, and with increased anticipatory activity before word onset. EEG data revealed increased pre-activation in left fronto-temporal regions, while MEG showed a tendency for greater sensorimotor engagement in response to low-predictability words, suggesting a possible motor-related component to linguistic anticipation. These findings provide new evidence that the brain dynamically integrates top-down predictions with bottom-up sensory input to facilitate language comprehension. To our knowledge, this is the first study to demonstrate these effects using naturalistic speech stimuli, bridging computational language models with neurophysiological data. Our findings provide novel insights for cognitive computational neuroscience, advancing the understanding of predictive processing in language and inspiring the development of neuroscience-inspired AI. Future research should explore the role of prediction and sensory precision in shaping neural responses and further refine models of language processing.
Figures
Reference graph
Works this paper leans on
-
[1]
[AIM24] AI@Meta. “Llama 3.2 Model Card”. In: (2024).url:https://github.com/ meta - llama / llama - models / blob / main / models / llama3 _ 2 / MODEL _ CARD.md. [AlK+24] Badr AlKhamissi et al. “The LLM Language Network: A Neuroscientific Ap- proachforIdentifyingCausallyTask-RelevantUnits”.In:arXivpreprintarXiv:2411.02280 (2024). [Ber17] Anna M Beres. “Tim...
arXiv 2024
-
[2]
Neurocomputational underpinnings of expected surprise
[Lec+22] Françoise Lecaignard et al. “Neurocomputational underpinnings of expected surprise”. In:Journal of Neuroscience42.3 (2022), pp. 474–486. [Mae+16] Burkhard Maess et al. “Prediction signatures in the brain: semantic pre- activation during language comprehension”. In:Frontiers in Human Neuro- science10 (2016), p
work page 2022
-
[5]
Minneapolis, Minnesota. 2019, p
work page 2019
-
[10]
[Pan+05] Dimitrios Pantazis et al. “A comparison of random field theory and permuta- tion methods for the statistical analysis of MEG data”. In:Neuroimage25.2 (2005), pp. 383–394. [Pas+02] Roberto Domingo Pascual-Marqui et al. “Standardized low-resolution brain electromagnetic tomography (sLORETA): technical details”. In:Methods Find Exp Clin Pharmacol24....
work page 2005
-
[13]
Neural network based successor representations to form cognitive maps of space and language
[Sto+22] Paul Stoewer et al. “Neural network based successor representations to form cognitive maps of space and language”. In:Scientific Reports12.1 (2022), p. 11233. [Sto+23] Paul Stoewer et al. “Neural network based formation of cognitive maps of se- mantic spaces and the putative emergence of abstract concepts”. In:Scientific Reports13.1 (2023), p
work page 2022
-
[32]
Analysis of Argument Structure Constructions in the Large Language Model BERT
Curran Associates, Inc., 2019, pp. 8024–8035.url:http://papers. neurips . cc / paper / 9015 - pytorch - an - imperative - style - high - performance-deep-learning-library.pdf. [Ped+11] F. Pedregosa et al. “Scikit-learn: Machine Learning in Python”. In:Journal of Machine Learning Research12 (2011), pp. 2825–2830. [PG20] Friedemann Pulvermüller and Luigi Gr...
work page Pith review arXiv 2011
-
[135]
Contextual feature extraction hierarchies converge in large language models and the brain
[Mis+24] Gavin Mischler et al. “Contextual feature extraction hierarchies converge in large language models and the brain”. In:Nature Machine Intelligence(2024), pp. 1–11. [Ope22] TB OpenAI.Chatgpt: Optimizing language models for dialogue. OpenAI
work page 2024
-
[267]
[GTP21] Luigi Grisoni, Rosario Tomasello, and Friedemann Pulvermüller. “Correlated brain indexes of semantic prediction and prediction error: Brain localization and category specificity”. In:Cerebral Cortex31.3 (2021), pp. 1553–1568. [Har+20] CharlesR.Harrisetal.“ArrayprogrammingwithNumPy”.In:Nature585.7825 (Sept. 2020), pp. 357–362.doi:10 . 1038 / s41586...
Show all 14 references
-
[423]
spaCy: Industrial-strength Natural Language Pro- cessing in Python
[Hon+20] Matthew Honnibal et al. “spaCy: Industrial-strength Natural Language Pro- cessing in Python”. In: (2020).doi:10.5281/zenodo.1212303. [KBW20] Gina R Kuperberg, Trevor Brothers, and Edward W Wlotko. “A tale of two positivities and the N400: Distinct neural signatures ar...
2020 arXiv
-
[591]
The more human-like the language model, the more surprisal is the best predictor of N400 amplitude
[MB22] James Michaelov and Ben Bergen. “The more human-like the language model, the more surprisal is the best predictor of N400 amplitude”. In:NeurIPS 2022 Workshop on Information-Theoretic Principles in Cognitive Systems
2022
-
[2015]
Speech motor cortex enables BCI cursor control and click
[Sin+24] Tyler Singer-Clark et al. “Speech motor cortex enables BCI cursor control and click”. In:bioRxiv(2024), pp. 2024–11. 18 [SK24] Achim Schilling and Patrick Krauss.The Bayesian brain: world models and conscious dimensions of auditory phantom perception
2024
-
[2022]
Strong Prediction: Language model surprisal ex- plainsmultipleN400effects
[Mic+24] James A Michaelov et al. “Strong Prediction: Language model surprisal ex- plainsmultipleN400effects”.In:Neurobiologyoflanguage5.1(2024),pp.107–
2024
-
[2024]
Adaptive ica for speech eegartifactremoval
[KSK23] Nikola Koelbl, Achim Schilling, and Patrick Krauss. “Adaptive ica for speech eegartifactremoval”.In:20235thInternationalConferenceonBio-engineering for Smart Technologies (BioSMART). IEEE. 2023, pp. 1–4. [KSK24] Hassane Kissane, Achim Schilling, and Patrick Krauss. “An...
2024 arXiv
-
[3644]
Coincidence detection and integration behavior in spiking neural networks
[Sto+24] Andreas Stoll et al. “Coincidence detection and integration behavior in spiking neural networks”. In:Cognitive Neurodynamics18.4 (2024), pp. 1753–1765. [Sur+23] Kishore Surendra et al. “Word class representations spontaneously emerge in a deep neural network trained o...
2024 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.