Pith. sign in

REVIEW 6 major objections 6 minor 14 references

The Predictive Brain: Neural Correlates of Word Expectancy Align with Large Language Model Prediction Probabilities

T0 review · 6 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper claims that during natural audiobook listening, the brain's word-expectancy signals—reflected in N400 amplitude and pre-onset activity—track BERT's continuous predictability scores rather than only a binary high/low distinction.

desk verdict A useful but statistically fragile replication of known N400-surprisal effects; the 'first naturalistic' claim is wrong, and the headline r² values are inflated by data-driven selection. read the letter →

arxiv 2506.08511 v1 pith:3KK3FS5Q submitted 2025-06-10 q-bio.NC

classification q-bio.NC
keywords predictivecodingN400BERTEEGMEGsurprisalnaturalisticspeechsemanticpredictionpotential
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that during naturalistic listening—an audiobook, not controlled sentences—the brain's word-expectancy machinery aligns with the probabilities assigned by a large language model. By assigning every noun a BERT masked-word probability and comparing that score against simultaneously recorded EEG and MEG, the authors find a graded relationship: the higher the model's predictability score, the smaller the N400-like neural response during word recognition and the larger the anticipatory activity just before the word begins. If true, this would mean predictive coding in language operates continuously in real speech, and that transformer-language-model probabilities are a workable quantitative proxy for the brain's top-down predictions.

What carries the argument

The load-bearing object is a per-noun predictability score produced by masking each noun in the audiobook text and reading BERT's softmax probability for the original word (averaged over sub-word tokens), which is then compared with word-locked neural responses. The neural side is the N400 component—the well-known negative deflection around 300–500 ms after a word that indexes semantic processing effort—plus pre-onset anticipatory activity. Forced alignment supplies precise word-onset times, and the analysis correlates mean activity in ten equal-sized predictability bins with the log of predictability scores, so the test is whether brain responses scale gradually with probability rather than only differing between high and low bins.

What would settle it

A reader could re-run the ten-bin correlation on the same audiobook data with the parietal EEG electrodes and left-frontal MEG sensors fixed in advance, adding word frequency, word length, and acoustic prominence as covariates; if the EEG correlation at 300–450 ms and the MEG correlation at 500–650 ms no longer reach significance, the claim that BERT predictability specifically drives the effect would be refuted.

Watch

Extended reading notes

Core claim

The central discovery is a continuous inverse correlation between BERT predictability and N400 amplitude in naturalistic speech, together with a positive correlation between predictability and pre-word-onset neural activity. In EEG, the post-onset effect appears at 300–450 ms over parietal electrodes ($p = 0.0006$, $r^2 = 0.79$ over ten predictability bins), and in MEG at 500–650 ms over left frontal sensors ($p = 0.0013$, $r^2 = 0.75$). Pre-onset activity also scales with predictability (EEG at $-100$ to $0$ ms: $p = 0.028$, $r^2 = 0.47$; MEG at $-350$ to $-250$ ms: $p = 0.0315$, $r^2 = 0.46$), and this pre-activation is negatively related to the later N400 amplitude, linking anticipation to reduced processing cost. Source reconstruction places the post-onset difference in parietal and sensorimotor cortex and the EEG pre-onset effect in left fronto-temporal regions, consistent with a left-lateralized language network. The authors further claim this is the first demonstration of these effects with naturalistic speech stimuli.

Load-bearing premise

The correlations rest on the assumption that the selected electrodes, MEG sensors, and time windows were not chosen from the same data used to test them, and that BERT's probability captures human predictability without needing to separate word frequency, word length, or acoustic prominence.

Editorial extensions

If this is right

  • If the claim holds, N400 amplitude in natural listening can be read as a graded neural surprisal signal, and language-model probabilities become a quantitative predictor of neural processing effort.
  • The effect generalizing from controlled visual sentence reading to continuous audiobook speech means predictive-coding accounts of language are not limited to artificial stimulus protocols.
  • The negative relation between pre-onset activity and N400 amplitude supports a direct anticipatory mechanism: stronger top-down pre-activation reduces bottom-up integration cost.
  • Convergent EEG and MEG correlations imply the predictability effect is robust across measurement modalities and is not a single-sensor artifact.
  • Because the correlation is computed over ten bins rather than a high/low split, the brain's response appears to track probability continuously, not just categorical expectancy violations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next test, not run in the paper, is a mixed-effects model on single trials that includes word frequency, word length, and acoustic prominence as covariates; if the BERT-probability slope survives, the case for context-specific prediction is much stronger.
  • The MEG tendency toward sensorimotor pre-onset engagement for low-predictability nouns suggests a testable motoric component of anticipation: compare nouns that differ in articulatory complexity while matching predictability.
  • Because BERT's masked probabilities are bidirectional, they include future context; comparing with a causal left-to-right language model would clarify whether the brain's anticipatory signal respects temporal order.
  • The pre-onset EEG window ($-100$ to $0$ ms) was inspected after the binary comparison showed no effect there, so a pre-registered replication with fixed windows is the cleanest way to confirm the anticipatory result.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

6 major / 6 minor

Summary. The paper reports an EEG/MEG study in which 29 participants listened to a German audiobook, and BERT-based predictability scores for nouns were correlated with neural responses. The central claims are that higher predictability is associated with reduced N400-like amplitudes (EEG: p=0.0006, r2=0.79 at 300–450 ms; MEG: p=0.0013, r2=0.75 at 500–650 ms) and with increased pre-onset activity (EEG: p=0.028, r2=0.47 at -100–0 ms; MEG: p=0.0315, r2=0.46 at -350 to -250 ms). These results are interpreted as evidence for continuous alignment between LLM word-expectancy probabilities and predictive brain processing in naturalistic speech, and the paper claims to be the first demonstration of such effects with naturalistic stimuli.

Significance. If robust, the central finding would extend previous N400-surprisal work to continuous, naturalistic listening with simultaneous EEG and MEG, and would strengthen the case that transformer-based language model probabilities provide a useful quantitative account of neural word-expectancy signals. The study has genuine strengths: a naturalistic stimulus, a plausible use of BERT-derived predictability, simultaneous EEG/MEG acquisition, and explicit attempts to check reproducibility with additional low-predictability subsets. However, the headline quantitative claims rest on analytic choices that were made after inspecting the same dataset that is then used for the reported p-values and r2 values. Because sensor selection, time-window selection, and the EEG pre-onset window were data-driven, the reported effect sizes are likely inflated, and the central claim that the brain's expectancy signals are 'continuously aligned' with BERT probabilities is not yet supported at the stated level of confidence. The qualitative direction of the N400 effect is consistent with prior literature, but the paper's specific quantitative evidence requires substantially stronger statistical validation.

major comments (6)
  1. [§2.3 and §3.3] The sensor and time-window selections used for the headline correlations are data-driven. In §2.3, the authors state that they 'identified peak responses in the time window from -1.0s until 2.0s and extracted the topographic distribution at the time point of the strongest negative responses,' and then selected parietal EEG electrodes and left frontal MEG sensors. In §3.3, the correlation analysis uses mean activity over these selected sensors and over time windows that were identified as significant in the same dataset. The reported p=0.0006/r2=0.79 and p=0.0013/r2=0.75 are therefore not out-of-sample estimates; selecting the spatial and temporal region that looks most favorable can inflate correlations substantially. The authors should either pre-specify sensors and time windows, or report results across the full sensor array and a broad fixed window, or use cross-validation/permutation procedures that account for the selection step.
  2. [§3.3] The EEG pre-onset window (-100 to 0 ms) is explicitly post hoc. The text states: 'in the EEG data, where no significant pre-onset effects were observed, we examined a pre-word interval at -100ms to assess early anticipatory processes.' This window was therefore tested after an initial null result, yet the reported p=0.028 is presented without correction for this additional test. The pre-onset EEG finding cannot be treated as confirmatory evidence; it should be labeled exploratory and the p-value should be adjusted, or an independent dataset/pre-registered analysis should be provided.
  3. [§3.3] The correlation analysis does not include any confounding regressors. In natural speech, BERT predictability is correlated with word frequency, word length, acoustic duration, speech rate, and possibly lexical neighborhood properties. The regression in §3.3 models only mean predictability per bin against mean neural activity, so the observed correlations could be driven by these confounds rather than by predictability per se. The authors should add control analyses that include word frequency, word length, and acoustic duration as covariates or as binning variables, or otherwise demonstrate that the predictability effect survives such controls.
  4. [§2.2 and §3.1] The handling of the class imbalance is not statistically transparent. The authors select a random subset of 194 low-predictability trials to match the 194 high-predictability trials, and then report that two additional low-predictability subsets show 'consistent' effects. However, no seed is given for the random selection, no distribution across many possible subsets is reported, and one of the additional EEG subsets shows only a trend (p=0.11), which is nonetheless interpreted as supporting reproducibility. The authors should report the full set of subset analyses, state whether they were preplanned, and quantify variability across random draws rather than selecting favorable subsets.
  5. [§2.3] The baseline correction used for the epoch analysis may materially affect the pre-onset findings. The epochs are baseline-corrected from -1.0 to 0.0 s, which includes the exact pre-onset intervals later analyzed (-350 to -250 ms and -100 to 0 ms). This means that the reported pre-onset 'activity' is measured relative to a baseline that already contains the anticipatory effect, potentially removing or distorting the very signal of interest. The authors should report whether the pre-onset correlations survive when using a baseline outside the analyzed window (e.g., -1000 to -600 ms) or when no baseline correction is applied.
  6. [§3.2] The source-space results are reported with strong language ('significant activity was observed') but no statistical test is described for the source reconstructions. It is unclear whether the source-space differences between low- and high-predictability nouns were tested with cluster-based permutation statistics, corrected for multiple comparisons across sources, or are purely descriptive. The authors need to specify the statistical procedure used in source space, or explicitly label these results as descriptive and not claim significance.
minor comments (6)
  1. [§3.3 and Figure 4] There is an inconsistency in the reported EEG post-stimulus p-value: the text states p=0.0006, while the Figure 4 caption states p=0.00006. The authors should correct this discrepancy and ensure all p-values/r2 values match between text and figures.
  2. [Introduction] Typo: 'many studies investigation predictability' should be 'many studies investigating predictability.'
  3. [§5 Acknowledgments] Typo: 'programmme' should be 'programme' (or 'program').
  4. [§2.2] The description of the semi-logarithmic scaling as 'reverses the softmax operation' is imprecise; taking the logarithm of a softmax probability does not invert the softmax mapping. The sentence should be rephrased as applying a log transform to approximate surprisal.
  5. [Discussion / Abstract] The claim that this is 'the first study to demonstrate these effects using naturalistic speech stimuli' is difficult to reconcile with the authors' own citation of Goldstein et al., who used naturalistic story listening and reported similar predictability-related neural responses around 400 ms. The novelty claim should be narrowed or supported by an explicit comparison of stimuli and analyses.
  6. [General] No data or code availability statement is provided, and the analysis pipeline is described only briefly. Given the data-driven nature of several analytic choices, sharing the analysis code and processed trial-level data would be important for assessing reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No definitional circularity: the BERT-neural correlations are empirical fits, not derivations from their inputs.

full rationale

I walked the claimed derivation chain and found no step in which a predicted quantity is defined in terms of the data used to test it in the sense of an equation or fitted parameter being renamed as a prediction. The BERT predictability scores are computed from the audio book text independently of the EEG/MEG recordings, and the neural responses are measured directly; the reported correlations are ordinary regression fits to binned mean amplitudes. The methodology does contain post hoc selections: parietal electrodes and left frontal sensors were chosen from the topographic peak of the same data, and the EEG pre-onset window (-100 to 0 ms) was analyzed only after the binary comparison showed no significant pre-onset effect. These choices can inflate p-values and r-squared values by double dipping, and they are a genuine statistical validity concern, but they do not make the central correlation equivalent to its inputs by construction: the monotonic relation across ten predictability bins is not forced by the window and channel selection. The self-citations (e.g., [Koe+24; Kra+24b] for protocol and [KSK23] for artifact removal) are methodological context, not load-bearing justifications of the central result. Therefore, under the definitional circularity standard, the paper is not circular.

Assumptions & free parameters 7 free parameters · 4 assumptions · 0 invented entities

The paper's quantitative claims depend on a chain of analytic choices (threshold, ROI, time windows, binning, baseline) that were set using the same data that is then tested, plus the domain assumption that BERT probabilities reflect human prediction. No independent validation of the predictability measure or control for lexical and acoustic confounds is provided.

free parameters (7)
  • low/high predictability threshold = 0.5
    Arbitrary split yields 1,182 low vs 194 high nouns; a different threshold would change the binary comparison and the balance correction.
  • EEG ROI channels = CP2, CPz, CP1, P2, Pz, P1
    Selected after inspecting the topographic distribution at the peak negative response in the same data.
  • MEG ROI sensors = A229, A212, A178, A154, A126, A230, A213, A179, A155, A127, A177, A153, A125 (left frontal)
    Selected after inspecting the topographic distribution in the same data.
  • post-stimulus analysis windows = EEG 300-450 ms; MEG 500-650 ms
    Identified from cluster-based permutation on the same dataset, then used for the correlation analysis.
  • pre-stimulus analysis windows = EEG -100 to 0 ms; MEG -350 to -250 ms
    The EEG window was only examined after no significant pre-onset effect was found in the binary test; the MEG window came from the same data contrast.
  • number of correlation bins = 10
    Chosen to keep about 137 trials per bin; the 10-point regression drives the reported r2 and p values.
  • baseline window = -1.0 to 0.0 s relative to word onset
    Chosen for baseline correction; it overlaps the pre-onset analysis window, potentially attenuating the anticipatory effect.
assumptions (4)
  • domain assumption BERT masked-token probability is a valid operationalization of human word predictability for the audiobook text.
    The authors do not compare BERT scores with human cloze judgments or behavioral ratings for the stimulus corpus; the entire study rests on this equivalence.
  • domain assumption Forced-alignment word onset timestamps from WebMAUS are accurate to within the temporal precision needed for ERP/ERF epoching.
    Epochs from -1.0 to +2.0 s assume alignments are correct; misalignment would blur the N400 and pre-onset windows.
  • domain assumption The 1-20 Hz bandpass, baseline correction from -1.0 to 0.0 s, and ICA artifact removal preserve the anticipatory pre-onset signals under study.
    The baseline window overlaps the pre-onset window of interest, so slow anticipatory components may be partially removed; the paper does not validate this.
  • ad hoc to paper The selected parietal EEG channels, left frontal MEG sensors, and specific time windows are representative without selection bias.
    These choices were made after inspecting the same dataset's topographic and temporal responses, making the confirmatory statistics partially dependent on the selection.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Predictive Brain: Neural Correlates of Word Expectancy Align with Large Language Model Prediction Probabilities." pith.science (2026). https://pith.science/paper/3KK3FS5Q

@misc{pith2026250608511,
  author       = {Pith},
  title        = {Pith review of: The Predictive Brain: Neural Correlates of Word Expectancy Align with Large Language Model Prediction Probabilities},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3KK3FS5Q}},
  note         = {Machine review of arXiv:2506.08511}
}
read the original abstract

Predictive coding theory suggests that the brain continuously anticipates upcoming words to optimize language processing, but the neural mechanisms remain unclear, particularly in naturalistic speech. Here, we simultaneously recorded EEG and MEG data from 29 participants while they listened to an audio book and assigned predictability scores to nouns using the BERT language model. Our results show that higher predictability is associated with reduced neural responses during word recognition, as reflected in lower N400 amplitudes, and with increased anticipatory activity before word onset. EEG data revealed increased pre-activation in left fronto-temporal regions, while MEG showed a tendency for greater sensorimotor engagement in response to low-predictability words, suggesting a possible motor-related component to linguistic anticipation. These findings provide new evidence that the brain dynamically integrates top-down predictions with bottom-up sensory input to facilitate language comprehension. To our knowledge, this is the first study to demonstrate these effects using naturalistic speech stimuli, bridging computational language models with neurophysiological data. Our findings provide novel insights for cognitive computational neuroscience, advancing the understanding of predictive processing in language and inspiring the development of neuroscience-inspired AI. Future research should explore the role of prediction and sensory precision in shaping neural responses and further refine models of language processing.

Figures

Figures reproduced from arXiv: 2506.08511 by the authors.

Figure 1
Figure 1. (A): Scheme for computing predictability scores for individual nouns in the audio [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. A: Grand average nouns ERP (EEG) including topographic map at peak 350 ms [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. RMS amplitudes in source space for EEG data in the time interval 300 ms-450 ms [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Correlation analysis between predictability scores of BERT and neural activity [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

14 extracted references · 11 canonical work pages

  1. [1]

    Llama 3.2 Model Card

    [AIM24] AI@Meta. “Llama 3.2 Model Card”. In: (2024).url:https://github.com/ meta - llama / llama - models / blob / main / models / llama3 _ 2 / MODEL _ CARD.md. [AlK+24] Badr AlKhamissi et al. “The LLM Language Network: A Neuroscientific Ap- proachforIdentifyingCausallyTask-RelevantUnits”.In:arXivpreprintarXiv:2411.02280 (2024). [Ber17] Anna M Beres. “Tim...

  2. [2]

    Neurocomputational underpinnings of expected surprise

    [Lec+22] Françoise Lecaignard et al. “Neurocomputational underpinnings of expected surprise”. In:Journal of Neuroscience42.3 (2022), pp. 474–486. [Mae+16] Burkhard Maess et al. “Prediction signatures in the brain: semantic pre- activation during language comprehension”. In:Frontiers in Human Neuro- science10 (2016), p

  3. [5]

    Minneapolis, Minnesota. 2019, p

  4. [10]

    A comparison of random field theory and permuta- tion methods for the statistical analysis of MEG data

    [Pan+05] Dimitrios Pantazis et al. “A comparison of random field theory and permuta- tion methods for the statistical analysis of MEG data”. In:Neuroimage25.2 (2005), pp. 383–394. [Pas+02] Roberto Domingo Pascual-Marqui et al. “Standardized low-resolution brain electromagnetic tomography (sLORETA): technical details”. In:Methods Find Exp Clin Pharmacol24....

  5. [13]

    Neural network based successor representations to form cognitive maps of space and language

    [Sto+22] Paul Stoewer et al. “Neural network based successor representations to form cognitive maps of space and language”. In:Scientific Reports12.1 (2022), p. 11233. [Sto+23] Paul Stoewer et al. “Neural network based formation of cognitive maps of se- mantic spaces and the putative emergence of abstract concepts”. In:Scientific Reports13.1 (2023), p

  6. [32]

    Analysis of Argument Structure Constructions in the Large Language Model BERT

    Curran Associates, Inc., 2019, pp. 8024–8035.url:http://papers. neurips . cc / paper / 9015 - pytorch - an - imperative - style - high - performance-deep-learning-library.pdf. [Ped+11] F. Pedregosa et al. “Scikit-learn: Machine Learning in Python”. In:Journal of Machine Learning Research12 (2011), pp. 2825–2830. [PG20] Friedemann Pulvermüller and Luigi Gr...

  7. [135]

    Contextual feature extraction hierarchies converge in large language models and the brain

    [Mis+24] Gavin Mischler et al. “Contextual feature extraction hierarchies converge in large language models and the brain”. In:Nature Machine Intelligence(2024), pp. 1–11. [Ope22] TB OpenAI.Chatgpt: Optimizing language models for dialogue. OpenAI

  8. [267]

    Correlated brain indexes of semantic prediction and prediction error: Brain localization and category specificity

    [GTP21] Luigi Grisoni, Rosario Tomasello, and Friedemann Pulvermüller. “Correlated brain indexes of semantic prediction and prediction error: Brain localization and category specificity”. In:Cerebral Cortex31.3 (2021), pp. 1553–1568. [Har+20] CharlesR.Harrisetal.“ArrayprogrammingwithNumPy”.In:Nature585.7825 (Sept. 2020), pp. 357–362.doi:10 . 1038 / s41586...

Show all 14 references
  1. [423]

    spaCy: Industrial-strength Natural Language Pro- cessing in Python

    [Hon+20] Matthew Honnibal et al. “spaCy: Industrial-strength Natural Language Pro- cessing in Python”. In: (2020).doi:10.5281/zenodo.1212303. [KBW20] Gina R Kuperberg, Trevor Brothers, and Edward W Wlotko. “A tale of two positivities and the N400: Distinct neural signatures ar...

  2. [591]

    The more human-like the language model, the more surprisal is the best predictor of N400 amplitude

    [MB22] James Michaelov and Ben Bergen. “The more human-like the language model, the more surprisal is the best predictor of N400 amplitude”. In:NeurIPS 2022 Workshop on Information-Theoretic Principles in Cognitive Systems

  3. [2015]

    Speech motor cortex enables BCI cursor control and click

    [Sin+24] Tyler Singer-Clark et al. “Speech motor cortex enables BCI cursor control and click”. In:bioRxiv(2024), pp. 2024–11. 18 [SK24] Achim Schilling and Patrick Krauss.The Bayesian brain: world models and conscious dimensions of auditory phantom perception

  4. [2022]

    Strong Prediction: Language model surprisal ex- plainsmultipleN400effects

    [Mic+24] James A Michaelov et al. “Strong Prediction: Language model surprisal ex- plainsmultipleN400effects”.In:Neurobiologyoflanguage5.1(2024),pp.107–

  5. [2024]

    Adaptive ica for speech eegartifactremoval

    [KSK23] Nikola Koelbl, Achim Schilling, and Patrick Krauss. “Adaptive ica for speech eegartifactremoval”.In:20235thInternationalConferenceonBio-engineering for Smart Technologies (BioSMART). IEEE. 2023, pp. 1–4. [KSK24] Hassane Kissane, Achim Schilling, and Patrick Krauss. “An...

  6. [3644]

    Coincidence detection and integration behavior in spiking neural networks

    [Sto+24] Andreas Stoll et al. “Coincidence detection and integration behavior in spiking neural networks”. In:Cognitive Neurodynamics18.4 (2024), pp. 1753–1765. [Sur+23] Kishore Surendra et al. “Word class representations spontaneously emerge in a deep neural network trained o...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.