REVIEW 3 major objections 4 minor 1 cited by
surprisal is Not a Theory
T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read The paper argues that surprisal scores drawn from different large language models encode divergent internal computations, so treating them as interchangeable evidence for Surprisal Theory is unsound.
desk verdict Useful critique of interchangeable surprisal, but the layerwise evidence doesn't close the gap to behavioral non-interchangeability. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key mechanism is the 'logit lens' and the companion 'lexical lens': probes that read off a model's next-word probability distribution and its top token guesses at each hidden layer, making internal computation visible. The paper uses these to compare GPT-2, Pythia, and RoBERTa, which are similar in parameter count but differ in training objective (decoder vs. masked encoder), tied vs. untied embeddings, and context directionality. The interpretive frame is multiple realizability combined with functional equivalence: if a computational-level theory is allowed to abstract over algorithms, then internal differences matter only when they change downstream processing. The paper argues they do
What would settle it
A decisive test would be to fit the same reading-time regression twice, once with final-layer surprisals from two models whose internal trajectories diverge (e.g., GPT-2 and Pythia) and once with the same final values but shuffled layerwise trajectories; if the coefficients and model fit are statistically indistinguishable, the paper's claim that different internal computations constitute different linking hypotheses would be undermined. Alternatively, showing that the divergence in layerwise trajectories never translates into differences in predicted reading times for any naturalistic corpus
Extended reading notes
Core claim
The central claim is that the 'representation-agnosticism' often invoked to justify using LLM surprisal as a computational-level measure is untenable. Different LLMs compute next-word probabilities through different latent structures and algorithms; the probabilities themselves are therefore not implementations of one same linking hypothesis. The authors introduce the criterion of functional equivalence: two models' surprisal values are interchangeable only if all subsequent processing is identical. Using logit-lens and lexical-lens probes, they show that although models like GPT-2 and Pythia correlate at 0.92 in final-layer probabilities, the layerwise evolution of those probabilities and t
Load-bearing premise
The argument rests on the premise that two models' surprisal values are interchangeable only when every step of their internal computation is functionally equivalent; if a computational-level theory is allowed to abstract away internal algorithms, then the observed layerwise differences are irrelevant to the interchangeability of the metric.
Editorial extensions
If this is right
- Researchers should stop treating surprisal estimates from different LLMs as interchangeable in reading-time and neurobehavioral analyses.
- Model selection in psycholinguistics must be justified by explicit cognitive commitments, not by predictive power alone.
- Surprisal Theory cannot claim to be a computational-level account while leaving representations and algorithms unexamined.
- Published comparisons across dozens of models may be conflating distinct linking hypotheses rather than measuring the same construct.
- The field should move toward explainable, cognitively motivated models, and reviewers should not routinely demand adding LLM surprisal as a covariate.
Reading between the lines
- If the authors' premise is right, meta-analyses pooling surprisal across model families may be averaging over qualitatively different computations, which could explain some conflicting results in the surprisal-reading-time literature.
- The same critique likely applies to other LLM-extracted psycholinguistic indices (e.g., attention weights, entropy), not just surprisal.
- A concrete extension: researchers could vary architecture while matching final surprisal values, to test whether layerwise trajectory differences actually change reading-time predictions—if they don't, the practical significance of this argument would shrink.
- The argument could be transformed into a positive methodological program: use layerwise probes to determine which representational commitments are necessary to replicate human processing, rather than treating any LLM as neutral.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that the common psycholinguistic practice of treating surprisal values extracted from different large language models as interchangeable is theoretically unjustified. It reports two analyses using GPT-2, Pythia-160M, and RoBERTa: first, layerwise logit-lens probabilities are correlated with human cloze probabilities (Experiment 1); second, the 'lexicality' of each model's top-k predictions is traced across layers (Experiment 2). The authors find that layerwise trajectories differ markedly across models and conclude that different architectures and training objectives instantiate different algorithmic and representational commitments, undermining representation-agnosticism in Surprisal Theory. The paper then develops a conceptual critique of using LLM-derived surprisal as a computational-level explanation and offers recommendations for theory-driven model selection.
Significance. The paper addresses a timely and important methodological question: whether results obtained with one LLM's surprisal values can be taken to test Surprisal Theory independently of the model that generated them. Its strengths include public data and scripts (OSF), an emphasis on falsifiable practice rather than fitted parameters, and a clear engagement with Marr's levels. If the central claim were established, it would make a meaningful contribution to computational psycholinguistics. However, the significance is conditional: the empirical results demonstrate differences in internal algorithmic trajectories, but the operationalized surprisal in the literature is the final-layer probability, and the paper itself reports that final-layer GPT-2 and Pythia probabilities correlate at 0.92. The gap between algorithmic divergence and non-interchangeable linking hypotheses is the load-bearing issue that must be addressed.
major comments (3)
- [§2.3 and §4.1] The central inference from divergent layerwise computations to non-interchangeable surprisal-based theories is not supported by the presented evidence. Study 1.1 reports that final-layer GPT-2 and Pythia probabilities correlate at 0.92, and final-layer correlations to cloze are similar across models. Surprisal as commonly operationalized is the final-layer negative log probability, and Marr's computational level explicitly abstracts over algorithmic implementation—a point the paper itself acknowledges in §4.1 ('multiple-realizability is intrinsic to computational-level theories'). The statement at the end of §4.1, 'If different LLMs change the linking hypothesis between probabilities and effort...', is exactly the untested premise. To support the paper's headline claim, the authors need to show that final surprisal values differ across models in ways that change behavioral predictions (e
- [§2.4–2.5 and §3.2] The claims of 'starkly different' trajectories (Figure 3) and 'massive fluctuations' in lexicality (Figure 4) rest on visual inspection of only three models, with no error bars, confidence intervals, or inferential statistics on the layerwise curves. This is particularly consequential because the qualitative reading of the curves is a key part of the argument that the models are not interchangeable. Additionally, the lexicality metric depends on the choice of k=10 in §3.2, but no sensitivity analysis is reported; different tokenizers across models could confound the comparison. I recommend quantifying trajectory differences (e.g., bootstrap CIs, pairwise curve-distance measures) and testing whether the lexicality patterns are robust to k and to tokenization.
- [§4.4] The discussion of the Natural Stories misalignment error is presented as evidence that the field's reliance on surprisal is theoretically fragile. This is an interesting argument, but the inference is not logically entailed: a correlation can survive a one-position shift for many reasons (e.g., autocorrelation in surprisal, smoothness of the cost function). The claim that the error 'should have led researchers to reconsider why exactly their surprisal analyses had been so successful' is speculative. This is not the load-bearing part of the paper, but it should be framed as an open question or a testable prediction rather than a necessary consequence.
minor comments (4)
- [§1] Typo: 'effors' should be 'efforts'. Also, the phrase 'these assumptions are is in fact not wholly valid' has a grammar error.
- [Title] The title 'surprisal is Not a Theory' overstates the paper's scope; the argument is specifically about LLM-derived surprisal and its use as a computational-level proxy. A more precise title would help manage reader expectations.
- [Figure 3] The figure would be more informative with confidence bands or per-context variability; as published, the eye is drawn to mean trajectories that may mask high variance.
- [§3.2 / Table 2] The tokenization differences across models make the qualitative examples in Table 2 hard to interpret; a brief note on subword segmentation would improve clarity.
Circularity Check
No significant circularity: the paper's claims rest on independent empirical probes of public models and corpora, not on fitted parameters or load-bearing self-citations.
full rationale
The paper contains no first-principles derivation, no fitted parameter that is later renamed a prediction, and no uniqueness theorem or ansatz imported from the authors' prior work. Its central evidence is empirical: layerwise logit-lens comparisons against Peelle et al. (2020) cloze norms (Experiment 1) and a new 'lexicality' metric applied to public models and the Provo Corpus (Experiment 2). These comparisons are externally checkable and do not reduce by construction to the conclusion. The authors' own result that final-layer GPT-2 and Pythia probabilities correlate at 0.92 and that final-layer cloze correlations are similar across models actually undercuts the strongest version of their 'not interchangeable' claim, which is an interpretive overreach rather than a circular step. The self-citations (Jacobs & McCarthy 2020, Jacobs et al. 2025, De Santo 2025, Jacobs & Grobol 2025) are contextual or auxiliary, not load-bearing: the median-split replication from Jacobs & McCarthy (2020) is re-demonstrated with the paper's own data. The functional-equivalence criterion is imported from Rogers (2021), an external source, and the contested inference—that different internal layerwise trajectories imply different computational-level linking hypotheses—is an unsupported empirical/philosophical argument, not a definitional equivalence. The conditional statement in §4.1 ('If different LLMs change the linking hypothesis between probabilities and effort...') is acknowledged as a premise, not presented as a derived result. Thus no circular step can be exhibited with the required specificity, and the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (1)
- k (top-k in lexicality metric) =
10
assumptions (3)
- domain assumption Marr's computational level requires that a linking hypothesis be transparent about representational and algorithmic commitments to be explanatory.
- domain assumption Functional equivalence requires that all processing after a given point be identical, so different latent representations make surprisal non-interchangeable.
- domain assumption The logit lens and lexicality metric accurately reflect the information used by a model to compute next-word probabilities.
Cite this review
Pith. "Pith review of surprisal is Not a Theory." pith.science (2026). https://pith.science/paper/WUATTFNU
@misc{pith2026260720208,
author = {Pith},
title = {Pith review of: surprisal is Not a Theory},
year = {2026},
howpublished = {\url{https://pith.science/paper/WUATTFNU}},
note = {Machine review of arXiv:2607.20208}
}
read the original abstract
Surprisal Theory is often characterized as a computational-level explanation per (Marr, 1982). We argue in this work that, even though a computational level narrative has been used to support "representation-agnostic research" within computational psycholinguistics, the movement toward black box systems embodied by large language models (LLMs) does not exempt modelers using the surprisal metric from the representational decisions required by computational-level characterizations. In fact, we argue that the uncritical use of LLM-surprisal obfuscates the representational and algorithmic-level commitments of different models. In three analyses, we show that the choice of algorithm and model architecture play significant roles in the computation of language model probabilities. We advise that researchers who wish to test Surprisal Theory re-evaluate the practice of treating large language model probabilities as interchangeable
Figures
Forward citations
Cited by 1 Pith paper
-
Surprisal Theory is Tautological (without Rational Grounding)
Unconstrained surprisal theory is a tautology: for any non-negative difficulty measure, a language model exists whose surprisal matches it affinely.
Reference graph
Works this paper leans on
-
[1]
Behavior Research Methods , author =
Moving beyond. Behavior Research Methods , author =. 2009 , pages =. doi:10.3758/BRM.41.4.977 , language =
-
[2]
Transactions of the Association for Computational Linguistics , volume =
Oh, Byung-Doh and Schuler, William , title =. Transactions of the Association for Computational Linguistics , volume =. 2023 , month =. doi:10.1162/tacl_a_00548 , url =
-
[3]
2022 , eprint=
On the Opportunities and Risks of Foundation Models , author=. 2022 , eprint=
2022
-
[4]
Behavioral and Brain Sciences , volume=
Machine yearning: LLMs do not capture formal linguistic structure and obscure neuroscientific inquiry , author=. Behavioral and Brain Sciences , volume=. 2026 , publisher=
2026
-
[5]
Behavioral and Brain Sciences , volume=
Are language models models? , author=. Behavioral and Brain Sciences , volume=. 2026 , publisher=
2026
-
[6]
Behavioral and Brain Sciences , volume=
Language models do not yet achieve explanatory adequacy because language is more than incremental prediction , author=. Behavioral and Brain Sciences , volume=. 2026 , publisher=
2026
-
[7]
Proceedings of the 30th Conference on Computational Natural Language Learning , pages=
Linguistic Profiling of Transformer Embedding Geometry , author=. Proceedings of the 30th Conference on Computational Natural Language Learning , pages=
-
[8]
2026 , eprint=
Trajectory Dynamics in Language Model Hidden States Predict Human Processing Costs Beyond Surprisal , author=. 2026 , eprint=
2026
Show all 104 references
-
[9]
Behavior Research Methods , author =
Cloze probability, predictability ratings, and computational estimates for 205. Behavior Research Methods , author =. 2023 , pages =. doi:10.3758/s13428-023-02261-8 , abstract =
2023 doi
-
[10]
Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics , pages=
Capturing Online SRC/ORC Effort with Memory Measures from a Minimalist Parser , author=. Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics , pages=
-
[11]
Anisotropy is
Machina, Anemily and Mercer, Robert , editor =. Anisotropy is. Proceedings of the 2024. 2024 , pages =. doi:10.18653/v1/2024.naacl-long.274 , abstract =
2024 doi
-
[12]
Transactions of the Association for Computational Linguistics , author =
Testing the. Transactions of the Association for Computational Linguistics , author =. 2023 , note =. doi:10.1162/tacl_a_00612 , abstract =
2023 doi
-
[13]
Nair, Sathvik and Resnik, Philip , editor =. Words,. Findings of the. 2023 , pages =. doi:10.18653/v1/2023.findings-emnlp.752 , abstract =
2023 doi
-
[14]
Language
Wilcox, Ethan and Meister, Clara and Cotterell, Ryan and Pimentel, Tiago , editor =. Language. Proceedings of the 2023. 2023 , pages =. doi:10.18653/v1/2023.emnlp-main.466 , abstract =
2023 doi
-
[15]
Analyzing
Dar, Guy and Geva, Mor and Gupta, Ankit and Berant, Jonathan , editor =. Analyzing. Proceedings of the 61st. 2023 , pages =. doi:10.18653/v1/2023.acl-long.893 , abstract =
2023 doi
-
[16]
Wilcox, Ethan and Vani, Pranali and Levy, Roger , editor =. A. Proceedings of the 59th. 2021 , pages =. doi:10.18653/v1/2021.acl-long.76 , abstract =
2021 doi
-
[17]
Surprisal
Oh, Byung-Doh and Clark, Christian and Schuler, William , editor =. Surprisal. Proceedings of the 59th. 2021 , pages =. doi:10.18653/v1/2021.acl-long.290 , abstract =
2021 doi
-
[18]
Proceedings of the 57th
Tenney, Ian and Das, Dipanjan and Pavlick, Ellie , editor =. Proceedings of the 57th. 2019 , pages =. doi:10.18653/v1/P19-1452 , abstract =
2019 doi
-
[19]
Predictive power of word surprisal for reading times is a linear function of language model quality , url =
Goodkind, Adam and Bicknell, Klinton , editor =. Predictive power of word surprisal for reading times is a linear function of language model quality , url =. 2018 , pages =. doi:10.18653/v1/W18-0102 , booktitle =
2018 doi
-
[20]
and Otten, Leun J
Frank, Stefan L. and Otten, Leun J. and Galli, Giulia and Vigliocco, Gabriella , editor =. Word surprisal predicts. Proceedings of the 51st. 2013 , pages =
2013
-
[21]
Robustness and processing difficulty models
Rauzy, Stéphane and Blache, Philippe , editor =. Robustness and processing difficulty models. Proceedings of the. 2012 , pages =
2012
-
[22]
Hale, John , year =. A. Second
-
[23]
Computational Brain & Behavior , author =
What’s. Computational Brain & Behavior , author =. 2025 , pages =. doi:10.1007/s42113-025-00237-9 , abstract =
2025 doi
-
[24]
Behavior Research Methods , author =
Completion norms for 3085. Behavior Research Methods , author =. 2020 , pages =. doi:10.3758/s13428-020-01351-1 , abstract =
2020 doi
-
[25]
Proceedings of the National Academy of Sciences , author =
Large-scale evidence for logarithmic effects of word predictability on reading time , volume =. Proceedings of the National Academy of Sciences , author =. 2024 , pages =. doi:10.1073/pnas.2307876121 , abstract =
2024 doi
-
[26]
International Conference on Machine Learning , author =
Pythia:. International Conference on Machine Learning , author =. 2023 , pages =
2023
-
[27]
Cognitive Psychology , author =
Limits on lexical prediction during reading , volume =. Cognitive Psychology , author =. 2016 , pages =. doi:10.1016/j.cogpsych.2016.06.002 , abstract =
2016 doi
-
[28]
Annual Review of Linguistics , author =
Predictability in. Annual Review of Linguistics , author =. 2025 , pages =. doi:10.1146/annurev-linguistics-011724-121517 , abstract =
2025 doi
-
[29]
Bengio, Yoshua and Ducharme, Réjean and Vincent, Pascal and Jauvin, Christian , year =. A
-
[30]
Psychological Review , author =
A. Psychological Review , author =. 2025 , file =. doi:10.31234/osf.io/eaf2z , abstract =
2025 doi
-
[31]
Eye movements reveal a dissociation between prediction and structural processing in language comprehension , language =
Timkey, William and Huang, Kuan-Jung and Oh, Byung-Doh and Prasad, Grusha and Arehalli, Suhas and Linzen, Tal and Dillon, Brian , file =. Eye movements reveal a dissociation between prediction and structural processing in language comprehension , language =
-
[32]
Language, Cognition and Neuroscience , author =
Surprisal in reading: language models predict the. Language, Cognition and Neuroscience , author =. 2025 , pages =. doi:10.1080/23273798.2025.2585303 , abstract =
2025
-
[33]
Language
Radford, Alec and Wu, Jeffrey and Child, Rewon and Luan, David and Amodei, Dario and Sutskever, Ilya , year =. Language
-
[34]
Cognition , author =
The effect of word predictability on reading time is logarithmic , volume =. Cognition , author =. 2013 , pages =. doi:10.1016/j.cognition.2013.02.013 , abstract =
2013 doi
-
[35]
Perspectives on Psychological Sciences , author =
How. Perspectives on Psychological Sciences , author =. 2021 , file =
2021
-
[36]
Cerebral Cortex , author =
Prediction. Cerebral Cortex , author =. 2016 , pages =. doi:10.1093/cercor/bhv075 , abstract =
2016 doi
-
[37]
Association for Computational Linguistics , author =
Geometric. Association for Computational Linguistics , author =. 2025 , file =
2025
-
[38]
Oh, Byung-Doh and Schuler, William , year =. Leading. Proceedings of the 2024. doi:10.18653/v1/2024.emnlp-main.202 , abstract =
2024 doi
- [39]
-
[40]
Neurobiology of Language , author =
Strong. Neurobiology of Language , author =. 2024 , pages =. doi:10.1162/nol_a_00105 , abstract =
2024 doi
-
[41]
Cognitive Science , author =
Lossy‐. Cognitive Science , author =. 2020 , pages =. doi:10.1111/cogs.12814 , abstract =
2020 doi
-
[42]
Reading and Writing , author =
Orthographic knowledge: clarifications, challenges, and future directions , volume =. Reading and Writing , author =. 2019 , pages =. doi:10.1007/s11145-018-9895-9 , abstract =
2019 doi
-
[43]
Vieira, Tim and LeBrun, Benjamin and Giulianelli, Mario and Gastaldi, Juan Luis and DuSell, Brian and Terilla, John and O'Donnell, Timothy J and Cotterell, Ryan , year =. From
- [44]
-
[45]
Computational Brain & Behavior , author =
Plausibility and. Computational Brain & Behavior , author =. 2024 , pages =. doi:10.1007/s42113-024-00196-7 , abstract =
2024 doi
-
[46]
Changing the
Rogers, Anna , year =. Changing the. Proceedings of the 59th. doi:10.18653/v1/2021.acl-long.170 , abstract =
2021 doi
-
[47]
Transactions of the Association for Computational Linguistics , author =
A. Transactions of the Association for Computational Linguistics , author =. 2020 , pages =. doi:10.1162/tacl_a_00349 , abstract =
2020 doi
-
[48]
Language, Cognition and Neuroscience , author =
Towards a computational(ist) neurobiology of language:. Language, Cognition and Neuroscience , author =. 2015 , pages =. doi:10.1080/23273798.2014.980750 , language =
2015
-
[49]
Predictive processing suppresses form-related words with overlapping onsets , abstract =
Haeuser, Katja I and Borovsky, Arielle , year =. Predictive processing suppresses form-related words with overlapping onsets , abstract =
-
[50]
Kuribayashi, Tatsuki and Oseki, Yohei and Ito, Takumi and Yoshida, Ryo and Asahara, Masayuki and Inui, Kentaro , year =. Lower. Proceedings of the 59th. doi:10.18653/v1/2021.acl-long.405 , abstract =
2021 doi
-
[51]
Psychological Bulletin , author =
On. Psychological Bulletin , author =. 1986 , file =
1986
-
[52]
Neural Computing Surveys , author =
Connectionist symbol processing:. Neural Computing Surveys , author =. 1998 , pages =
1998
-
[53]
Open Philosophy , author =
The. Open Philosophy , author =. 2019 , pages =. doi:10.1515/opphil-2019-0018 , abstract =
2019 doi
-
[54]
Psychological
Putnam, Hilary , year =. Psychological. Art,
-
[55]
Synthese , author =
Special. Synthese , author =. 1974 , pages =
1974
-
[56]
Castro, Pablo Samuel , month = oct, year =. The. doi:10.48550/arXiv.2510.16175 , abstract =
-
[57]
and Gauthier, Jon and Hu, Jennifer and Qian, Peng and Levy, Roger P , year =
Wilcox, Ethan G. and Gauthier, Jon and Hu, Jennifer and Qian, Peng and Levy, Roger P , year =. On the
-
[58]
Journal of Memory and Language , author =
Large-scale benchmark yields no evidence that language model surprisal explains syntactic disambiguation difficulty , volume =. Journal of Memory and Language , author =. 2024 , pages =. doi:10.1016/j.jml.2024.104510 , language =
2024
-
[59]
Perspectives on Psychological Science , author =
Theory. Perspectives on Psychological Science , author =. 2021 , file =
2021
-
[60]
Cummins, R , year =. “. In
-
[61]
Psychological Inquiry , author =
Theory. Psychological Inquiry , author =. 2020 , pages =. doi:10.1080/1047840X.2020.1853477 , language =
2020
-
[62]
Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina , year =
-
[63]
Language and Linguistics Compass , author =
Information‐theoretical. Language and Linguistics Compass , author =. 2016 , pages =. doi:10.1111/lnc3.12196 , abstract =
2016 doi
-
[64]
Journal of Memory and Language , author =
Uncovering patterns of semantic predictability in sentence processing , copyright =. Journal of Memory and Language , author =. 2025 , file =. doi:10.31234/osf.io/znkpg , abstract =
2025 doi
-
[65]
Using the
Press, Ofir and Wolf, Lior , year =. Using the. Proceedings of the 15th. doi:10.18653/v1/E17-2025 , abstract =
2025 doi
-
[66]
and McCarthy, Arya D
Jacobs, Cassandra L. and McCarthy, Arya D. , year =. The human unlikeness of neural language models in next-word prediction , url =. Proceedings of the. doi:10.18653/v1/2020.winlp-1.29 , language =
2020 doi
-
[67]
Behavior Research Methods , author =
The. Behavior Research Methods , author =. 2018 , pages =. doi:10.3758/s13428-017-0908-4 , abstract =
2018 doi
-
[68]
Journal of Experimental Psychology: Human Perception and Performance , author =
Performance theories for sentence coding:. Journal of Experimental Psychology: Human Perception and Performance , author =. 1976 , pages =. doi:10.1037/0096-1523.2.1.56 , language =
1976 doi
-
[69]
and Just, Marcel Adam , year =
Kieras, David E. and Just, Marcel Adam , year =. New methods in reading comprehension research , isbn =
-
[70]
The Journalism Quarterly , author =
". The Journalism Quarterly , author =. 1953 , pages =
1953
-
[71]
Cloze but no cigar:
Smith, Nathaniel J and Levy, Roger , year =. Cloze but no cigar:. Cognitive
-
[72]
Journal of Memory and Language , author =
Context-based facilitation of semantic access follows both logarithmic and linear functions of stimulus probability , volume =. Journal of Memory and Language , author =. 2022 , pages =. doi:10.1016/j.jml.2021.104311 , abstract =
2022
-
[73]
Vision - a computational investigation into the human representation and pr , isbn =
Marr, David (late Professor Of Psychology At The Massachuse , year =. Vision - a computational investigation into the human representation and pr , isbn =
-
[74]
Cognition , author =
Expectation-based syntactic comprehension , volume =. Cognition , author =. 2008 , pages =. doi:10.1016/j.cognition.2007.05.006 , abstract =
2008 doi
-
[75]
Bell System Technical Journal , author =
A. Bell System Technical Journal , author =. 1948 , pages =. doi:10.1002/j.1538-7305.1948.tb01338.x , language =
1948
-
[76]
Noisy-context surprisal as a human sentence processing cost model , url =
Futrell, Richard and Levy, Roger , year =. Noisy-context surprisal as a human sentence processing cost model , url =. Proceedings of the 15th. doi:10.18653/v1/E17-1065 , abstract =
-
[77]
Open Mind , author =
The. Open Mind , author =. 2023 , file =
2023
-
[78]
Cognitive Science , author =
Evaluation of an. Cognitive Science , author =. 2024 , pages =. doi:10.1111/cogs.13500 , abstract =
2024 doi
-
[79]
Computational Brain & Behavior , author =
On. Computational Brain & Behavior , author =. 2023 , pages =. doi:10.1007/s42113-022-00166-x , abstract =
2023 doi
-
[80]
Towards a
Jacobs, Cassandra L and Grobol, Morgan , year =. Towards a. Cognitive
-
[81]
Tanenhaus, Michael K , year =. On-. The on-
-
[82]
Brain Research , author =
A linking hypothesis for eyetracking and mousetracking in the visual world paradigm , volume =. Brain Research , author =. 2025 , pages =. doi:10.1016/j.brainres.2025.149477 , abstract =
2025
-
[83]
Proceedings of the American Philosophical Society , author =
A. Proceedings of the American Philosophical Society , author =. 1960 , pages =
1960
-
[84]
Language Research Reports , author =
Augmented. Language Research Reports , author =. 1971 , pages =
1971
-
[85]
Association for Computational Linguistics , author =
Computation of the. Association for Computational Linguistics , author =. 1991 , file =
1991
-
[86]
Journal of Eye Movement Research , author =
Parsing costs as predictors of reading difficulty:. Journal of Eye Movement Research , author =. 2008 , file =. doi:10.16910/jemr.2.1.1 , abstract =
2008 doi
-
[87]
Cognition , author =
Data from eye-tracking corpora as evidence for theories of syntactic processing complexity , volume =. Cognition , author =. 2008 , pages =. doi:10.1016/j.cognition.2008.07.008 , abstract =
2008 doi
-
[88]
Syntactic complexity , copyright =
Frazier, Lyn , editor =. Syntactic complexity , copyright =. Natural. 1985 , doi =
1985
-
[89]
Cognitive Psychology , author =
Making and correcting errors during sentence comprehension:. Cognitive Psychology , author =. 1982 , pages =. doi:10.1016/0010-0285(82)90008-1 , language =
1982 doi
-
[90]
Journal of Memory and Language , author =
Word predictability effects are linear, not logarithmic:. Journal of Memory and Language , author =. 2021 , pages =. doi:10.1016/j.jml.2020.104174 , abstract =
2021
-
[91]
Language, Cognition and Neuroscience , author =
Constraint satisfaction in large language models , volume =. Language, Cognition and Neuroscience , author =. 2024 , pages =. doi:10.1080/23273798.2024.2364339 , abstract =
2024
-
[92]
Language and Cognitive Processes , author =
Probabilistic constraints and syntactic ambiguity resolution , volume =. Language and Cognitive Processes , author =. 1994 , pages =. doi:10.1080/01690969408402115 , language =
1994 doi
-
[93]
Cognitive Science , author =
A. Cognitive Science , author =. 1999 , pages =. doi:10.1207/s15516709cog2304_8 , abstract =
1999 doi
-
[94]
Nature Reviews Psychology , author =
Psychological models and their distractors , volume =. Nature Reviews Psychology , author =. 2022 , pages =. doi:10.1038/s44159-022-00031-5 , language =
2022 doi
-
[95]
and Lieder, Falk and Goodman, Noah D
Griffiths, Thomas L. and Lieder, Falk and Goodman, Noah D. , year = 2015, journal =. Rational. doi:10.1111/tops.12142 , url =
2015 doi
-
[96]
Poggio, Tomaso , year = 2012, journal =. The. doi:10.1068/p7299 , abstract =
2012 doi
-
[97]
, year = 1984, month = jun, publisher =
Pylyshyn, Zenon W. , year = 1984, month = jun, publisher =. Computation and. doi:10.7551/mitpress/2004.001.0001 , url =
1984 doi
-
[98]
Syntactic Complexity , booktitle =
Frazier, Lyn , editor =. Syntactic Complexity , booktitle =. doi:10.1017/CBO9780511597855.005 , url =
-
[99]
Campanelli, Luca and Dyke, Julie Van and Marton, Klara , year = 2018, month = jan, journal =. The
2018
-
[100]
A Decomposition of Surprisal Tracks the
Li, Jiaxuan and Futrell, Richard , year = 2023, url =. A Decomposition of Surprisal Tracks the. Annual
2023
- [101]
-
[102]
Brain Research , series =
Multiple Effects of Sentential Constraint on Word Processing , author =. Brain Research , series =. doi:10.1016/j.brainres.2006.06.101 , url =
2006 doi
-
[103]
and Blank, Idan and Vishnevetsky, Anastasia and Piantadosi, Steven and Fedorenko, Evelina , editor =
Futrell, Richard and Gibson, Edward and Tily, Harry J. and Blank, Idan and Vishnevetsky, Anastasia and Piantadosi, Steven and Fedorenko, Evelina , editor =. The. Proceedings of the
-
[104]
To model human linguistic prediction, make LLMs less superhuman , journal =
Byung-Doh Oh and Tal Linzen , keywords =. To model human linguistic prediction, make LLMs less superhuman , journal =. 2026 , issn =. doi:https://doi.org/10.1016/j.tics.2026.05.008 , url =
2026 doi
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.