Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

The paper argues that surprisal scores drawn from different large language models encode divergent internal computations, so treating them as interchangeable evidence for Surprisal Theory is unsound.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 10:27 UTC pith:WUATTFNU

load-bearing objection Useful critique of interchangeable surprisal, but the layerwise evidence doesn't close the gap to behavioral non-interchangeability. the 3 major comments →

arxiv 2607.20208 v1 pith:WUATTFNU submitted 2026-07-22 cs.CL

surprisal is Not a Theory

classification cs.CL
keywords surprisallarge language modelspsycholinguisticsmultiple realizabilitycomputational-level theoryrepresentation-agnosticismreading timesmodel interchangeability
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The authors aim to show that LLM-derived surprisal is not a single, theory-neutral metric despite its apparent uniformity at the output layer. They argue that each language model instantiates a distinct linking hypothesis between probabilities and processing effort, because models differ in training objective, architecture, and the representations built along the way. Two experiments—tracing next-word probability trajectories across layers and measuring the word-likeness ('lexicality') of models' top predictions—find strikingly different internal behaviors across GPT-2, Pythia, and RoBERTa, even where final correlations to human cloze or reading behavior look similar. If the paper is right, pooling surprisal from many models or choosing models by predictive power alone conflates different algorithmic commitments and undermines claims that Surprisal Theory is a computational-level explanation.

Core claim

The central claim is that the 'representation-agnosticism' often invoked to justify using LLM surprisal as a computational-level measure is untenable. Different LLMs compute next-word probabilities through different latent structures and algorithms; the probabilities themselves are therefore not implementations of one same linking hypothesis. The authors introduce the criterion of functional equivalence: two models' surprisal values are interchangeable only if all subsequent processing is identical. Using logit-lens and lexical-lens probes, they show that although models like GPT-2 and Pythia correlate at 0.92 in final-layer probabilities, the layerwise evolution of those probabilities and t

What carries the argument

The key mechanism is the 'logit lens' and the companion 'lexical lens': probes that read off a model's next-word probability distribution and its top token guesses at each hidden layer, making internal computation visible. The paper uses these to compare GPT-2, Pythia, and RoBERTa, which are similar in parameter count but differ in training objective (decoder vs. masked encoder), tied vs. untied embeddings, and context directionality. The interpretive frame is multiple realizability combined with functional equivalence: if a computational-level theory is allowed to abstract over algorithms, then internal differences matter only when they change downstream processing. The paper argues they do

Load-bearing premise

The argument rests on the premise that two models' surprisal values are interchangeable only when every step of their internal computation is functionally equivalent; if a computational-level theory is allowed to abstract away internal algorithms, then the observed layerwise differences are irrelevant to the interchangeability of the metric.

What would settle it

A decisive test would be to fit the same reading-time regression twice, once with final-layer surprisals from two models whose internal trajectories diverge (e.g., GPT-2 and Pythia) and once with the same final values but shuffled layerwise trajectories; if the coefficients and model fit are statistically indistinguishable, the paper's claim that different internal computations constitute different linking hypotheses would be undermined. Alternatively, showing that the divergence in layerwise trajectories never translates into differences in predicted reading times for any naturalistic corpus

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Researchers should stop treating surprisal estimates from different LLMs as interchangeable in reading-time and neurobehavioral analyses.
  • Model selection in psycholinguistics must be justified by explicit cognitive commitments, not by predictive power alone.
  • Surprisal Theory cannot claim to be a computational-level account while leaving representations and algorithms unexamined.
  • Published comparisons across dozens of models may be conflating distinct linking hypotheses rather than measuring the same construct.
  • The field should move toward explainable, cognitively motivated models, and reviewers should not routinely demand adding LLM surprisal as a covariate.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the authors' premise is right, meta-analyses pooling surprisal across model families may be averaging over qualitatively different computations, which could explain some conflicting results in the surprisal-reading-time literature.
  • The same critique likely applies to other LLM-extracted psycholinguistic indices (e.g., attention weights, entropy), not just surprisal.
  • A concrete extension: researchers could vary architecture while matching final surprisal values, to test whether layerwise trajectory differences actually change reading-time predictions—if they don't, the practical significance of this argument would shrink.
  • The argument could be transformed into a positive methodological program: use layerwise probes to determine which representational commitments are necessary to replicate human processing, rather than treating any LLM as neutral.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper argues that the common psycholinguistic practice of treating surprisal values extracted from different large language models as interchangeable is theoretically unjustified. It reports two analyses using GPT-2, Pythia-160M, and RoBERTa: first, layerwise logit-lens probabilities are correlated with human cloze probabilities (Experiment 1); second, the 'lexicality' of each model's top-k predictions is traced across layers (Experiment 2). The authors find that layerwise trajectories differ markedly across models and conclude that different architectures and training objectives instantiate different algorithmic and representational commitments, undermining representation-agnosticism in Surprisal Theory. The paper then develops a conceptual critique of using LLM-derived surprisal as a computational-level explanation and offers recommendations for theory-driven model selection.

Significance. The paper addresses a timely and important methodological question: whether results obtained with one LLM's surprisal values can be taken to test Surprisal Theory independently of the model that generated them. Its strengths include public data and scripts (OSF), an emphasis on falsifiable practice rather than fitted parameters, and a clear engagement with Marr's levels. If the central claim were established, it would make a meaningful contribution to computational psycholinguistics. However, the significance is conditional: the empirical results demonstrate differences in internal algorithmic trajectories, but the operationalized surprisal in the literature is the final-layer probability, and the paper itself reports that final-layer GPT-2 and Pythia probabilities correlate at 0.92. The gap between algorithmic divergence and non-interchangeable linking hypotheses is the load-bearing issue that must be addressed.

major comments (3)
  1. [§2.3 and §4.1] The central inference from divergent layerwise computations to non-interchangeable surprisal-based theories is not supported by the presented evidence. Study 1.1 reports that final-layer GPT-2 and Pythia probabilities correlate at 0.92, and final-layer correlations to cloze are similar across models. Surprisal as commonly operationalized is the final-layer negative log probability, and Marr's computational level explicitly abstracts over algorithmic implementation—a point the paper itself acknowledges in §4.1 ('multiple-realizability is intrinsic to computational-level theories'). The statement at the end of §4.1, 'If different LLMs change the linking hypothesis between probabilities and effort...', is exactly the untested premise. To support the paper's headline claim, the authors need to show that final surprisal values differ across models in ways that change behavioral predictions (e
  2. [§2.4–2.5 and §3.2] The claims of 'starkly different' trajectories (Figure 3) and 'massive fluctuations' in lexicality (Figure 4) rest on visual inspection of only three models, with no error bars, confidence intervals, or inferential statistics on the layerwise curves. This is particularly consequential because the qualitative reading of the curves is a key part of the argument that the models are not interchangeable. Additionally, the lexicality metric depends on the choice of k=10 in §3.2, but no sensitivity analysis is reported; different tokenizers across models could confound the comparison. I recommend quantifying trajectory differences (e.g., bootstrap CIs, pairwise curve-distance measures) and testing whether the lexicality patterns are robust to k and to tokenization.
  3. [§4.4] The discussion of the Natural Stories misalignment error is presented as evidence that the field's reliance on surprisal is theoretically fragile. This is an interesting argument, but the inference is not logically entailed: a correlation can survive a one-position shift for many reasons (e.g., autocorrelation in surprisal, smoothness of the cost function). The claim that the error 'should have led researchers to reconsider why exactly their surprisal analyses had been so successful' is speculative. This is not the load-bearing part of the paper, but it should be framed as an open question or a testable prediction rather than a necessary consequence.
minor comments (4)
  1. [§1] Typo: 'effors' should be 'efforts'. Also, the phrase 'these assumptions are is in fact not wholly valid' has a grammar error.
  2. [Title] The title 'surprisal is Not a Theory' overstates the paper's scope; the argument is specifically about LLM-derived surprisal and its use as a computational-level proxy. A more precise title would help manage reader expectations.
  3. [Figure 3] The figure would be more informative with confidence bands or per-context variability; as published, the eye is drawn to mean trajectories that may mask high variance.
  4. [§3.2 / Table 2] The tokenization differences across models make the qualitative examples in Table 2 hard to interpret; a brief note on subword segmentation would improve clarity.

Circularity Check

0 steps flagged

No significant circularity: the paper's claims rest on independent empirical probes of public models and corpora, not on fitted parameters or load-bearing self-citations.

full rationale

The paper contains no first-principles derivation, no fitted parameter that is later renamed a prediction, and no uniqueness theorem or ansatz imported from the authors' prior work. Its central evidence is empirical: layerwise logit-lens comparisons against Peelle et al. (2020) cloze norms (Experiment 1) and a new 'lexicality' metric applied to public models and the Provo Corpus (Experiment 2). These comparisons are externally checkable and do not reduce by construction to the conclusion. The authors' own result that final-layer GPT-2 and Pythia probabilities correlate at 0.92 and that final-layer cloze correlations are similar across models actually undercuts the strongest version of their 'not interchangeable' claim, which is an interpretive overreach rather than a circular step. The self-citations (Jacobs & McCarthy 2020, Jacobs et al. 2025, De Santo 2025, Jacobs & Grobol 2025) are contextual or auxiliary, not load-bearing: the median-split replication from Jacobs & McCarthy (2020) is re-demonstrated with the paper's own data. The functional-equivalence criterion is imported from Rogers (2021), an external source, and the contested inference—that different internal layerwise trajectories imply different computational-level linking hypotheses—is an unsupported empirical/philosophical argument, not a definitional equivalence. The conditional statement in §4.1 ('If different LLMs change the linking hypothesis between probabilities and effort...') is acknowledged as a premise, not presented as a derived result. Thus no circular step can be exhibited with the required specificity, and the appropriate finding is no significant circularity.

Axiom & Free-Parameter Ledger

1 free parameters · 3 axioms · 0 invented entities

No theoretical entities are posited. The paper introduces a new metric (lexicality) and a hand-chosen analysis parameter (k=10). Its argument relies on domain assumptions about Marr's levels and functional equivalence, plus the validity of interpretability probes.

free parameters (1)
  • k (top-k in lexicality metric) = 10
    The lexicality metric is defined for the top k guesses; k=10 is chosen without a robustness analysis, and results could differ with k.
axioms (3)
  • domain assumption Marr's computational level requires that a linking hypothesis be transparent about representational and algorithmic commitments to be explanatory.
    The paper's claim that LLM-surprisal is not a good computational-level theory depends on this specific reading of Marr (1982), invoked in §4.1. Alternative readings allow computational-level theories to be implementation-agnostic.
  • domain assumption Functional equivalence requires that all processing after a given point be identical, so different latent representations make surprisal non-interchangeable.
    Stated in §1 after citing Rogers (2021); this is the load-bearing premise for why different layerwise computations undermine representation-agnosticism.
  • domain assumption The logit lens and lexicality metric accurately reflect the information used by a model to compute next-word probabilities.
    The empirical arguments in §2-3 assume these interpretability probes are valid windows into the models' computations.

pith-pipeline@v1.3.0-alltime-deepseek · 21384 in / 10049 out tokens · 87559 ms · 2026-08-01T10:27:03.487327+00:00 · methodology

0 comments
read the original abstract

Surprisal Theory is often characterized as a computational-level explanation per (Marr, 1982). We argue in this work that, even though a computational level narrative has been used to support "representation-agnostic research" within computational psycholinguistics, the movement toward black box systems embodied by large language models (LLMs) does not exempt modelers using the surprisal metric from the representational decisions required by computational-level characterizations. In fact, we argue that the uncritical use of LLM-surprisal obfuscates the representational and algorithmic-level commitments of different models. In three analyses, we show that the choice of algorithm and model architecture play significant roles in the computation of language model probabilities. We advise that researchers who wish to test Surprisal Theory re-evaluate the practice of treating large language model probabilities as interchangeable

Figures

Figures reproduced from arXiv: 2607.20208 by Andr\'es Bux\'o-Lugo, Aniello De Santo, Cassandra L. Jacobs, Morgan Grobol, Ryan J. Hubbard.

Figure 1
Figure 1. Figure 1: Spearman’s correlation between word model-estimated surprisal and cloze prob￾ability, across models and layers. correlations are stronger at higher cloze probability levels (𝐵upper = 0.04, 𝑡(35456) = 45.78, 𝑝 < 0.001) ,the item-level correspondences are highly vari￾abl, replicating a prior finding from Jacobs and McCarthy (2020) [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Demonstration of correlation between language model and human cloze prob￾ability in Peelle et al.’s (2020) completion norms for GPT-2’s first and final layers. Dark points represent cloze level-specific means; lighter points represent item-specific values. Solid line represents 𝑦 = 𝑥, indicating a perfect correlation. Model probabilities are faceted by layer (first on the left vs. last on the right). how t… view at source ↗
Figure 3
Figure 3. Figure 3: Next-word probabilities through the logit lens across layers and models. 11 [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Layerwise lexicality changes for 𝑘 = 10 over the 2398 preambles of the Provo Corpus (Luke & Christianson, 2016, 2018). 16 [PITH_FULL_IMAGE:figures/full_fig_p016_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Surprisal Theory is Tautological (without Rational Grounding)

    cs.CL 2026-07 conditional novelty 6.0

    Unconstrained surprisal theory is a tautology: for any non-negative difficulty measure, a language model exists whose surprisal matches it affinely.

Reference graph

Works this paper leans on

104 extracted references · 20 canonical work pages · cited by 1 Pith paper

  1. [1]

    Behavior Research Methods , author =

    Moving beyond. Behavior Research Methods , author =. 2009 , pages =. doi:10.3758/BRM.41.4.977 , language =

  2. [2]

    Transactions of the Association for Computational Linguistics , volume =

    Oh, Byung-Doh and Schuler, William , title =. Transactions of the Association for Computational Linguistics , volume =. 2023 , month =. doi:10.1162/tacl_a_00548 , url =

  3. [3]

    2022 , eprint=

    On the Opportunities and Risks of Foundation Models , author=. 2022 , eprint=

  4. [4]

    Behavioral and Brain Sciences , volume=

    Machine yearning: LLMs do not capture formal linguistic structure and obscure neuroscientific inquiry , author=. Behavioral and Brain Sciences , volume=. 2026 , publisher=

  5. [5]

    Behavioral and Brain Sciences , volume=

    Are language models models? , author=. Behavioral and Brain Sciences , volume=. 2026 , publisher=

  6. [6]

    Behavioral and Brain Sciences , volume=

    Language models do not yet achieve explanatory adequacy because language is more than incremental prediction , author=. Behavioral and Brain Sciences , volume=. 2026 , publisher=

  7. [7]

    Proceedings of the 30th Conference on Computational Natural Language Learning , pages=

    Linguistic Profiling of Transformer Embedding Geometry , author=. Proceedings of the 30th Conference on Computational Natural Language Learning , pages=

  8. [8]

    2026 , eprint=

    Trajectory Dynamics in Language Model Hidden States Predict Human Processing Costs Beyond Surprisal , author=. 2026 , eprint=

  9. [9]

    Behavior Research Methods , author =

    Cloze probability, predictability ratings, and computational estimates for 205. Behavior Research Methods , author =. 2023 , pages =. doi:10.3758/s13428-023-02261-8 , abstract =

  10. [10]

    Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics , pages=

    Capturing Online SRC/ORC Effort with Memory Measures from a Minimalist Parser , author=. Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics , pages=

  11. [11]

    Anisotropy is

    Machina, Anemily and Mercer, Robert , editor =. Anisotropy is. Proceedings of the 2024. 2024 , pages =. doi:10.18653/v1/2024.naacl-long.274 , abstract =

  12. [12]

    Transactions of the Association for Computational Linguistics , author =

    Testing the. Transactions of the Association for Computational Linguistics , author =. 2023 , note =. doi:10.1162/tacl_a_00612 , abstract =

  13. [13]

    Nair, Sathvik and Resnik, Philip , editor =. Words,. Findings of the. 2023 , pages =. doi:10.18653/v1/2023.findings-emnlp.752 , abstract =

  14. [14]

    Language

    Wilcox, Ethan and Meister, Clara and Cotterell, Ryan and Pimentel, Tiago , editor =. Language. Proceedings of the 2023. 2023 , pages =. doi:10.18653/v1/2023.emnlp-main.466 , abstract =

  15. [15]

    Analyzing

    Dar, Guy and Geva, Mor and Gupta, Ankit and Berant, Jonathan , editor =. Analyzing. Proceedings of the 61st. 2023 , pages =. doi:10.18653/v1/2023.acl-long.893 , abstract =

  16. [16]

    Wilcox, Ethan and Vani, Pranali and Levy, Roger , editor =. A. Proceedings of the 59th. 2021 , pages =. doi:10.18653/v1/2021.acl-long.76 , abstract =

  17. [17]

    Surprisal

    Oh, Byung-Doh and Clark, Christian and Schuler, William , editor =. Surprisal. Proceedings of the 59th. 2021 , pages =. doi:10.18653/v1/2021.acl-long.290 , abstract =

  18. [18]

    Proceedings of the 57th

    Tenney, Ian and Das, Dipanjan and Pavlick, Ellie , editor =. Proceedings of the 57th. 2019 , pages =. doi:10.18653/v1/P19-1452 , abstract =

  19. [19]

    Predictive power of word surprisal for reading times is a linear function of language model quality , url =

    Goodkind, Adam and Bicknell, Klinton , editor =. Predictive power of word surprisal for reading times is a linear function of language model quality , url =. 2018 , pages =. doi:10.18653/v1/W18-0102 , booktitle =

  20. [20]

    and Otten, Leun J

    Frank, Stefan L. and Otten, Leun J. and Galli, Giulia and Vigliocco, Gabriella , editor =. Word surprisal predicts. Proceedings of the 51st. 2013 , pages =

  21. [21]

    Robustness and processing difficulty models

    Rauzy, Stéphane and Blache, Philippe , editor =. Robustness and processing difficulty models. Proceedings of the. 2012 , pages =

  22. [22]

    Hale, John , year =. A. Second

  23. [23]

    Computational Brain & Behavior , author =

    What’s. Computational Brain & Behavior , author =. 2025 , pages =. doi:10.1007/s42113-025-00237-9 , abstract =

  24. [24]

    Behavior Research Methods , author =

    Completion norms for 3085. Behavior Research Methods , author =. 2020 , pages =. doi:10.3758/s13428-020-01351-1 , abstract =

  25. [25]

    Proceedings of the National Academy of Sciences , author =

    Large-scale evidence for logarithmic effects of word predictability on reading time , volume =. Proceedings of the National Academy of Sciences , author =. 2024 , pages =. doi:10.1073/pnas.2307876121 , abstract =

  26. [26]

    International Conference on Machine Learning , author =

    Pythia:. International Conference on Machine Learning , author =. 2023 , pages =

  27. [27]

    Cognitive Psychology , author =

    Limits on lexical prediction during reading , volume =. Cognitive Psychology , author =. 2016 , pages =. doi:10.1016/j.cogpsych.2016.06.002 , abstract =

  28. [28]

    Annual Review of Linguistics , author =

    Predictability in. Annual Review of Linguistics , author =. 2025 , pages =. doi:10.1146/annurev-linguistics-011724-121517 , abstract =

  29. [29]

    Bengio, Yoshua and Ducharme, Réjean and Vincent, Pascal and Jauvin, Christian , year =. A

  30. [30]

    Psychological Review , author =

    A. Psychological Review , author =. 2025 , file =. doi:10.31234/osf.io/eaf2z , abstract =

  31. [31]

    Eye movements reveal a dissociation between prediction and structural processing in language comprehension , language =

    Timkey, William and Huang, Kuan-Jung and Oh, Byung-Doh and Prasad, Grusha and Arehalli, Suhas and Linzen, Tal and Dillon, Brian , file =. Eye movements reveal a dissociation between prediction and structural processing in language comprehension , language =

  32. [32]

    Language, Cognition and Neuroscience , author =

    Surprisal in reading: language models predict the. Language, Cognition and Neuroscience , author =. 2025 , pages =. doi:10.1080/23273798.2025.2585303 , abstract =

  33. [33]

    Language

    Radford, Alec and Wu, Jeffrey and Child, Rewon and Luan, David and Amodei, Dario and Sutskever, Ilya , year =. Language

  34. [34]

    Cognition , author =

    The effect of word predictability on reading time is logarithmic , volume =. Cognition , author =. 2013 , pages =. doi:10.1016/j.cognition.2013.02.013 , abstract =

  35. [35]

    Perspectives on Psychological Sciences , author =

    How. Perspectives on Psychological Sciences , author =. 2021 , file =

  36. [36]

    Cerebral Cortex , author =

    Prediction. Cerebral Cortex , author =. 2016 , pages =. doi:10.1093/cercor/bhv075 , abstract =

  37. [37]

    Association for Computational Linguistics , author =

    Geometric. Association for Computational Linguistics , author =. 2025 , file =

  38. [38]

    Oh, Byung-Doh and Schuler, William , year =. Leading. Proceedings of the 2024. doi:10.18653/v1/2024.emnlp-main.202 , abstract =

  39. [39]

    doi:10.48550/arXiv.1907.11692 , abstract =

    Liu, Yinhan and Ott, Myle and Goyal, Naman and Du, Jingfei and Joshi, Mandar and Chen, Danqi and Levy, Omer and Lewis, Mike and Zettlemoyer, Luke and Stoyanov, Veselin , month = jul, year =. doi:10.48550/arXiv.1907.11692 , abstract =

  40. [40]

    Neurobiology of Language , author =

    Strong. Neurobiology of Language , author =. 2024 , pages =. doi:10.1162/nol_a_00105 , abstract =

  41. [41]

    Cognitive Science , author =

    Lossy‐. Cognitive Science , author =. 2020 , pages =. doi:10.1111/cogs.12814 , abstract =

  42. [42]

    Reading and Writing , author =

    Orthographic knowledge: clarifications, challenges, and future directions , volume =. Reading and Writing , author =. 2019 , pages =. doi:10.1007/s11145-018-9895-9 , abstract =

  43. [43]

    Vieira, Tim and LeBrun, Benjamin and Giulianelli, Mario and Gastaldi, Juan Luis and DuSell, Brian and Terilla, John and O'Donnell, Timothy J and Cotterell, Ryan , year =. From

  44. [44]

    To model human linguistic prediction, make

    Oh, Byung-Doh and Linzen, Tal , month = oct, year =. To model human linguistic prediction, make. doi:10.48550/arXiv.2510.05141 , abstract =

  45. [45]

    Computational Brain & Behavior , author =

    Plausibility and. Computational Brain & Behavior , author =. 2024 , pages =. doi:10.1007/s42113-024-00196-7 , abstract =

  46. [46]

    Changing the

    Rogers, Anna , year =. Changing the. Proceedings of the 59th. doi:10.18653/v1/2021.acl-long.170 , abstract =

  47. [47]

    Transactions of the Association for Computational Linguistics , author =

    A. Transactions of the Association for Computational Linguistics , author =. 2020 , pages =. doi:10.1162/tacl_a_00349 , abstract =

  48. [48]

    Language, Cognition and Neuroscience , author =

    Towards a computational(ist) neurobiology of language:. Language, Cognition and Neuroscience , author =. 2015 , pages =. doi:10.1080/23273798.2014.980750 , language =

  49. [49]

    Predictive processing suppresses form-related words with overlapping onsets , abstract =

    Haeuser, Katja I and Borovsky, Arielle , year =. Predictive processing suppresses form-related words with overlapping onsets , abstract =

  50. [50]

    Kuribayashi, Tatsuki and Oseki, Yohei and Ito, Takumi and Yoshida, Ryo and Asahara, Masayuki and Inui, Kentaro , year =. Lower. Proceedings of the 59th. doi:10.18653/v1/2021.acl-long.405 , abstract =

  51. [51]

    Psychological Bulletin , author =

    On. Psychological Bulletin , author =. 1986 , file =

  52. [52]

    Neural Computing Surveys , author =

    Connectionist symbol processing:. Neural Computing Surveys , author =. 1998 , pages =

  53. [53]

    Open Philosophy , author =

    The. Open Philosophy , author =. 2019 , pages =. doi:10.1515/opphil-2019-0018 , abstract =

  54. [54]

    Psychological

    Putnam, Hilary , year =. Psychological. Art,

  55. [55]

    Synthese , author =

    Special. Synthese , author =. 1974 , pages =

  56. [56]

    Castro, Pablo Samuel , month = oct, year =. The. doi:10.48550/arXiv.2510.16175 , abstract =

  57. [57]

    and Gauthier, Jon and Hu, Jennifer and Qian, Peng and Levy, Roger P , year =

    Wilcox, Ethan G. and Gauthier, Jon and Hu, Jennifer and Qian, Peng and Levy, Roger P , year =. On the

  58. [58]

    Journal of Memory and Language , author =

    Large-scale benchmark yields no evidence that language model surprisal explains syntactic disambiguation difficulty , volume =. Journal of Memory and Language , author =. 2024 , pages =. doi:10.1016/j.jml.2024.104510 , language =

  59. [59]

    Perspectives on Psychological Science , author =

    Theory. Perspectives on Psychological Science , author =. 2021 , file =

  60. [60]

    Cummins, R , year =. “. In

  61. [61]

    Psychological Inquiry , author =

    Theory. Psychological Inquiry , author =. 2020 , pages =. doi:10.1080/1047840X.2020.1853477 , language =

  62. [62]

    Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina , year =

  63. [63]

    Language and Linguistics Compass , author =

    Information‐theoretical. Language and Linguistics Compass , author =. 2016 , pages =. doi:10.1111/lnc3.12196 , abstract =

  64. [64]

    Journal of Memory and Language , author =

    Uncovering patterns of semantic predictability in sentence processing , copyright =. Journal of Memory and Language , author =. 2025 , file =. doi:10.31234/osf.io/znkpg , abstract =

  65. [65]

    Using the

    Press, Ofir and Wolf, Lior , year =. Using the. Proceedings of the 15th. doi:10.18653/v1/E17-2025 , abstract =

  66. [66]

    and McCarthy, Arya D

    Jacobs, Cassandra L. and McCarthy, Arya D. , year =. The human unlikeness of neural language models in next-word prediction , url =. Proceedings of the. doi:10.18653/v1/2020.winlp-1.29 , language =

  67. [67]

    Behavior Research Methods , author =

    The. Behavior Research Methods , author =. 2018 , pages =. doi:10.3758/s13428-017-0908-4 , abstract =

  68. [68]

    Journal of Experimental Psychology: Human Perception and Performance , author =

    Performance theories for sentence coding:. Journal of Experimental Psychology: Human Perception and Performance , author =. 1976 , pages =. doi:10.1037/0096-1523.2.1.56 , language =

  69. [69]

    and Just, Marcel Adam , year =

    Kieras, David E. and Just, Marcel Adam , year =. New methods in reading comprehension research , isbn =

  70. [70]

    The Journalism Quarterly , author =

    ". The Journalism Quarterly , author =. 1953 , pages =

  71. [71]

    Cloze but no cigar:

    Smith, Nathaniel J and Levy, Roger , year =. Cloze but no cigar:. Cognitive

  72. [72]

    Journal of Memory and Language , author =

    Context-based facilitation of semantic access follows both logarithmic and linear functions of stimulus probability , volume =. Journal of Memory and Language , author =. 2022 , pages =. doi:10.1016/j.jml.2021.104311 , abstract =

  73. [73]

    Vision - a computational investigation into the human representation and pr , isbn =

    Marr, David (late Professor Of Psychology At The Massachuse , year =. Vision - a computational investigation into the human representation and pr , isbn =

  74. [74]

    Cognition , author =

    Expectation-based syntactic comprehension , volume =. Cognition , author =. 2008 , pages =. doi:10.1016/j.cognition.2007.05.006 , abstract =

  75. [75]

    Bell System Technical Journal , author =

    A. Bell System Technical Journal , author =. 1948 , pages =. doi:10.1002/j.1538-7305.1948.tb01338.x , language =

  76. [76]

    Noisy-context surprisal as a human sentence processing cost model , url =

    Futrell, Richard and Levy, Roger , year =. Noisy-context surprisal as a human sentence processing cost model , url =. Proceedings of the 15th. doi:10.18653/v1/E17-1065 , abstract =

  77. [77]

    Open Mind , author =

    The. Open Mind , author =. 2023 , file =

  78. [78]

    Cognitive Science , author =

    Evaluation of an. Cognitive Science , author =. 2024 , pages =. doi:10.1111/cogs.13500 , abstract =

  79. [79]

    Computational Brain & Behavior , author =

    On. Computational Brain & Behavior , author =. 2023 , pages =. doi:10.1007/s42113-022-00166-x , abstract =

  80. [80]

    Towards a

    Jacobs, Cassandra L and Grobol, Morgan , year =. Towards a. Cognitive

Showing first 80 references.