Pith. sign in

REVIEW 2 major objections 4 minor 68 references

Surprisal theory is tautological when the language model is left unconstrained, because any non-negative difficulty measure can be turned into a language model whose surprisal reproduces it exactly (up to a context-dependent additive consta

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 07:02 UTC pith:2AAF4HI2

load-bearing objection The tautology construction is sound and worth taking seriously, but the paper's empirical claim that the corpus assumption is falsified is not entailed by its own formal apparatus. the 2 major comments →

arxiv 2607.21574 v1 pith:2AAF4HI2 submitted 2026-07-23 cs.CL

Surprisal Theory is Tautological (without Rational Grounding)

classification cs.CL
keywords surprisal theorylanguage modelstautologyfalsifiabilityprocessing difficultycorpus assumptionscaling implicationpsycholinguistics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that surprisal theory, as usually stated, is a tautology: for any non-negative measure of processing difficulty, there exists a language model whose surprisal equals that difficulty plus a context-dependent constant. The construction is explicit: define a model whose next-unit probabilities are proportional to exp(-difficulty). Under a mild boundedness condition this is a genuine probability distribution. Consequently, unless the language model is constrained from outside the reading-time data, any empirical fit of surprisal to reading times is uninformative. The paper identifies the corpus assumption—that the model is the distribution generating the training corpus—as the constraint that historically made the theory testable, and argues recent scaling results undermine that assumption, so the theory needs a 'rationalist' grounding in independently motivated properties of the comprehender.

Core claim

The paper's central claim is that the equation 'difficulty = -a log q + b' is vacuous when q is free. For any non-negative difficulty measure d, setting q_d(u|context) = exp(-d(u|context)) / sum_{u'} exp(-d(u'|context)) gives -log q_d = d + log Z(context), which is exactly the affine form with slope 1. Since every pattern of difficulty can be reproduced by some language model, no experiment can distinguish surprisal theory from any other theory of difficulty unless the family of models and the criterion for choosing within it are constrained by something other than the behavioral data. The paper proves a mild sufficient condition (bounded sentence-final difficulty) under which the constructi

What carries the argument

Equation (6), the softmax-like construction q_d(unit|context) = exp(-d(unit|context)) / partition function, which converts any difficulty measure into a language model. The identity it induces, -log q_d = d + log Z(context), is exactly the affine surprisal equation with slope 1 and context-dependent intercept. Proposition 1 supplies a sufficient condition (bounded EOS difficulty) for the construction to be a tight probability distribution with finite expected length; the fitness measure M(q) and the consistency theorem translate the corpus assumption into the 'scaling implication' that better corpus fit yields better psychometric fit.

Load-bearing premise

That recent scaling results falsify the corpus assumption—rather than merely showing that today's largest models are still misspecified—assumes those models approximate the true text-generating distribution itself and not just the best available fit within the model family; the paper itself calls this an assumption, not a theorem.

What would settle it

Measure sentence-final processing difficulty over contexts of increasing length: if wrap-up cost grows with context length, the bounded-difficulty condition behind Proposition 1 fails and the construction may leak probability mass, so the claim that every difficulty pattern is reproducible would require repair. Conversely, if models trained on the same corpus genre as the reading-time data show monotonically improving psychometric fit with size, the empirical case against the corpus assumption would be undercut.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Any reported correlation between a language model's surprisal and human reading times is uninformative about surprisal theory itself unless the model family and selection criterion are justified independently of that behavioral data.
  • The corpus assumption, not surprisal theory per se, is what made the framework empirically testable; if it is false, the empirical basis of two decades of supporting work is called into question.
  • Scaling results showing that larger corpus-trained models fit reading times worse constitute evidence against the corpus assumption, provided the models are close enough to the true text distribution.
  • The theory can be made falsifiable again by deriving the language model from independently motivated constraints such as memory limitations or processing goals—the rationalist route—rather than from corpus fit.
  • The construction applies to any theory that defines difficulty purely as a function of a probability distribution, so the tautology problem generalizes beyond surprisal.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the tautology argument implies that head-to-head comparisons of language models in psycholinguistics should be read as evidence about inductive biases of model classes, not as evidence for or against surprisal theory.
  • Editorial extension: a testable consequence of the rationalist proposal is that psychometric fit should improve along a dimension defined by cognitive plausibility (e.g., memory-bounded architectures) rather than by corpus fit; this could be evaluated by comparing fits across architectures matched for corpus performance.
  • Editorial extension: the paper's logic also suggests that behavioral-data fine-tuning of language models cannot escape the tautology, because the model is then chosen using the very data to be explained—an implication the paper gestures at but does not develop.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper argues that surprisal theory—the claim that processing difficulty d is an affine function of surprisal under some language model q—is unfalsifiable/tautological when q is unconstrained. The central construction is Eq. (6): for any non-negative difficulty measure d, define q_d by the softmax exp(-d)/Z; then -log q_d = d + log Z, so Eq. (1) holds with a=1 and b=-log Z. Proposition 1 gives a mild sufficient condition under which q_d is a genuine tight language model. The paper then formalizes a 'scaling implication' (Proposition 2): under MLE consistency, a fitness measure M(q) converges to M(q*), the KL projection. It uses recent findings (Oh and Schuler, 2023; Oh et al., 2024; Lin and Schuler, 2025) that larger language models produce worse reading-time fits to argue that the corpus assumption is falsified, concluding that surprisal theory requires a rationalist grounding of q independent of the behavioral data.

Significance. The central formal observation is clean and, as far as I can tell, correct: Eq. (6) is a valid construction, and Proposition 1 supplies a precise sufficient condition for tightness. The paper makes a valuable conceptual contribution by isolating the sense in which unconstrained surprisal theory is vacuous and by framing the 'grounding problem' as a choice of a constrained model family Q. The formal proofs in the appendix are self-contained and the paper is unusually candid about its assumptions and limitations. The empirical half of the paper, however, is substantially weaker than the formal half. The inference from the Oh-Schuler reversal to the falsification of the corpus assumption depends on properties of the fitness measure M that are not stated in Definition 1, and on an unformalized identification of model scale with MLE consistency. Those gaps are load-bearing for the paper's empirical conclusion, though not for the tautology construction itself.

major comments (2)
  1. [§4.2, Definition 1, Eq. (8)] The claim that the corpus assumption predicts that p_C maximizes M is not derived from Eq. (8). In Definition 1, M(q) = Σ p(uu) g(d(u|u), -log q(u|u)) with no regression coefficients or context baselines. If d satisfies Eq. (4), d = -a log p_C + b(u_{<t}), then the softmax model q_d of Eq. (6) satisfies -log q_d = d + log Z(u_{<t}), so its surprisal coincides with d up to an additive constant, whereas p_C's surprisal is offset by b and scaled by a. For any reasonable g that rewards closeness—e.g., squared error—q_d will score strictly higher than p_C. Hence a rise-and-fall of M as q_N approaches p_C is compatible with Eq. (4) holding exactly for p_C; it does not falsify the corpus assumption unless M is defined as the best fit after estimating a and b (e.g., via regression with context controls, as footnote 6 suggests but Definition 1 does not implement).
  2. [§4.2 / Proposition 2] Proposition 2 concerns a fixed model family Q and MLE as N → ∞. The Oh-Schuler findings concern a sequence of models of increasing size/capacity trained on different data amounts; the paper provides no theorem or formal argument that this sequence converges to p_C or to the KL projection under the conditions of Theorem 1. The statement in §4.2 that 'as models approach p_C' a reversal is evidence against the corpus assumption relies on an additional, unformalized identification of model scale with estimation quality. Footnote 21 acknowledges one aspect (KL projection vs p_C), but not the model-size/N gap. Without this link, the empirical case against the corpus assumption is not fully made.
minor comments (4)
  1. [Abstract / §1] There is a spacing artifact in the abstract: 'consistent withsomelan- guage model' should be 'consistent with some language model.'
  2. [Definition 1, Eq. (8)] The notation u is used both for the next unit and for the context string; this is confusing in the double sum. Consider using a distinct symbol, e.g., x for the unit and w for the context.
  3. [Appendix B.1] There is a typo: 'noteE qd[|y|]≤...' should read 'Note that E_{q_d}[|y|] ≤ ...'.
  4. [Footnote 16] The statement 'if d(u|u)≤C' should specify whether the first argument is a unit in Σ_E or a string; the notation is overloaded.

Circularity Check

0 steps flagged

No significant circularity: the vacuity construction is the paper's object of study, its self-citations are independent support, and the flagged gaps are support/correctness concerns rather than circular reductions.

full rationale

The paper's central claim is a meta-argument demonstrating that unconstrained surprisal theory is vacuous, and its derivation chain does not reduce to its own inputs. The construction in Eq. (6), q_d(u|u) ∝ exp(-d(u|u)), delivers -log q_d = d + log Z by pure algebra; the paper presents this by-construction equality as the demonstrated tautology of the target theory, not as a hidden premise or as an empirical prediction. The self-definitional circularity it exposes belongs to surprisal theory itself, which is exactly the paper's point — the paper commits no circularity by exhibiting it. The technical lemmas used to upgrade q_d to a genuine language model are cited from the author's own prior work (Du et al. 2023, Prop. 4.3, for conditional-collection tightness in the proof of Prop. 1; Opedal et al. 2024, Prop. 1, for the sum of prefix probabilities in Lemma 1's proof), but both are general, published, parameter-free statements with no reference to difficulty measures or surprisal; per the review rules such citations are real, independent support and do not raise the circularity score. The affine form with context-dependent b, which lets the construction absorb log Z, is justified by Shain et al. (2024), external large-scale reading-time evidence, and by standard regression practice; no ansatz is smuggled in as a self-cited axiom. What the manuscript itself flags is missing support, not circularity: footnote 21 concedes that the inference from Oh–Schuler reversals to falsification of the corpus assumption 'is an assumption, not a theorem,' and the assertion in §4.2 that the corpus assumption predicts p_C maximizes M is not derived from Eq. (4) plus Definition 1 when b varies by context; the Limitations section similarly admits the global-MLE idealization and the unverified mildness of Prop. 1's conditions. These are validity/strength concerns about the empirical conclusion, not reductions of the derivation to its premises, and they do not touch the formal tautology result. Score 2 reflects only the presence of minor self-citations in the derivation chain; no circular step was identified.

Axiom & Free-Parameter Ledger

1 free parameters · 7 axioms · 0 invented entities

The proof of the central vacuity result rests only on standard probability theory plus the mild bounded-wrap-up condition and an external tightness characterization. The stronger claims about the corpus assumption additionally rely on external scaling studies and the well-specification assumption flagged in fn. 21. No new physical or mechanistic entities are introduced.

free parameters (1)
  • Affine parameters a and b(u<t) in Eq. (1) = a=1, b=-log Z in the tautology construction
    These are the degrees of freedom of the theory under critique. The paper shows they can always be chosen to fit any d; they are not fitted by the author to data, but they are exactly why unconstrained surprisal is vacuous.
axioms (7)
  • standard math Tightness of conditional collections is characterized by lower bounds on EOS probability with divergent sum (Du et al. 2023, Prop. 4.3).
    Used in Appendix B.1 to upgrade the pointwise construction of q_d to a genuine language model with total mass 1.
  • domain assumption Processing difficulty d is a single non-negative function of unit and context shared across readers.
    The paper adopts the standard linear mixed-effects abstraction (footnote 4); if d were reader-specific, Eq. (6) would produce a different q per reader, and the single-q tautology would not apply to raw reading times.
  • domain assumption d(EOS|u) is uniformly bounded by a constant C, i.e., sentence-wrap-up cost does not grow without bound with context length.
    Needed for Proposition 1(b) and finite expected length. The author says it likely holds for realistic difficulty measures but has not verified it on psychometric datasets (Limitations).
  • domain assumption The corpus distribution p_C has finite expected length and per-unit probabilities are bounded below by ε.
    Required for the fitness measure M (Definition 1), Lemma 1, and the uniform domination condition in Theorem 1; standard for softmax models with compact parameter sets.
  • standard math Q is compact in the pointwise topology and the empirical MLE converges to the KL projection of p_C onto Q.
    Theorem A.1 is proved under this regularity condition. The paper concedes that assuming a globally optimal estimator is an idealization and is likely unattainable for non-convex neural training (Limitations).
  • domain assumption The largest transformer LMs in scaling studies approximate p_C itself, not merely the KL projection q*.
    Needed to interpret the inverse-U in psychometric fitness as falsifying the corpus assumption; footnote 21 calls this 'an assumption, not a theorem.'
  • domain assumption The external scaling results (Oh & Schuler 2023; Oh et al. 2024, 2025; Lin & Schuler 2025) are correct and not attributable to data leakage or corpus mismatch.
    The paper explicitly says the argument relies on these findings being correct and acknowledges the case is not fully closed (§1.3 fn. 13 and Limitations).

pith-pipeline@v1.3.0-alltime-deepseek · 17969 in / 19233 out tokens · 191071 ms · 2026-08-01T07:02:43.429426+00:00 · methodology

0 comments
read the original abstract

Surprisal theory holds that the human processing difficulty of a linguistic unit in context is an affine function of its surprisal under some language model. I argue this claim is a tautology without further constraint: for any non-negative difficulty measure over units in context, there exists a language model whose surprisal is an affine function of it under mild technical conditions. Therefore, because any pattern of difficulty is consistent with some language model, without an additional constraint on the language model, surprisal theory makes no falsifiable predictions. The tautology was long obscured by an assumption implicit in two decades of psycholinguistic work---that the relevant language model is the distribution that generated the training corpus, so that improving corpus fit improves predictions of human behavior. Recent empirical work has undermined this assumption, demonstrating that better corpus models can be worse predictors of processing difficulty. I conclude that breaking the tautology requires a rationalist intervention, i.e., the relevant language model must be derived from a non-empirically motivated model of the comprehender, which could be based on, for instance, memory constraints or processing goals, and that, thus, does not depend on the behavioral data surprisal theory is meant to explain.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

68 extracted references · 10 canonical work pages

  1. [1]

    Gerry T. M. Altmann and Yuki Kamide. 1999. https://doi.org/10.1016/S0010-0277(99)00059-1 Incremental interpretation at verbs: Restricting the domain of subsequent reference . Cognition, 73(3):247--264

  2. [2]

    Pascal Bergstr \"a er, Ryan Cotterell, and Anthony Widjaja Lin. 2026. https://openreview.net/forum?id=Yxz92UuPLQ Transformers are inherently succinct . In Proceedings of the International Conference on Learning Representations

  3. [3]

    Patrick Billingsley. 2012. https://www.wiley.com/en-us/Probability+and+Measure-p-9781118122372 Probability and Measure , anniversary edition. John Wiley & Sons

  4. [4]

    Blum and Ronald L

    Avrim L. Blum and Ronald L. Rivest. 1992. https://doi.org/10.1016/0893-6080(92)90010-G Training a 3-node neural network is NP -complete . Neural Networks, 5(1):117--127

  5. [5]

    Booth and Richard A

    Taylor L. Booth and Richard A. Thompson. 1973. https://doi.org/10.1109/T-C.1973.223746 Applying probability measures to abstract languages . IEEE Transactions on Computers, C-22(5):442--450

  6. [6]

    Marisa Ferrara Boston, John Hale, Reinhold Kliegl, Umesh Patil, and Shravan Vasishth. 2008. https://doi.org/10.16910/jemr.2.1.1 Parsing costs as predictors of reading difficulty: An evaluation using the P otsdam sentence corpus . Journal of Eye Movement Research, 2(1):1--12

  7. [7]

    Robert N. Brandon. 1978. https://www.sciencedirect.com/science/article/abs/pii/0039368178900055 Adaptation and evolutionary theory . Studies in History and Philosophy of Science, 9(3):181--206

  8. [8]

    Samuel Butler. 1882. https://archive.org/details/evolutionoldnew00butlgoog Evolution, Old and New , 2nd edition. Hardwicke and Bogue, London

  9. [9]

    Hubbard, and Cassandra L

    Andr \'e s Bux \'o -Lugo, Aniello De Santo , Morgan Grobol, Ryan J. Hubbard, and Cassandra L. Jacobs. 2026. https://arxiv.org/abs/2607.20208 surprisal is N ot a T heory . ArXiv:2607.20208

  10. [10]

    Hu, Jing Liu, Jaap Jumelet, Tal Linzen, Aaron Mueller, Candace Ross, Raj Sanjay Shah, Alex Warstadt, Ethan Gotlieb Wilcox, and Adina Williams

    Lucas Charpentier, Leshem Choshen, Ryan Cotterell, Mustafa Omer Gul, Michael Y. Hu, Jing Liu, Jaap Jumelet, Tal Linzen, Aaron Mueller, Candace Ross, Raj Sanjay Shah, Alex Warstadt, Ethan Gotlieb Wilcox, and Adina Williams. 2025. https://doi.org/10.18653/v1/2025.babylm-main.28 Findings of the third B aby LM challenge: Accelerating language modeling researc...

  11. [11]

    Noam Chomsky. 1965. https://archive.org/details/aspectsoftheoryo0000chom Aspects of the Theory of Syntax . MIT Press, Cambridge, MA

  12. [12]

    Noam Chomsky. 1966. https://doi.org/10.1017/CBO9780511803116 Cartesian Linguistics: A Chapter in the History of Rationalist Thought . Harper & Row, New York

  13. [13]

    Charles Darwin. 1859. https://archive.org/details/onoriginofspecie00darw On the Origin of Species by Means of Natural Selection . John Murray, London

  14. [14]

    Andrea De Varda and Marco Marelli. 2024. https://doi.org/10.18653/v1/2024.cmcl-1.3 Locally biased transformers better align with human reading times . In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pages 30--36

  15. [15]

    Gary S. Dell. 1986. https://doi.org/10.1037/0033-295X.93.3.283 A spreading-activation theory of retrieval in sentence production . Psychological Review, 93(3):283--321

  16. [16]

    Vera Demberg and Frank Keller. 2008. https://doi.org/10.1016/j.cognition.2008.07.008 Data from eye-tracking corpora as evidence for theories of syntactic processing complexity . Cognition, 109(2):193--210

  17. [17]

    Li Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, and Ryan Cotterell. 2023. https://doi.org/10.18653/v1/2023.acl-long.543 A measure-theoretic characterization of tight language models . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, pages 9744--9770

  18. [18]

    Frank and Rens Bod

    Stefan L. Frank and Rens Bod. 2011. https://doi.org/10.1177/0956797611409589 Insensitivity of the human sentence-processing system to hierarchical structure . Psychological Science, 22(6):829--834

  19. [19]

    Frank, Leun J

    Stefan L. Frank, Leun J. Otten, Giulia Galli, and Gabriella Vigliocco. 2015. https://doi.org/10.1016/j.bandl.2014.10.006 The ERP response to the amount of information conveyed by words in sentences . Brain and Language, 140:1--11

  20. [20]

    Richard Futrell, Edward Gibson, and Roger P. Levy. 2020. https://doi.org/10.1111/cogs.12814 Lossy-context surprisal: An information-theoretic model of memory effects in sentence processing . Cognitive Science, 44(3):e12814

  21. [21]

    Mario Giulianelli, Andreas Opedal, and Ryan Cotterell. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.682 Generalized measures of anticipation and responsivity in online language processing . In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 11648--11669

  22. [22]

    Adam Goodkind and Klinton Bicknell. 2018. https://doi.org/10.18653/v1/W18-0102 Predictive power of word surprisal for reading times is a linear function of language model quality . In Proceedings of the 8th Workshop on Cognitive Modeling and Computational Linguistics, pages 10--18

  23. [23]

    Levy, and Edward Gibson

    Michael Hahn, Richard Futrell, Roger P. Levy, and Edward Gibson. 2022. https://doi.org/10.1073/pnas.2122602119 A resource-rational model of human processing of recursive linguistic structure . Proceedings of the National Academy of Sciences, 119(43):e2122602119

  24. [24]

    John Hale. 2001. https://aclanthology.org/N01-1021/ A probabilistic E arley parser as a psycholinguistic model . Proceedings of the Second Meeting of the North American Chapter of the Association for Computational Linguistics, pages 159--166

  25. [25]

    John Hale. 2003. https://doi.org/10.1023/A:1022492123056 The information conveyed by words in sentences . Journal of Psycholinguistic Research, 32(2):101--123

  26. [26]

    Sepp Hochreiter and J \"u rgen Schmidhuber. 1997. https://doi.org/10.1162/neco.1997.9.8.1735 Long short-term memory . Neural Computation, 9(8):1735--1780

  27. [27]

    Hu, Aaron Mueller, Candace Ross, Adina Williams, Tal Linzen, Chengxu Zhuang, Ryan Cotterell, Leshem Choshen, Alex Warstadt, and Ethan Gotlieb Wilcox

    Michael Y. Hu, Aaron Mueller, Candace Ross, Adina Williams, Tal Linzen, Chengxu Zhuang, Ryan Cotterell, Leshem Choshen, Alex Warstadt, and Ethan Gotlieb Wilcox. 2024. https://aclanthology.org/2024.conll-babylm.1/ Findings of the second B aby LM challenge: Sample-efficient pretraining on developmentally plausible corpora . In Proceedings of the 2nd BabyLM ...

  28. [28]

    Herbert Jaeger. 2001. https://www.ai.rug.nl/minds/uploads/EchoStatesTechRep.pdf The ``echo state'' approach to analysing and training recurrent neural networks . GMD Report, 148

  29. [29]

    Samuel Kiegeland, V \' e steinn Sn bjarnarson, Tim Vieira, and Ryan Cotterell. 2026. https://aclanthology.org/2026.acl-long.1485/ On the proper treatment of units in surprisal theory . In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 32202--32224

  30. [30]

    Samuel Kiegeland, Ethan Wilcox, Afra Amini, David Robert Reich, and Ryan Cotterell. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.526 Reverse-engineering the reader . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 9367--9389

  31. [31]

    Kuperberg and T

    Gina R. Kuperberg and T. Florian Jaeger. 2016. https://doi.org/10.1080/23273798.2015.1102299 What do we mean by prediction in language comprehension? Language, Cognition and Neuroscience, 31(1):32--59

  32. [32]

    Tatsuki Kuribayashi, Yohei Oseki, Souhaib Ben Taieb, Kentaro Inui, and Timothy Baldwin. 2025. https://arxiv.org/abs/2502.01615 Large language models are human-like internally . Transactions of the Association for Computational Linguistics, 13:1743--1766

  33. [33]

    Roger Levy. 2008. https://doi.org/10.1016/j.cognition.2007.05.006 Expectation-based syntactic comprehension . Cognition, 106(3):1126--1177

  34. [34]

    Jiaoda Li and Ryan Cotterell. 2025. https://arxiv.org/abs/2505.23623 Characterizing the expressivity of fixed-precision transformer language models . In Advances in Neural Information Processing Systems, volume 38

  35. [35]

    Yi-Chien Lin and William Schuler. 2025. https://arxiv.org/abs/2506.11338 Surprisal from larger transformer-based language models predicts fMRI data more poorly . arXiv preprint

  36. [36]

    Brian MacWhinney. 2000. https://doi.org/10.1162/coli.2000.26.4.657 The CHILDES Project: Tools for Analyzing Talk , 3rd edition. Lawrence Erlbaum Associates, Mahwah, NJ

  37. [37]

    Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz

    Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993. https://aclanthology.org/J93-2004/ Building a large annotated corpus of E nglish: The P enn T reebank . Computational Linguistics, 19(2):313--330

  38. [38]

    Kate McCurdy and Michael Hahn. 2024. https://doi.org/10.18653/v1/2024.conll-1.4 Lossy context surprisal predicts task-dependent patterns in relative clause processing . In Proceedings of the 28th Conference on Computational Natural Language Learning, pages 36--45

  39. [39]

    Clara Meister, Tiago Pimentel, Thomas Hikaru Clark, Ryan Cotterell, and Roger P. Levy. 2022. https://arxiv.org/abs/2203.17213 Analyzing wrap-up effects through an information-theoretic lens . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing

  40. [40]

    Mills and John H

    Susan K. Mills and John H. Beatty. 1979. https://doi.org/10.1086/288865 The propensity interpretation of fitness . Philosophy of Science, 46(2):263--286

  41. [41]

    Byung-Doh Oh and William Schuler. 2023. https://doi.org/10.1162/tacl_a_00548 Why does surprisal from larger transformer-based language models provide a poorer fit to human reading times? Transactions of the Association for Computational Linguistics, 11:336--350

  42. [42]

    Byung-Doh Oh, Shisen Yue, and William Schuler. 2024. https://doi.org/10.18653/v1/2024.eacl-long.162 Frequency explains the inverse correlation of large language models' size, training data amount, and surprisal's fit to reading times . In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics, pages 2644--2663

  43. [43]

    Byung-Doh Oh, Hongao Zhu, and William Schuler. 2025. https://arxiv.org/abs/2506.01172 The inverse scaling effect of pre-trained language model surprisal is not due to data leakage . In Findings of the Association for Computational Linguistics: ACL 2025

  44. [44]

    Andreas Opedal, Eleanor Chodroff, Ryan Cotterell, and Ethan Gotlieb Wilcox. 2024. https://aclanthology.org/2024.emnlp-main.179/ On the role of context in reading time prediction . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing

  45. [45]

    Robert Henry Peters. 1976. https://doi.org/10.1086/283045 Tautology in evolution and ecology . The American Naturalist, 110(971):1--12

  46. [46]

    Pickering and Simon Garrod

    Martin J. Pickering and Simon Garrod. 2004. https://doi.org/10.1017/S0140525X04000056 Toward a mechanistic psychology of dialogue . Behavioral and Brain Sciences, 27(2):169--190

  47. [47]

    Karl R. Popper. 1959. https://archive.org/details/logicofscientifi0000popp The Logic of Scientific Discovery . Hutchinson, London

  48. [48]

    Karl R. Popper. 1974. https://doi.org/10.1007/978-94-010-1863-0_5 Darwinism as a metaphysical research programme . In Paul Arthur Schilpp, editor, The Philosophy of Karl Popper, pages 133--143. Open Court, La Salle, IL

  49. [49]

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf Language models are unsupervised multitask learners . Technical report, OpenAI

  50. [50]

    Henry Scheff \'e . 1947. https://doi.org/10.1214/aoms/1177730390 A useful convergence theorem for probability distributions . The Annals of Mathematical Statistics, 18(3):434--438

  51. [51]

    Cory Shain, Clara Meister, Tiago Pimentel, Ryan Cotterell, and Roger P. Levy. 2024. https://doi.org/10.1073/pnas.2307876121 Large-scale evidence for logarithmic effects of word predictability on reading time . Proceedings of the National Academy of Sciences, 121(10):e2307876121

  52. [52]

    Sinnott, L

    Edmund W. Sinnott, L. C. Dunn, and Theodosius Dobzhansky. 1958. https://archive.org/details/principlesofgene0000sinn Principles of Genetics , 5th edition. McGraw-Hill, New York

  53. [53]

    Smith and Roger Levy

    Nathaniel J. Smith and Roger Levy. 2013. https://doi.org/10.1016/j.cognition.2013.02.013 The effect of word predictability on reading time is logarithmic . Cognition, 128(3):302--319

  54. [54]

    V \' e steinn Sn bjarnarson, Samuel Kiegeland, Tianyu Liu, Reda Boumasmoud, Ryan Cotterell, and Tim Vieira. 2026. https://arxiv.org/abs/2603.05193 Transducing language models . In Proceedings of the International Conference on Learning Representations

  55. [55]

    Tanenhaus, Michael J

    Michael K. Tanenhaus, Michael J. Spivey-Knowlton, Kathleen M. Eberhard, and Julie C. Sedivy. 1995. https://doi.org/10.1126/science.7777863 Integration of visual and linguistic information in spoken language comprehension . Science, 268(5217):1632--1634

  56. [56]

    William Timkey and Tal Linzen. 2023. https://arxiv.org/abs/2310.16142 A language model with limited memory capacity captures interference in human sentence processing . In Findings of the Association for Computational Linguistics: EMNLP 2023

  57. [57]

    Eleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re, Tianyang Xu, Mario Giulianelli, Karolina Sta \'n czak, and Ryan Cotterell. 2026. https://aclanthology.org/2026.acl-long.575/ Probing for reading times . In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12618--12642

  58. [58]

    van der Vaart

    Aad W. van der Vaart. 1998. https://doi.org/10.1017/CBO9780511802256 Asymptotic Statistics . Cambridge University Press

  59. [59]

    Gomez, ukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. https://papers.nips.cc/paper/7181-attention-is-all-you-need Attention is all you need . In Advances in Neural Information Processing Systems, volume 30

  60. [60]

    Conrad Hal Waddington. 1960. https://archive.org/details/evolutionafterd01taxs Evolutionary adaptation . In Sol Tax, editor, Evolution After Darwin, volume 1, pages 381--402. University of Chicago Press, Chicago

  61. [61]

    Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Gotlieb Wilcox, Chengxu Zhuang, Juan Ciro, Rafael Mosquera, Bhargavi Paranjabe, Adina Williams, Tal Linzen, and Ryan Cotterell. 2023. https://doi.org/10.18653/v1/2023.conll-babylm.1 Findings of the B aby LM challenge: Sample-efficient pretraining on developmentally plausible corpora . In Proceedings of t...

  62. [62]

    Halbert White. 1982. https://doi.org/10.2307/1912526 Maximum likelihood estimation of misspecified models . Econometrica, 50(1):1--25

  63. [63]

    Halbert White. 1989. https://doi.org/10.1162/neco.1989.1.4.425 Learning in artificial neural networks: A statistical perspective . Neural Computation, 1(4):425--464

  64. [64]

    Ethan Gotlieb Wilcox, Jon Gauthier, Jennifer Hu, Peng Qian, and Roger Levy. 2020. https://escholarship.org/uc/item/738338tm On the predictive power of neural language models for human real-time comprehension behavior . In Proceedings of the 42nd Annual Meeting of the Cognitive Science Society, pages 1707--1713

  65. [65]

    Ethan Gotlieb Wilcox, Tiago Pimentel, Clara Meister, and Ryan Cotterell. 2024. https://doi.org/10.1016/j.cognition.2024.105765 An information-theoretic analysis of targeted regressions during reading . Cognition, 249:105765

  66. [66]

    Ethan Gotlieb Wilcox, Tiago Pimentel, Clara Meister, Ryan Cotterell, and Roger P. Levy. 2023. https://doi.org/10.1162/tacl_a_00612 Testing the predictions of surprisal theory in 11 languages . Transactions of the Association for Computational Linguistics, 11:1451--1470

  67. [67]

    Andy Yang, Anej Svete, Jiaoda Li, Anthony Widjaja Lin, Jonathan Rawski, Ryan Cotterell, and David Chiang. 2026. https://arxiv.org/abs/2510.27118 Probability distributions computed by autoregressive transformers . In Proceedings of the International Conference on Learning Representations

  68. [68]

    Ryo Yoshida, Shinnosuke Isono, Taiga Someya, Yohei Oseki, and Tatsuki Kuribayashi. 2026. https://aclanthology.org/2026.acl-long.1694/ An existence proof for neural language models that can explain garden-path effects via surprisal . In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 36...