Pith. sign in

REVIEW 2 major objections 7 minor 91 references

Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events

T0 review · 2 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Language models do not reliably differentiate impossible events from merely improbable ones: in the most adversarial condition, every model tested assigns higher probability to the impossible sentence at or above chance.

desk verdict A well-executed adversarial evaluation showing LLMs at or below chance on possible-vs-impossible distinctions, but the English passive stimuli have a potential instrument-reading confound that a referee should pin down. read the letter →

arxiv 2506.06808 v1 pith:GQ6OYDSC submitted 2025-06-07 cs.CL cs.AI

classification cs.CLcs.AI
keywords languagemodelseventpossibilitytypicalitysemanticrelatednessminimalpairsanimacyviolationsworldknowledgescaling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether language models can reliably assign higher probability to sentences describing possible events than to sentences describing impossible ones. Using minimal pairs that vary whether the critical word is typical or atypical and semantically related or unrelated to its context, the authors find that accuracy is high when the possible sentence is typical and the impossible word is unrelated, but it collapses when the possible event is atypical and the impossible word is related, with all 35 models tested performing at or below chance. The failure appears in both English and Mandarin and does not disappear with scale; in fact, smaller models often outperform larger ones. The authors conclude that previous evidence of model world knowledge may reflect sensitivity to typicality and contextual relatedness rather than a robust ability to distinguish the impossible from the merely unlikely.

What carries the argument

The experimental machinery is a minimal-pairs probability comparison: each task pairs a sentence describing a possible event with a sentence describing an impossible event, differing only by a single critical word, and a model is scored correct if it assigns the possible sentence a higher probability. Possibility is operationalized through animacy violations, typicality through human plausibility ratings, and semantic relatedness through latent semantic analysis, with English and Mandarin stimuli drawn from prior psycholinguistic studies. The critical measure is the proportion of pairs on which the model prefers the possible sentence, computed for 35 language models and across training checkpoints of the Pythia suite.

What would settle it

Run the same minimal-pair probability comparison on a large, independently normed set of sentences whose impossibility is verified by human raters to admit no literal reading, and check whether any model family scores above chance on the condition where the possible sentence is atypical and unrelated while the impossible sentence is related; if any does, the paper's universal claim would be refuted.

Watch

Extended reading notes

Core claim

The central discovery is that language models' apparent ability to tell possible from impossible events is conditional on the possible event being typical and on semantic relatedness aligning with possibility. In the most adversarial condition, comparing a possible-but-atypical sentence with an unrelated critical word against an impossible sentence with a related critical word, every model tested assigns higher probability to the impossible sentence half or more of the time. This holds across model families and in both English and Mandarin. A training-trajectory analysis using the Pythia suite shows that performance on this contrast never rises above chance over the course of training, indicating that simply scaling up models does not repair the weakness. The paper interprets these results as evidence that language models lean on typicality and contextual relatedness as prediction cues instead of on robust world knowledge.

Load-bearing premise

The load-bearing premise is that the sentences marked impossible, such as 'the cure was discovered by the stamp,' are genuinely impossible rather than merely figurative or the start of a plausible longer sentence; the entire accuracy metric depends on this labeling being valid for the models being tested.

Editorial extensions

If this is right

  • When the possible event is typical and the impossible critical word is unrelated to context, language models perform well, with accuracies often above 90 percent.
  • Making the possible event atypical, making the impossible word semantically related, or both, produces significant drops in accuracy in both English and Mandarin.
  • Regression analyses show that the typicality and semantic relatedness of both the possible and the impossible critical word independently predict whether the model gets a pair right, even after controlling for word frequency.
  • Larger models do not systematically outperform smaller ones on this task, and on the most adversarial comparison the Pythia models never exceed chance accuracy over the full course of training.
  • These results imply that previously reported abilities of language models to distinguish possible from impossible events may be driven by sensitivity to typicality and semantic relatedness rather than by robust event understanding.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the paper's account is right, then countermeasures such as explicitly instructing a model to reason literally or training on data that decorrelates relatedness from possibility might restore some robustness, but scaling alone will not.
  • The same logic may apply to other types of impossibility beyond animacy violations, such as physical or temporal contradictions, so the failure mode is probably broader than the tested stimuli.
  • In practical deployments where an atypical but possible event must be distinguished from an impossible one, such as medical triage or planning, these results suggest that default probability estimates from language models can be actively misleading.
  • A testable extension would be to measure whether adding a short literal-reading prompt or a reasoning chain changes the ordering of probabilities on the most adversarial pairs, which would separate a deficit in stored world knowledge from a deficit in default prediction.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 7 minor

Summary. The paper introduces the 'Sherlock Holmes Task': given minimal pairs that differ by one critical word, a language model should assign higher probability to a possible-but-atypical sentence than to an impossible sentence. Using English stimuli from Vega-Mendoza et al. (2021) and Mandarin stimuli from Chow and Phillips (2013), the authors evaluate 35 base language models across conditions that cross possibility, typicality, and contextual semantic relatedness. They find that typical-vs-impossible accuracy is high, but accuracy drops sharply when the possible event is atypical and the impossible event is contextually related; in the PAU versus IAR condition, every tested model scores at or below chance in both languages. A mixed-effects regression on English items shows significant effects of relatedness, typicality, and word frequency, and Pythia scaling curves show that performance on PAU versus IAR never rises above chance during training.

Significance. The empirical pattern is striking and is well supported by the accuracy tables and mixed-effects regressions. The paper's main contribution is to separate typicality from possibility and to identify a failure mode that persists across model families, languages, and scale; this is an important corrective to claims that language models have robust event understanding. The authors provide open data and code, use a standard evaluation harness, and strengthen the analysis with item-level regression and training-curve data. However, the central inference depends on the 'impossible' stimuli being genuinely impossible, and the passive-by instrumental reading is a potential confound that must be resolved before the headline claim can be accepted.

major comments (2)
  1. [Section 3 and Limitations; Table 1; Figure 2] The 'impossible' stimuli are all animacy-violating passive sentences of the form 'X was V-ed by NP.' In English, 'by NP' can introduce an instrument as well as an agent, so sentences such as 'The cure for the disease was discovered by the medication' admit a possible instrumental reading ('discovered by means of the medication'). The Limitations section explicitly considers figurative and sentence-continuation readings but does not mention instrument readings. The norming study by Vega-Mendoza et al. (2021) collected plausibility ratings, not possibility judgments under all available readings, so it does not rule out this confound. Because the headline result in Section 5.4 (all models at or below chance on PAU versus IAR) is measured against labels of impossibility, a nontrivial fraction of instrument-reading items would make the reported below-chance accuracy reflect correct interpretation of a possible event rather than a failure to distinguish impossible from improbable. The authors should address this by excluding items whose critical noun can take an instrumental reading, by norming the stimuli for possibility under any reading, or by using active-voice controls.
  2. [Section 5.4 and Appendix A, Tables 2-3] The paper states that 'all models tested' perform at or below chance on the PAU versus IAR comparison, and the tables indeed show every reported accuracy at or below 0.50. However, the paper does not report per-model confidence intervals or significance tests against chance, so the strength of the claim rests on the aggregate mixed-effects analysis and visual inspection. Given that the between-model variance in this condition is nontrivial (e.g., English PAU versus IAR ranges from 0.168 to 0.426), a per-model binomial test or a model-level confidence interval would make the 'all models' claim more precise and would also clarify whether the below-chance pattern is statistically distinguishable from chance for each model.
minor comments (7)
  1. [Title page] The affiliation block contains a typo: 'Deparmtent' should be 'Department'.
  2. [Section 1] The phrase 'Clothesis the obvious best answer here' appears to be missing a space or verb; it should read 'Clothes is the obvious best answer here.'
  3. [Section 7.2] The word 'datasests' should be 'datasets.'
  4. [Limitations] The word 'indetical' should be 'identical.'
  5. [References] The reference to Jones et al. (2022) contains the typo 'Distrubutional' and should be 'Distributional.'
  6. [Appendix A, Table 2] In the model column, 'ai-forever/mGPT0.865' lacks a space between the model name and the accuracy value.
  7. [Experiment 3 and Table 1] The paper uses 'typicality' as an umbrella term but in Experiment 3 the typicality predictor is operationalized as human plausibility ratings from Vega-Mendoza et al. (2021); this should be stated explicitly where the regression is introduced.

Circularity Check

0 steps flagged · score 1.0 of 10

No meaningful circularity: the central result is read directly from model probabilities on externally normed psycholinguistic stimuli, with no fitted parameter renamed as a prediction.

full rationale

The paper's derivation chain is self-contained against external evidence. The 'impossible' versus 'possible but atypical' labels come from human psycholinguistic stimuli (Vega-Mendoza et al., 2021; Chow and Phillips, 2013), and the outcome measure is the raw probability each pretrained model assigns to the two sentences in each minimal pair. No parameter is fitted to the target comparison, and no score is constructed so as to make the below-chance result true by definition; the reported accuracies are direct counts of which sentence receives higher model probability. The paper does cite prior work by the same authors (e.g., Michaelov and Bergen, 2022; Michaelov et al., 2024; Jones et al., 2022), but these citations are used only to motivate the semantic-relatedness hypothesis and to position the result against prior findings; the central claim does not reduce to those citations. The acknowledged concern that animacy-violating passives may admit figurative or instrumental readings is a validity/correctness threat about stimulus labeling, not a circularity in the paper's derivation: the labels are imported from external norming rather than defined in terms of model outputs. Therefore the appropriate finding is a low non-circularity score, with no specific circular step to exhibit.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new physical entities, forces, or fitted constants. Its central claim rests on domain assumptions about what counts as impossible, how relatedness and typicality are measured, and how model probabilities should be compared.

assumptions (4)
  • domain assumption Sentences with animacy violations (e.g., 'the stamp discovered the cure') are impossible events.
    Section 2.3 and the Limitations state that impossibility is operationalized as animacy violations. The paper notes these can sometimes be read figuratively or as plausible sentence beginnings, and mitigates with a period and prior norming data.
  • domain assumption The relatedness and typicality labels from Vega-Mendoza et al. (2021) and Chow and Phillips (2013) are valid operationalizations for language model evaluation.
    English relatedness uses Latent Semantic Analysis values and typicality uses human plausibility ratings from the original studies; Mandarin stimuli inherit classifications from Chow and Phillips (2013). The paper does not independently norm the Mandarin stimuli.
  • domain assumption Comparing full-sentence probabilities assigned by base language models is a valid way to measure event understanding.
    Section 3 adopts the BLiMP-style minimal-pair approach, comparing which of two sentences receives higher probability. This assumes that probability ranking reflects the model's event knowledge rather than, for example, surface frequency artifacts.
  • standard math Logistic mixed-effects models and likelihood ratio tests provide valid inference for item-level accuracy differences.
    Section 4.3 and Experiment 3 rely on these statistical tools with maximal random effects structures that converge; the Mandarin comparison used a Wald test when the regression did not converge.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events." pith.science (2026). https://pith.science/paper/GQ6OYDSC

@misc{pith2026250606808,
  author       = {Pith},
  title        = {Pith review of: Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GQ6OYDSC}},
  note         = {Machine review of arXiv:2506.06808}
}
read the original abstract

Can language models reliably predict that possible events are more likely than merely improbable ones? By teasing apart possibility, typicality, and contextual relatedness, we show that despite the results of previous work, language models' ability to do this is far from robust. In fact, under certain conditions, all models tested - including Llama 3, Gemma 2, and Mistral NeMo - perform at worse-than-chance level, assigning higher probabilities to impossible sentences such as 'the car was given a parking ticket by the brake' than to merely unlikely sentences such as 'the car was given a parking ticket by the explorer'.

Figures

Figures reproduced from arXiv: 2506.06808 by the authors.

Figure 1
Figure 1. Language model scores for all tasks comparing sentences describing typical events (Possible-Typical [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Language model scores for all tasks comparing sentences describing possible but atypical (Possible [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Pythia language model scores at all English tasks (from Experiments 1–2) over the course of training. [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

91 extracted references · 26 canonical work pages

  1. [1]

    01.AI , Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Guoyin Wang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yanpeng Li, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zh...

  2. [2]

    Mostafa Abdou, Vinit Ravishankar, Maria Barrett, Yonatan Belinkov, Desmond Elliott, and Anders S gaard. 2020. https://doi.org/10.18653/v1/2020.acl-main.679 The Sensitivity of Language Models and Humans to Winograd Schema Perturbations . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 7590--7604, Online. A...

  3. [3]

    Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Leandro von Werra , and Thomas Wolf. 2024. SmolLM - blazingly fast and remarkably powerful

  4. [4]

    Bender, Timnit Gebru, Angelina McMillan-Major , and Shmargaret Shmitchell

    Emily M. Bender, Timnit Gebru, Angelina McMillan-Major , and Shmargaret Shmitchell. 2021. https://doi.org/10.1145/3442188.3445922 On the Dangers of Stochastic Parrots : Can Language Models Be Too Big ? 🦜 . In Proceedings of the 2021 ACM Conference on Fairness , Accountability , and Transparency , FAccT '21, pages 610--623, New York, NY, USA. Association f...

  5. [5]

    Alina Beygelzimer, Yann Dauphin, Percy Liang, and Jennifer Wortman Vaughan. 2021. Introducing the NeurIPS 2021 Paper Checklist

  6. [6]

    Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O'Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, Usvsn Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal. 2023. Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling . In Proceedings of the 40th Intern...

  7. [7]

    Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick, Jason Phang, Aviya Skowron, Samson Tan, Xiangru Tang, Kevin A

    Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Jeffrey Hsu, Mimansa Jaiswal, Wilson Y. Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick,...

  8. [8]

    BigScience Workshop , Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili \'c , Daniel Hesslow, Roman Castagn \'e , Alexandra Sasha Luccioni, Fran c ois Yvon, Matthias Gall \'e , Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Beno \^i t Sagot, Niklas Muennighoff, Albert Villano...

Show all 91 references
  1. [9]

    Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. https://doi.org/10.1609/aaai.v34i05.6239 PIQA : Reasoning about Physical Commonsense in Natural Language . Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):7432--7439

  2. [10]

    Broderick, Andrew J

    Michael P. Broderick, Andrew J. Anderson, Giovanni M. Di Liberto, Michael J. Crosse, and Edmund C. Lalor. 2018. https://doi.org/10.1016/j.cub.2018.01.080 Electrophysiological Correlates of Semantic Dissimilarity Reflect the Comprehension of Natural , Narrative Speech . Current...

  3. [11]

    Marine Carpuat, Marie-Catherine de Marneffe , Ivan Vladimir Meza Ruiz, Jesse Dodge, Margot Mieskes, ARR Editors-in-Chief , and ACL Ethics Committee . 2024. The ARR Responsible NLP Research checklist. http://aclrollingreview.org/responsibleNLPresearch/

  4. [12]

    Wing-Yee Chow and Colin Phillips. 2013. https://doi.org/10.1016/j.brainres.2013.02.016 No semantic illusions in the `` Semantic P600 '' phenomenon: ERP evidence from Mandarin Chinese . Brain Research, 1506:76--93

  5. [13]

    Chwilla, Herman H

    Dorothee J. Chwilla, Herman H. J. Kolk, and Constance T. W. M. Vissers. 2007. https://doi.org/10.1016/j.brainres.2007.09.014 Immediate integration of novel meanings: N400 support for an embodied view of language comprehension . Brain Research, 1183:109--123

  6. [14]

    Ernest Davis and Gary Marcus. 2015. https://doi.org/10.1145/2701413 Commonsense reasoning and commonsense knowledge in artificial intelligence . Communications of the ACM, 58(9):92--103

  7. [15]

    Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. 2019. https://doi.org/10.18653/v1/D19-1224 Show Your Work : Improved Reporting of Experimental Results . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and ...

  8. [16]

    Arthur Conan Doyle. 1890. The Sign Of Four . Spencer Blackett, London

  9. [17]

    Mengnan Du, Fengxiang He, Na Zou, Dacheng Tao, and Xia Hu. 2022. https://doi.org/10.48550/arXiv.2208.11857 Shortcut Learning of Large Language Models in Natural Language Understanding : A Survey . Preprint, arXiv:2208.11857

  10. [18]

    Mengnan Du, Varun Manjunatha, Rajiv Jain, Ruchi Deshpande, Franck Dernoncourt, Jiuxiang Gu, Tong Sun, and Xia Hu. 2021. https://doi.org/10.18653/v1/2021.naacl-main.71 Towards Interpreting and Mitigating Shortcut Learning Behavior of NLU models . In Proceedings of the 2021 Conf...

  11. [19]

    S. T. Dumais, G. W. Furnas, T. K. Landauer, S. Deerwester, and R. Harshman. 1988. https://doi.org/10.1145/57167.57214 Using latent semantic analysis to improve access to textual information . In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems - CHI '...

  12. [20]

    Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Sch \"u tze, and Yoav Goldberg. 2021. https://doi.org/10.1162/tacl_a_00410 Measuring and Improving Consistency in Pretrained Language Models . Transactions of the Association for Computati...

  13. [21]

    Allyson Ettinger, Naomi Feldman, Philip Resnik, and Colin Phillips. 2016. Modeling N400 amplitude using vector space models of word representation. In Proceedings of the 38th Annual Conference of the Cognitive Science Society , Philadelphia, USA

  14. [22]

    Evelina Fedorenko, Idan Asher Blank, Matthew Siegelman, and Zachary Mineroff. 2020. https://doi.org/10.1016/j.cognition.2020.104348 Lack of selectivity for syntax relative to word meanings throughout the language network . Cognition, 203:104348

  15. [23]

    Jessica Zosa Forde and Michela Paganini. 2019. The Scientific Method in the Science of Machine Learning . In Debugging Machine Learning Models Workshop at ICLR

  16. [24]

    Frank and Roel M

    Stefan L. Frank and Roel M. Willems. 2017. https://doi.org/10.1080/23273798.2017.1323109 Word predictability and semantic similarity show distinct patterns of brain activity during language comprehension . Language, Cognition and Neuroscience, 32(9):1192--1203

  17. [25]

    Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021. https://doi.org/10.5281/zenodo.53...

  18. [26]

    Wichmann

    Robert Geirhos, J \"o rn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. 2020. https://doi.org/10.1038/s42256-020-00257-z Shortcut learning in deep neural networks . Nature Machine Intelligence, 2(11):665--673

  19. [27]

    Gemma Team , Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivi \`e re, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, L \'e onard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Bo...

  20. [28]

    Gemma Team , Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, L \'e onard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ram \'e , Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charli...

  21. [29]

    James J. Gibson. 1966. The Senses Considered as Perceptual Systems. Houghton Mifflin, Boston

  22. [30]

    James J. Gibson. 1979. The Ecological Approach to Visual Perception. Houghton Mifflin, Boston

  23. [31]

    Arthur M Glenberg and David A Robertson. 2000. https://doi.org/10.1006/jmla.2000.2714 Symbol Grounding and Meaning : A Comparison of High-Dimensional and Embodied Theories of Meaning . Journal of Memory and Language, 43(3):379--401

  24. [32]

    H. P. Grice. 1989. Studies in the Way of Words. Harvard University Press, Cambridge, Mass

  25. [33]

    Paul Grice

    H. Paul Grice. 1975. Logic and conversation. In Peter Cole and Jerry Morgan, editors, Syntax and Semantics 3: Speech Acts , pages 41--58

  26. [34]

    Dirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tusha...

  27. [35]

    Odd Erik Gundersen, Kevin Coakley, Christine Kirkpatrick, and Yolanda Gil. 2023. https://doi.org/10.48550/arXiv.2204.07610 Sources of Irreproducibility in Machine Learning : A Review . Preprint, arXiv:2204.07610

  28. [36]

    Odd Erik Gundersen and Sigbj rn Kjensmo. 2018. https://doi.org/10.1609/aaai.v32i1.11503 State of the Art : Reproducibility in Artificial Intelligence . Proceedings of the AAAI Conference on Artificial Intelligence, 32(1)

  29. [37]

    Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018. https://doi.org/10.18653/v1/N18-2017 Annotation Artifacts in Natural Language Inference Data . In Proceedings of the 2018 Conference of the North American Chapter of the Ass...

  30. [38]

    Michael Hanna, Yonatan Belinkov, and Sandro Pezzelle. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.744 When Language Models Fall in Love : Animacy Processing in Transformer Language Models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Pr...

  31. [39]

    Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. 2018. https://doi.org/10.1609/aaai.v32i1.11694 Deep Reinforcement Learning That Matters . Proceedings of the AAAI Conference on Artificial Intelligence, 32(1)

  32. [40]

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon ...

  33. [41]

    Jennifer Hu and Michael Frank. 2024. Auxiliary task demands mask the capabilities of smaller language models. In First Conference on Language Modeling

  34. [42]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L \'e lio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Tho...

  35. [43]

    Jones, Tyler A

    Cameron R. Jones, Tyler A. Chang, Seana Coulson, James A. Michaelov, Sean Trott, and Benjamin K. Bergen. 2022. Distrubutional Semantics Still Can 't Account for Affordances . Proceedings of the Annual Meeting of the Cognitive Science Society, 44(44)

  36. [44]

    Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. https://doi.org/10.48550/arXiv.2001.08361 Scaling Laws for Neural Language Models . Preprint, arXiv:2001.08361

  37. [45]

    Cantrell, Kenny Peng, Thanh Hien Pham, Christopher A

    Sayash Kapoor, Emily M. Cantrell, Kenny Peng, Thanh Hien Pham, Christopher A. Bail, Odd Erik Gundersen, Jake M. Hofman, Jessica Hullman, Michael A. Lones, Momin M. Malik, Priyanka Nanayakkara, Russell A. Poldrack, Inioluwa Deborah Raji, Michael Roberts, Matthew J. Salganik, Ma...

  38. [46]

    Ivanova, Giulia Rambelli, Emmanuele Chersoni, Jingyuan Selena She, Zawad Chowdhury, Evelina Fedorenko, and Alessandro Lenci

    Carina Kauf, Anna A. Ivanova, Giulia Rambelli, Emmanuele Chersoni, Jingyuan Selena She, Zawad Chowdhury, Evelina Fedorenko, and Alessandro Lenci. 2023. https://doi.org/10.1111/cogs.13386 Event Knowledge in Large Language Models : The Gap Between the Impossible and the Unlikely...

  39. [47]

    Pride Kavumba, Benjamin Heinzerling, Ana Brassard, and Kentaro Inui. 2021. https://doi.org/10.18653/v1/2021.naacl-main.304 Learning to Learn to be Right for the Right Reasons . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computati...

  40. [48]

    Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei , and Juan Carlos Niebles. 2017. https://doi.org/10.1109/ICCV.2017.83 Dense- Captioning Events in Videos . In 2017 IEEE International Conference on Computer Vision ( ICCV ) , pages 706--715

  41. [49]

    Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012. The Winograd Schema Challenge . In Thirteenth International Conference on the Principles of Knowledge Representation and Reasoning

  42. [50]

    Valentin Li \'e vin, Christoffer Egeberg Hother, Andreas Geert Motzfeldt, and Ole Winther. 2024. https://doi.org/10.1016/j.patter.2024.100943 Can large language models reason about medical questions? Patterns, 5(3)

  43. [51]

    Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mon...

  44. [52]

    Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016. https://doi.org/10.1162/tacl_a_00115 Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies . Transactions of the Association for Computational Linguistics, 4:521--535

  45. [53]

    Llama Team . 2024. The Llama 3 Herd of Models

  46. [54]

    Benjamin Marie, Atsushi Fujita, and Raphael Rubino. 2021. https://doi.org/10.18653/v1/2021.acl-long.566 Scientific Credibility of Machine Translation Research : A Meta-Evaluation of 769 Papers . In Proceedings of the 59th Annual Meeting of the Association for Computational Lin...

  47. [55]

    Rebecca Marvin and Tal Linzen. 2018. https://doi.org/10.18653/v1/D18-1151 Targeted Syntactic Evaluation of Language Models . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages 1192--1202, Brussels, Belgium. Association for Computa...

  48. [56]

    Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D

    R. Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D. Hardy, and Thomas L. Griffiths. 2024. https://doi.org/10.1073/pnas.2322420121 Embers of autoregression show how large language models are shaped by the problem they are trained to solve . Proceedings of the National Academy ...

  49. [57]

    Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. https://doi.org/10.18653/v1/P19-1334 Right for the Wrong Reasons : Diagnosing Syntactic Heuristics in Natural Language Inference . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages...

  50. [58]

    Urbach, Mary Hare, Ken McRae, and Jeffrey L

    Ross Metusalem, Marta Kutas, Thomas P. Urbach, Mary Hare, Ken McRae, and Jeffrey L. Elman. 2012. https://doi.org/10.1016/j.jml.2012.01.001 Generalized event knowledge activation during online sentence comprehension . Journal of Memory and Language, 66(4):545--567

  51. [59]

    Michaelov, Megan D

    James A. Michaelov, Megan D. Bardolph, Cyma K. Van Petten, Benjamin K. Bergen, and Seana Coulson. 2024. https://doi.org/10.1162/nol_a_00105 Strong Prediction : Language Model Surprisal Explains Multiple N400 Effects . Neurobiology of Language, 5(1):107--135

  52. [60]

    Michaelov and Benjamin K

    James A. Michaelov and Benjamin K. Bergen. 2022. Collateral facilitation in humans and language models. In Proceedings of the 26th Conference on Computational Natural Language Learning ( CoNLL ) , pages 13--26, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computat...

  53. [61]

    Michaelov, Seana Coulson, and Benjamin Bergen

    James A. Michaelov, Seana Coulson, and Benjamin Bergen. 2023. Can Peanuts Fall in Love with Distributional Semantics ? In Proceedings of the Annual Meeting of the Cognitive Science Society , volume 45

  54. [62]

    Kanishka Misra, Allyson Ettinger, and Julia Rayz. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.415 Exploring BERT 's Sensitivity to Lexical Cues using Tests from Semantic Priming . In Findings of the Association for Computational Linguistics : EMNLP 2020 , pages 4625-...

  55. [63]

    Nieuwland and Jos JA Van Berkum

    Mante S. Nieuwland and Jos JA Van Berkum. 2006. When peanuts fall in love: N400 evidence for the power of discourse. Journal of cognitive neuroscience, 18(7):1098--1111

  56. [64]

    Kuperberg

    Martin Paczynski and Gina R. Kuperberg. 2012. https://doi.org/10.1016/j.jml.2012.07.003 Multiple influences of semantic memory on sentence processing: Distinct effects of semantic relatedness on violations of real-world event/state knowledge and animacy selection restrictions ...

  57. [65]

    Mehdi Parviz, Mark Johnson, Blake Johnson, and Jon Brock. 2011. Using Language Models and Latent Semantic Analysis to Characterise the N400m Neural Response . In Proceedings of the Australasian Language Technology Association Workshop 2011 , pages 38--46, Canberra, Australia

  58. [66]

    Qwen, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le ...

  59. [67]

    Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks,...

  60. [68]

    Elizabeth Kumar, Aaron Horowitz, and Andrew Selbst

    Inioluwa Deborah Raji, I. Elizabeth Kumar, Aaron Horowitz, and Andrew Selbst. 2022. https://doi.org/10.1145/3531146.3533158 The Fallacy of AI Functionality . In Proceedings of the 2022 ACM Conference on Fairness , Accountability , and Transparency , FAccT '22, pages 959--972, ...

  61. [69]

    Nils Reimers and Iryna Gurevych. 2017. https://doi.org/10.18653/v1/D17-1035 Reporting Score Distributions Makes a Difference : Performance Study of LSTM-networks for Sequence Tagging . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , ...

  62. [70]

    Anna Rogers, Timothy Baldwin, and Kobi Leins. 2021. https://doi.org/10.18653/v1/2021.findings-emnlp.414 ' Just What do You Think You 're Doing , Dave ?' A Checklist for Responsible Data Use in NLP . In Findings of the Association for Computational Linguistics : EMNLP 2021 , pa...

  63. [71]

    Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. https://doi.org/10.1609/aaai.v34i05.6399 WinoGrande : An Adversarial Winograd Schema Challenge at Scale . Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):8732--8740

  64. [72]

    Patrick Schramowski, Wolfgang Stammer, Stefano Teso, Anna Brugger, Franziska Herbert, Xiaoting Shao, Hans-Georg Luigs, Anne-Katrin Mahlein, and Kristian Kersting. 2020. https://doi.org/10.1038/s42256-020-0212-3 Making deep neural networks right for the right scientific reasons...

  65. [73]

    Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, and Praneeth Netrapalli. 2020. The Pitfalls of Simplicity Bias in Neural Networks . In Advances in Neural Information Processing Systems , volume 33, pages 9573--9585. Curran Associates, Inc

  66. [74]

    Oleh Shliazhko, Alena Fenogenova, Maria Tikhonova, Anastasia Kozlova, Vladislav Mikhailov, and Tatiana Shavrina. 2024. https://doi.org/10.1162/tacl_a_00633 mGPT : Few-Shot Learners Go Multilingual . Transactions of the Association for Computational Linguistics, 12:58--79

  67. [75]

    Pfohl, Heather Cole-Lewis , Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H

    Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R. Pfohl, Heather Cole-Lewis , Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H. Chen, Nigam H. Shah, Sami Lachgar, Philip Andrew ...

  68. [76]

    Dan Sperber. 1986. Relevance: Communication and Cognition. Language and Thought Series. Harvard University Press, Cambridge, Massachusetts

  69. [77]

    Michal Stefanik. 2022. https://doi.org/10.18653/v1/2022.naacl-srw.6 Methods for Estimating and Improving Robustness of Language Models . In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics : Human Language Techno...

  70. [78]

    Shane Storks, Qiaozi Gao, and Joyce Y. Chai. 2020. https://doi.org/10.48550/arXiv.1904.01172 Recent Advances in Natural Language Inference : A Survey of Benchmarks , Resources , and Approaches . Preprint, arXiv:1904.01172

  71. [79]

    Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. https://doi.org/10.18653/v1/N19-1421 CommonsenseQA : A Question Answering Challenge Targeting Commonsense Knowledge . In Proceedings of the 2019 Conference of the North American Chapter of the Associatio...

  72. [80]

    Takahisa Uchida, Nicolas Lair, Hiroshi Ishiguro, and Peter Ford Dominey. 2021. https://doi.org/10.1162/nol_a_00026 A Model of Online Temporal-Spatial Integration for Immediacy and Overrule in Discourse Comprehension . Neurobiology of Language, 2(1):83--105

  73. [81]

    Urbach and Marta Kutas

    Thomas P. Urbach and Marta Kutas. 2010. https://doi.org/10.1016/j.jml.2010.03.008 Quantifiers more or less quantify on-line: ERP evidence for partial incremental interpretation . Journal of Memory and Language, 63(2):158--179

  74. [82]

    Langlotz, Jason Hom, Sergios Gatidis, John Pauly, and Akshay S

    Dave Van Veen, Cara Van Uden, Louis Blankemeier, Jean-Benoit Delbrouck, Asad Aali, Christian Bluethgen, Anuj Pareek, Malgorzata Polacin, Eduardo Pontes Reis, Anna Seehofnerov \'a , Nidhi Rohatgi, Poonam Hosamani, William Collins, Neera Ahuja, Curtis P. Langlotz, Jason Hom, Ser...

  75. [83]

    Pickering, and Mante S

    Mariana Vega-Mendoza , Martin J. Pickering, and Mante S. Nieuwland. 2021. https://doi.org/10.1016/j.neuropsychologia.2020.107724 Concurrent use of animacy and event-knowledge during comprehension: Evidence from event-related potentials . Neuropsychologia, 152:107724

  76. [84]

    Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020. https://doi.org/10.1162/tacl_a_00321 BLiMP : The Benchmark of Linguistic Minimal Pairs for English . Transactions of the Association for Computational Linguistics, ...

  77. [85]

    Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019. https://doi.org/10.1162/tacl_a_00290 Neural Network Acceptability Judgments . Transactions of the Association for Computational Linguistics, 7:625--641

  78. [86]

    Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

    Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. Emergent Abilities of Large Language Mod...

  79. [87]

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze...

  80. [88]

    Keren Ye and Adriana Kovashka. 2021. https://doi.org/10.1609/aaai.v35i4.16428 A Case Study of the Shortcut Effects in Visual Commonsense Reasoning . Proceedings of the AAAI Conference on Artificial Intelligence, 35(4):3181--3189

  81. [89]

    Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018. https://doi.org/10.18653/v1/D18-1009 SWAG : A Large-Scale Adversarial Dataset for Grounded Commonsense Inference . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages...

  82. [90]

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. https://doi.org/10.18653/v1/P19-1472 HellaSwag : Can a Machine Really Finish Your Sentence ? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791--4...

  83. [91]

    Hongming Zhang, Xinran Zhao, and Yangqiu Song. 2020. https://doi.org/10.18653/v1/2020.acl-main.508 WinoWhy : A Deep Diagnosis of Essential Commonsense Knowledge for Answering Winograd Schema Challenge . In Proceedings of the 58th Annual Meeting of the Association for Computati...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.