REVIEW 2 major objections 7 minor 91 references
Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events
T0 review · 2 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Language models do not reliably differentiate impossible events from merely improbable ones: in the most adversarial condition, every model tested assigns higher probability to the impossible sentence at or above chance.
desk verdict A well-executed adversarial evaluation showing LLMs at or below chance on possible-vs-impossible distinctions, but the English passive stimuli have a potential instrument-reading confound that a referee should pin down. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The experimental machinery is a minimal-pairs probability comparison: each task pairs a sentence describing a possible event with a sentence describing an impossible event, differing only by a single critical word, and a model is scored correct if it assigns the possible sentence a higher probability. Possibility is operationalized through animacy violations, typicality through human plausibility ratings, and semantic relatedness through latent semantic analysis, with English and Mandarin stimuli drawn from prior psycholinguistic studies. The critical measure is the proportion of pairs on which the model prefers the possible sentence, computed for 35 language models and across training checkpoints of the Pythia suite.
What would settle it
Run the same minimal-pair probability comparison on a large, independently normed set of sentences whose impossibility is verified by human raters to admit no literal reading, and check whether any model family scores above chance on the condition where the possible sentence is atypical and unrelated while the impossible sentence is related; if any does, the paper's universal claim would be refuted.
Extended reading notes
Core claim
The central discovery is that language models' apparent ability to tell possible from impossible events is conditional on the possible event being typical and on semantic relatedness aligning with possibility. In the most adversarial condition, comparing a possible-but-atypical sentence with an unrelated critical word against an impossible sentence with a related critical word, every model tested assigns higher probability to the impossible sentence half or more of the time. This holds across model families and in both English and Mandarin. A training-trajectory analysis using the Pythia suite shows that performance on this contrast never rises above chance over the course of training, indicating that simply scaling up models does not repair the weakness. The paper interprets these results as evidence that language models lean on typicality and contextual relatedness as prediction cues instead of on robust world knowledge.
Load-bearing premise
The load-bearing premise is that the sentences marked impossible, such as 'the cure was discovered by the stamp,' are genuinely impossible rather than merely figurative or the start of a plausible longer sentence; the entire accuracy metric depends on this labeling being valid for the models being tested.
Editorial extensions
If this is right
- When the possible event is typical and the impossible critical word is unrelated to context, language models perform well, with accuracies often above 90 percent.
- Making the possible event atypical, making the impossible word semantically related, or both, produces significant drops in accuracy in both English and Mandarin.
- Regression analyses show that the typicality and semantic relatedness of both the possible and the impossible critical word independently predict whether the model gets a pair right, even after controlling for word frequency.
- Larger models do not systematically outperform smaller ones on this task, and on the most adversarial comparison the Pythia models never exceed chance accuracy over the full course of training.
- These results imply that previously reported abilities of language models to distinguish possible from impossible events may be driven by sensitivity to typicality and semantic relatedness rather than by robust event understanding.
Reading between the lines
- If the paper's account is right, then countermeasures such as explicitly instructing a model to reason literally or training on data that decorrelates relatedness from possibility might restore some robustness, but scaling alone will not.
- The same logic may apply to other types of impossibility beyond animacy violations, such as physical or temporal contradictions, so the failure mode is probably broader than the tested stimuli.
- In practical deployments where an atypical but possible event must be distinguished from an impossible one, such as medical triage or planning, these results suggest that default probability estimates from language models can be actively misleading.
- A testable extension would be to measure whether adding a short literal-reading prompt or a reasoning chain changes the ordering of probabilities on the most adversarial pairs, which would separate a deficit in stored world knowledge from a deficit in default prediction.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the 'Sherlock Holmes Task': given minimal pairs that differ by one critical word, a language model should assign higher probability to a possible-but-atypical sentence than to an impossible sentence. Using English stimuli from Vega-Mendoza et al. (2021) and Mandarin stimuli from Chow and Phillips (2013), the authors evaluate 35 base language models across conditions that cross possibility, typicality, and contextual semantic relatedness. They find that typical-vs-impossible accuracy is high, but accuracy drops sharply when the possible event is atypical and the impossible event is contextually related; in the PAU versus IAR condition, every tested model scores at or below chance in both languages. A mixed-effects regression on English items shows significant effects of relatedness, typicality, and word frequency, and Pythia scaling curves show that performance on PAU versus IAR never rises above chance during training.
Significance. The empirical pattern is striking and is well supported by the accuracy tables and mixed-effects regressions. The paper's main contribution is to separate typicality from possibility and to identify a failure mode that persists across model families, languages, and scale; this is an important corrective to claims that language models have robust event understanding. The authors provide open data and code, use a standard evaluation harness, and strengthen the analysis with item-level regression and training-curve data. However, the central inference depends on the 'impossible' stimuli being genuinely impossible, and the passive-by instrumental reading is a potential confound that must be resolved before the headline claim can be accepted.
major comments (2)
- [Section 3 and Limitations; Table 1; Figure 2] The 'impossible' stimuli are all animacy-violating passive sentences of the form 'X was V-ed by NP.' In English, 'by NP' can introduce an instrument as well as an agent, so sentences such as 'The cure for the disease was discovered by the medication' admit a possible instrumental reading ('discovered by means of the medication'). The Limitations section explicitly considers figurative and sentence-continuation readings but does not mention instrument readings. The norming study by Vega-Mendoza et al. (2021) collected plausibility ratings, not possibility judgments under all available readings, so it does not rule out this confound. Because the headline result in Section 5.4 (all models at or below chance on PAU versus IAR) is measured against labels of impossibility, a nontrivial fraction of instrument-reading items would make the reported below-chance accuracy reflect correct interpretation of a possible event rather than a failure to distinguish impossible from improbable. The authors should address this by excluding items whose critical noun can take an instrumental reading, by norming the stimuli for possibility under any reading, or by using active-voice controls.
- [Section 5.4 and Appendix A, Tables 2-3] The paper states that 'all models tested' perform at or below chance on the PAU versus IAR comparison, and the tables indeed show every reported accuracy at or below 0.50. However, the paper does not report per-model confidence intervals or significance tests against chance, so the strength of the claim rests on the aggregate mixed-effects analysis and visual inspection. Given that the between-model variance in this condition is nontrivial (e.g., English PAU versus IAR ranges from 0.168 to 0.426), a per-model binomial test or a model-level confidence interval would make the 'all models' claim more precise and would also clarify whether the below-chance pattern is statistically distinguishable from chance for each model.
minor comments (7)
- [Title page] The affiliation block contains a typo: 'Deparmtent' should be 'Department'.
- [Section 1] The phrase 'Clothesis the obvious best answer here' appears to be missing a space or verb; it should read 'Clothes is the obvious best answer here.'
- [Section 7.2] The word 'datasests' should be 'datasets.'
- [Limitations] The word 'indetical' should be 'identical.'
- [References] The reference to Jones et al. (2022) contains the typo 'Distrubutional' and should be 'Distributional.'
- [Appendix A, Table 2] In the model column, 'ai-forever/mGPT0.865' lacks a space between the model name and the accuracy value.
- [Experiment 3 and Table 1] The paper uses 'typicality' as an umbrella term but in Experiment 3 the typicality predictor is operationalized as human plausibility ratings from Vega-Mendoza et al. (2021); this should be stated explicitly where the regression is introduced.
Circularity Check
No meaningful circularity: the central result is read directly from model probabilities on externally normed psycholinguistic stimuli, with no fitted parameter renamed as a prediction.
full rationale
The paper's derivation chain is self-contained against external evidence. The 'impossible' versus 'possible but atypical' labels come from human psycholinguistic stimuli (Vega-Mendoza et al., 2021; Chow and Phillips, 2013), and the outcome measure is the raw probability each pretrained model assigns to the two sentences in each minimal pair. No parameter is fitted to the target comparison, and no score is constructed so as to make the below-chance result true by definition; the reported accuracies are direct counts of which sentence receives higher model probability. The paper does cite prior work by the same authors (e.g., Michaelov and Bergen, 2022; Michaelov et al., 2024; Jones et al., 2022), but these citations are used only to motivate the semantic-relatedness hypothesis and to position the result against prior findings; the central claim does not reduce to those citations. The acknowledged concern that animacy-violating passives may admit figurative or instrumental readings is a validity/correctness threat about stimulus labeling, not a circularity in the paper's derivation: the labels are imported from external norming rather than defined in terms of model outputs. Therefore the appropriate finding is a low non-circularity score, with no specific circular step to exhibit.
Assumptions & free parameters
assumptions (4)
- domain assumption Sentences with animacy violations (e.g., 'the stamp discovered the cure') are impossible events.
- domain assumption The relatedness and typicality labels from Vega-Mendoza et al. (2021) and Chow and Phillips (2013) are valid operationalizations for language model evaluation.
- domain assumption Comparing full-sentence probabilities assigned by base language models is a valid way to measure event understanding.
- standard math Logistic mixed-effects models and likelihood ratio tests provide valid inference for item-level accuracy differences.
Cite this review
Pith. "Pith review of Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events." pith.science (2026). https://pith.science/paper/GQ6OYDSC
@misc{pith2026250606808,
author = {Pith},
title = {Pith review of: Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events},
year = {2026},
howpublished = {\url{https://pith.science/paper/GQ6OYDSC}},
note = {Machine review of arXiv:2506.06808}
}
read the original abstract
Can language models reliably predict that possible events are more likely than merely improbable ones? By teasing apart possibility, typicality, and contextual relatedness, we show that despite the results of previous work, language models' ability to do this is far from robust. In fact, under certain conditions, all models tested - including Llama 3, Gemma 2, and Mistral NeMo - perform at worse-than-chance level, assigning higher probabilities to impossible sentences such as 'the car was given a parking ticket by the brake' than to merely unlikely sentences such as 'the car was given a parking ticket by the explorer'.
Figures
Reference graph
Works this paper leans on
-
[1]
01.AI , Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Guoyin Wang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yanpeng Li, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zh...
-
[2]
Mostafa Abdou, Vinit Ravishankar, Maria Barrett, Yonatan Belinkov, Desmond Elliott, and Anders S gaard. 2020. https://doi.org/10.18653/v1/2020.acl-main.679 The Sensitivity of Language Models and Humans to Winograd Schema Perturbations . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 7590--7604, Online. A...
-
[3]
Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Leandro von Werra , and Thomas Wolf. 2024. SmolLM - blazingly fast and remarkably powerful
2024
-
[4]
Bender, Timnit Gebru, Angelina McMillan-Major , and Shmargaret Shmitchell
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major , and Shmargaret Shmitchell. 2021. https://doi.org/10.1145/3442188.3445922 On the Dangers of Stochastic Parrots : Can Language Models Be Too Big ? 🦜 . In Proceedings of the 2021 ACM Conference on Fairness , Accountability , and Transparency , FAccT '21, pages 610--623, New York, NY, USA. Association f...
arXiv 2021
-
[5]
Alina Beygelzimer, Yann Dauphin, Percy Liang, and Jennifer Wortman Vaughan. 2021. Introducing the NeurIPS 2021 Paper Checklist
2021
-
[6]
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O'Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, Usvsn Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal. 2023. Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling . In Proceedings of the 40th Intern...
2023
-
[7]
Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Jeffrey Hsu, Mimansa Jaiswal, Wilson Y. Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick,...
-
[8]
BigScience Workshop , Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili \'c , Daniel Hesslow, Roman Castagn \'e , Alexandra Sasha Luccioni, Fran c ois Yvon, Matthias Gall \'e , Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Beno \^i t Sagot, Niklas Muennighoff, Albert Villano...
Show all 91 references
-
[9]
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. https://doi.org/10.1609/aaai.v34i05.6239 PIQA : Reasoning about Physical Commonsense in Natural Language . Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):7432--7439
2020 doi
-
[10]
Broderick, Andrew J
Michael P. Broderick, Andrew J. Anderson, Giovanni M. Di Liberto, Michael J. Crosse, and Edmund C. Lalor. 2018. https://doi.org/10.1016/j.cub.2018.01.080 Electrophysiological Correlates of Semantic Dissimilarity Reflect the Comprehension of Natural , Narrative Speech . Current...
2018 doi
-
[11]
Marine Carpuat, Marie-Catherine de Marneffe , Ivan Vladimir Meza Ruiz, Jesse Dodge, Margot Mieskes, ARR Editors-in-Chief , and ACL Ethics Committee . 2024. The ARR Responsible NLP Research checklist. http://aclrollingreview.org/responsibleNLPresearch/
2024
-
[12]
Wing-Yee Chow and Colin Phillips. 2013. https://doi.org/10.1016/j.brainres.2013.02.016 No semantic illusions in the `` Semantic P600 '' phenomenon: ERP evidence from Mandarin Chinese . Brain Research, 1506:76--93
2013 doi
-
[13]
Chwilla, Herman H
Dorothee J. Chwilla, Herman H. J. Kolk, and Constance T. W. M. Vissers. 2007. https://doi.org/10.1016/j.brainres.2007.09.014 Immediate integration of novel meanings: N400 support for an embodied view of language comprehension . Brain Research, 1183:109--123
2007 doi
-
[14]
Ernest Davis and Gary Marcus. 2015. https://doi.org/10.1145/2701413 Commonsense reasoning and commonsense knowledge in artificial intelligence . Communications of the ACM, 58(9):92--103
2015 doi
-
[15]
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. 2019. https://doi.org/10.18653/v1/D19-1224 Show Your Work : Improved Reporting of Experimental Results . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and ...
2019 doi
-
[16]
Arthur Conan Doyle. 1890. The Sign Of Four . Spencer Blackett, London
- [17]
-
[18]
Mengnan Du, Varun Manjunatha, Rajiv Jain, Ruchi Deshpande, Franck Dernoncourt, Jiuxiang Gu, Tong Sun, and Xia Hu. 2021. https://doi.org/10.18653/v1/2021.naacl-main.71 Towards Interpreting and Mitigating Shortcut Learning Behavior of NLU models . In Proceedings of the 2021 Conf...
2021 doi
-
[19]
S. T. Dumais, G. W. Furnas, T. K. Landauer, S. Deerwester, and R. Harshman. 1988. https://doi.org/10.1145/57167.57214 Using latent semantic analysis to improve access to textual information . In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems - CHI '...
1988
-
[20]
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Sch \"u tze, and Yoav Goldberg. 2021. https://doi.org/10.1162/tacl_a_00410 Measuring and Improving Consistency in Pretrained Language Models . Transactions of the Association for Computati...
2021 doi
-
[21]
Allyson Ettinger, Naomi Feldman, Philip Resnik, and Colin Phillips. 2016. Modeling N400 amplitude using vector space models of word representation. In Proceedings of the 38th Annual Conference of the Cognitive Science Society , Philadelphia, USA
2016
-
[22]
Evelina Fedorenko, Idan Asher Blank, Matthew Siegelman, and Zachary Mineroff. 2020. https://doi.org/10.1016/j.cognition.2020.104348 Lack of selectivity for syntax relative to word meanings throughout the language network . Cognition, 203:104348
2020
-
[23]
Jessica Zosa Forde and Michela Paganini. 2019. The Scientific Method in the Science of Machine Learning . In Debugging Machine Learning Models Workshop at ICLR
2019
-
[24]
Frank and Roel M
Stefan L. Frank and Roel M. Willems. 2017. https://doi.org/10.1080/23273798.2017.1323109 Word predictability and semantic similarity show distinct patterns of brain activity during language comprehension . Language, Cognition and Neuroscience, 32(9):1192--1203
2017
-
[25]
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021. https://doi.org/10.5281/zenodo.53...
2021 doi
-
[26]
Wichmann
Robert Geirhos, J \"o rn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. 2020. https://doi.org/10.1038/s42256-020-00257-z Shortcut learning in deep neural networks . Nature Machine Intelligence, 2(11):665--673
2020 doi
- [27]
- [28]
-
[29]
James J. Gibson. 1966. The Senses Considered as Perceptual Systems. Houghton Mifflin, Boston
1966
-
[30]
James J. Gibson. 1979. The Ecological Approach to Visual Perception. Houghton Mifflin, Boston
1979
-
[31]
Arthur M Glenberg and David A Robertson. 2000. https://doi.org/10.1006/jmla.2000.2714 Symbol Grounding and Meaning : A Comparison of High-Dimensional and Embodied Theories of Meaning . Journal of Memory and Language, 43(3):379--401
2000
-
[32]
H. P. Grice. 1989. Studies in the Way of Words. Harvard University Press, Cambridge, Mass
1989
-
[33]
Paul Grice
H. Paul Grice. 1975. Logic and conversation. In Peter Cole and Jerry Morgan, editors, Syntax and Semantics 3: Speech Acts , pages 41--58
1975
-
[34]
Dirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tusha...
2024
- [35]
-
[36]
Odd Erik Gundersen and Sigbj rn Kjensmo. 2018. https://doi.org/10.1609/aaai.v32i1.11503 State of the Art : Reproducibility in Artificial Intelligence . Proceedings of the AAAI Conference on Artificial Intelligence, 32(1)
2018 doi
-
[37]
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018. https://doi.org/10.18653/v1/N18-2017 Annotation Artifacts in Natural Language Inference Data . In Proceedings of the 2018 Conference of the North American Chapter of the Ass...
2018 doi
-
[38]
Michael Hanna, Yonatan Belinkov, and Sandro Pezzelle. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.744 When Language Models Fall in Love : Animacy Processing in Transformer Language Models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Pr...
2023 doi
-
[39]
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. 2018. https://doi.org/10.1609/aaai.v32i1.11694 Deep Reinforcement Learning That Matters . Proceedings of the AAAI Conference on Artificial Intelligence, 32(1)
2018 doi
-
[40]
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon ...
2022
-
[41]
Jennifer Hu and Michael Frank. 2024. Auxiliary task demands mask the capabilities of smaller language models. In First Conference on Language Modeling
2024
- [42]
-
[43]
Jones, Tyler A
Cameron R. Jones, Tyler A. Chang, Seana Coulson, James A. Michaelov, Sean Trott, and Benjamin K. Bergen. 2022. Distrubutional Semantics Still Can 't Account for Affordances . Proceedings of the Annual Meeting of the Cognitive Science Society, 44(44)
2022
-
[44]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. https://doi.org/10.48550/arXiv.2001.08361 Scaling Laws for Neural Language Models . Preprint, arXiv:2001.08361
-
[45]
Cantrell, Kenny Peng, Thanh Hien Pham, Christopher A
Sayash Kapoor, Emily M. Cantrell, Kenny Peng, Thanh Hien Pham, Christopher A. Bail, Odd Erik Gundersen, Jake M. Hofman, Jessica Hullman, Michael A. Lones, Momin M. Malik, Priyanka Nanayakkara, Russell A. Poldrack, Inioluwa Deborah Raji, Michael Roberts, Matthew J. Salganik, Ma...
2024 doi
-
[46]
Ivanova, Giulia Rambelli, Emmanuele Chersoni, Jingyuan Selena She, Zawad Chowdhury, Evelina Fedorenko, and Alessandro Lenci
Carina Kauf, Anna A. Ivanova, Giulia Rambelli, Emmanuele Chersoni, Jingyuan Selena She, Zawad Chowdhury, Evelina Fedorenko, and Alessandro Lenci. 2023. https://doi.org/10.1111/cogs.13386 Event Knowledge in Large Language Models : The Gap Between the Impossible and the Unlikely...
2023 doi
-
[47]
Pride Kavumba, Benjamin Heinzerling, Ana Brassard, and Kentaro Inui. 2021. https://doi.org/10.18653/v1/2021.naacl-main.304 Learning to Learn to be Right for the Right Reasons . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computati...
2021 doi
-
[48]
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei , and Juan Carlos Niebles. 2017. https://doi.org/10.1109/ICCV.2017.83 Dense- Captioning Events in Videos . In 2017 IEEE International Conference on Computer Vision ( ICCV ) , pages 706--715
2017 doi
-
[49]
Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012. The Winograd Schema Challenge . In Thirteenth International Conference on the Principles of Knowledge Representation and Reasoning
2012
-
[50]
Valentin Li \'e vin, Christoffer Egeberg Hother, Andreas Geert Motzfeldt, and Ole Winther. 2024. https://doi.org/10.1016/j.patter.2024.100943 Can large language models reason about medical questions? Patterns, 5(3)
2024
-
[51]
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mon...
2022 doi
-
[52]
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016. https://doi.org/10.1162/tacl_a_00115 Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies . Transactions of the Association for Computational Linguistics, 4:521--535
2016 doi
-
[53]
Llama Team . 2024. The Llama 3 Herd of Models
2024
-
[54]
Benjamin Marie, Atsushi Fujita, and Raphael Rubino. 2021. https://doi.org/10.18653/v1/2021.acl-long.566 Scientific Credibility of Machine Translation Research : A Meta-Evaluation of 769 Papers . In Proceedings of the 59th Annual Meeting of the Association for Computational Lin...
2021 doi
-
[55]
Rebecca Marvin and Tal Linzen. 2018. https://doi.org/10.18653/v1/D18-1151 Targeted Syntactic Evaluation of Language Models . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages 1192--1202, Brussels, Belgium. Association for Computa...
2018 doi
-
[56]
Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D
R. Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D. Hardy, and Thomas L. Griffiths. 2024. https://doi.org/10.1073/pnas.2322420121 Embers of autoregression show how large language models are shaped by the problem they are trained to solve . Proceedings of the National Academy ...
2024 doi
-
[57]
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. https://doi.org/10.18653/v1/P19-1334 Right for the Wrong Reasons : Diagnosing Syntactic Heuristics in Natural Language Inference . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages...
2019 doi
-
[58]
Urbach, Mary Hare, Ken McRae, and Jeffrey L
Ross Metusalem, Marta Kutas, Thomas P. Urbach, Mary Hare, Ken McRae, and Jeffrey L. Elman. 2012. https://doi.org/10.1016/j.jml.2012.01.001 Generalized event knowledge activation during online sentence comprehension . Journal of Memory and Language, 66(4):545--567
2012 doi
-
[59]
Michaelov, Megan D
James A. Michaelov, Megan D. Bardolph, Cyma K. Van Petten, Benjamin K. Bergen, and Seana Coulson. 2024. https://doi.org/10.1162/nol_a_00105 Strong Prediction : Language Model Surprisal Explains Multiple N400 Effects . Neurobiology of Language, 5(1):107--135
2024 doi
-
[60]
Michaelov and Benjamin K
James A. Michaelov and Benjamin K. Bergen. 2022. Collateral facilitation in humans and language models. In Proceedings of the 26th Conference on Computational Natural Language Learning ( CoNLL ) , pages 13--26, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computat...
2022
-
[61]
Michaelov, Seana Coulson, and Benjamin Bergen
James A. Michaelov, Seana Coulson, and Benjamin Bergen. 2023. Can Peanuts Fall in Love with Distributional Semantics ? In Proceedings of the Annual Meeting of the Cognitive Science Society , volume 45
2023
-
[62]
Kanishka Misra, Allyson Ettinger, and Julia Rayz. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.415 Exploring BERT 's Sensitivity to Lexical Cues using Tests from Semantic Priming . In Findings of the Association for Computational Linguistics : EMNLP 2020 , pages 4625-...
2020 doi
-
[63]
Nieuwland and Jos JA Van Berkum
Mante S. Nieuwland and Jos JA Van Berkum. 2006. When peanuts fall in love: N400 evidence for the power of discourse. Journal of cognitive neuroscience, 18(7):1098--1111
2006
-
[64]
Kuperberg
Martin Paczynski and Gina R. Kuperberg. 2012. https://doi.org/10.1016/j.jml.2012.07.003 Multiple influences of semantic memory on sentence processing: Distinct effects of semantic relatedness on violations of real-world event/state knowledge and animacy selection restrictions ...
2012 doi
-
[65]
Mehdi Parviz, Mark Johnson, Blake Johnson, and Jon Brock. 2011. Using Language Models and Latent Semantic Analysis to Characterise the N400m Neural Response . In Proceedings of the Australasian Language Technology Association Workshop 2011 , pages 38--46, Canberra, Australia
2011
- [66]
- [67]
-
[68]
Elizabeth Kumar, Aaron Horowitz, and Andrew Selbst
Inioluwa Deborah Raji, I. Elizabeth Kumar, Aaron Horowitz, and Andrew Selbst. 2022. https://doi.org/10.1145/3531146.3533158 The Fallacy of AI Functionality . In Proceedings of the 2022 ACM Conference on Fairness , Accountability , and Transparency , FAccT '22, pages 959--972, ...
2022
-
[69]
Nils Reimers and Iryna Gurevych. 2017. https://doi.org/10.18653/v1/D17-1035 Reporting Score Distributions Makes a Difference : Performance Study of LSTM-networks for Sequence Tagging . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , ...
2017 doi
-
[70]
Anna Rogers, Timothy Baldwin, and Kobi Leins. 2021. https://doi.org/10.18653/v1/2021.findings-emnlp.414 ' Just What do You Think You 're Doing , Dave ?' A Checklist for Responsible Data Use in NLP . In Findings of the Association for Computational Linguistics : EMNLP 2021 , pa...
2021 doi
-
[71]
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. https://doi.org/10.1609/aaai.v34i05.6399 WinoGrande : An Adversarial Winograd Schema Challenge at Scale . Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):8732--8740
2020 doi
-
[72]
Patrick Schramowski, Wolfgang Stammer, Stefano Teso, Anna Brugger, Franziska Herbert, Xiaoting Shao, Hans-Georg Luigs, Anne-Katrin Mahlein, and Kristian Kersting. 2020. https://doi.org/10.1038/s42256-020-0212-3 Making deep neural networks right for the right scientific reasons...
2020 doi
-
[73]
Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, and Praneeth Netrapalli. 2020. The Pitfalls of Simplicity Bias in Neural Networks . In Advances in Neural Information Processing Systems , volume 33, pages 9573--9585. Curran Associates, Inc
2020
-
[74]
Oleh Shliazhko, Alena Fenogenova, Maria Tikhonova, Anastasia Kozlova, Vladislav Mikhailov, and Tatiana Shavrina. 2024. https://doi.org/10.1162/tacl_a_00633 mGPT : Few-Shot Learners Go Multilingual . Transactions of the Association for Computational Linguistics, 12:58--79
2024 doi
-
[75]
Pfohl, Heather Cole-Lewis , Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R. Pfohl, Heather Cole-Lewis , Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H. Chen, Nigam H. Shah, Sami Lachgar, Philip Andrew ...
2025
-
[76]
Dan Sperber. 1986. Relevance: Communication and Cognition. Language and Thought Series. Harvard University Press, Cambridge, Massachusetts
1986
-
[77]
Michal Stefanik. 2022. https://doi.org/10.18653/v1/2022.naacl-srw.6 Methods for Estimating and Improving Robustness of Language Models . In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics : Human Language Techno...
2022 doi
- [78]
-
[79]
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. https://doi.org/10.18653/v1/N19-1421 CommonsenseQA : A Question Answering Challenge Targeting Commonsense Knowledge . In Proceedings of the 2019 Conference of the North American Chapter of the Associatio...
2019 doi
-
[80]
Takahisa Uchida, Nicolas Lair, Hiroshi Ishiguro, and Peter Ford Dominey. 2021. https://doi.org/10.1162/nol_a_00026 A Model of Online Temporal-Spatial Integration for Immediacy and Overrule in Discourse Comprehension . Neurobiology of Language, 2(1):83--105
2021 doi
-
[81]
Urbach and Marta Kutas
Thomas P. Urbach and Marta Kutas. 2010. https://doi.org/10.1016/j.jml.2010.03.008 Quantifiers more or less quantify on-line: ERP evidence for partial incremental interpretation . Journal of Memory and Language, 63(2):158--179
2010 doi
-
[82]
Langlotz, Jason Hom, Sergios Gatidis, John Pauly, and Akshay S
Dave Van Veen, Cara Van Uden, Louis Blankemeier, Jean-Benoit Delbrouck, Asad Aali, Christian Bluethgen, Anuj Pareek, Malgorzata Polacin, Eduardo Pontes Reis, Anna Seehofnerov \'a , Nidhi Rohatgi, Poonam Hosamani, William Collins, Neera Ahuja, Curtis P. Langlotz, Jason Hom, Ser...
2024 doi
-
[83]
Pickering, and Mante S
Mariana Vega-Mendoza , Martin J. Pickering, and Mante S. Nieuwland. 2021. https://doi.org/10.1016/j.neuropsychologia.2020.107724 Concurrent use of animacy and event-knowledge during comprehension: Evidence from event-related potentials . Neuropsychologia, 152:107724
2021
-
[84]
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020. https://doi.org/10.1162/tacl_a_00321 BLiMP : The Benchmark of Linguistic Minimal Pairs for English . Transactions of the Association for Computational Linguistics, ...
2020 doi
-
[85]
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019. https://doi.org/10.1162/tacl_a_00290 Neural Network Acceptability Judgments . Transactions of the Association for Computational Linguistics, 7:625--641
2019 doi
-
[86]
Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. Emergent Abilities of Large Language Mod...
2022
- [87]
-
[88]
Keren Ye and Adriana Kovashka. 2021. https://doi.org/10.1609/aaai.v35i4.16428 A Case Study of the Shortcut Effects in Visual Commonsense Reasoning . Proceedings of the AAAI Conference on Artificial Intelligence, 35(4):3181--3189
2021 doi
-
[89]
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018. https://doi.org/10.18653/v1/D18-1009 SWAG : A Large-Scale Adversarial Dataset for Grounded Commonsense Inference . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages...
2018 doi
-
[90]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. https://doi.org/10.18653/v1/P19-1472 HellaSwag : Can a Machine Really Finish Your Sentence ? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791--4...
2019 doi
-
[91]
Hongming Zhang, Xinran Zhao, and Yangqiu Song. 2020. https://doi.org/10.18653/v1/2020.acl-main.508 WinoWhy : A Deep Diagnosis of Essential Commonsense Knowledge for Answering Winograd Schema Challenge . In Proceedings of the 58th Annual Meeting of the Association for Computati...
2020 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.