Pith. sign in

REVIEW 2 major objections 5 minor 53 references

Logical forms complement probability in understanding language model (and human) performance

T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Logical form, not just input probability, predicts how well language models reason.

desk verdict Solid empirical study of logical form and LLM reasoning, but the headline necessity-rejection bias is confounded with lexical choices and needs a paraphrase control. read the letter →

arxiv 2502.09589 v2 pith:C7YBFUM7 submitted 2025-02-13 cs.CL cs.LO

classification cs.CLcs.LO
keywords logicalformmodallogiclargelanguagemodelsperplexitysyllogismaffirmationbiasnecessitymodalityhumanreasoning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that knowing the logical form of a reasoning question—its modality (plain, must, may) and its argument structure (disjunctive syllogism, modus ponens, modus tollens)—predicts how well language models answer it, even after accounting for the model's perplexity on the input. Using a controlled set of 24,000 yes/no syllogism questions in propositional and modal logic, the authors find that LLMs systematically answer "yes" under possibility (may) but "no" under necessity (must), and that this rejection bias explains why accuracy on necessity questions is low. Human participants show a similar preference for modus ponens but no such rejection of necessity, suggesting the bias is a model artifact rather than a general reasoning pattern. If correct, the claim matters because it moves the explanation of LLM reasoning performance beyond input probability to the structure of the logical form itself.

What carries the argument

The controlled dataset of 24 logical-form templates (three modalities × four argument forms × valid/invalid sequents) rendered as natural-language yes/no questions, together with a soft-accuracy metric based on the relative probabilities of Yes and No tokens and linear mixed-effects models that separate the contributions of modality, argument form, and perplexity.

What would settle it

Re-running the modality comparison with alternative English paraphrases of the negated modal premises—for instance, "It is not certain that P" instead of "It's uncertain whether P", and "P is not guaranteed" instead of "It's impossible that P"—and checking whether the necessity rejection bias persists; if the bias disappears or reverses, the effect is lexical rather than logical-form-driven.

Watch

Extended reading notes

Core claim

The central claim is that logical form is a complementary, statistically significant predictor of LLM logic-reasoning accuracy alongside input probability: modality and argument form explain variance in soft accuracy beyond perplexity, and in particular LLMs exhibit an affirmation bias under possibility but a rejection bias under necessity, the latter being absent in human behavior. The paper establishes this through a controlled dataset and mixed-effects regressions, and shows with nonce-word controls that perplexity alone cannot account for performance differences.

Load-bearing premise

The natural-language templates are assumed to isolate logical form as the cause of the performance differences, so if the wording differences, such as "uncertain whether" versus "impossible that," drive the observed modality effects instead of the modal logic itself, the central interpretation would no longer hold.

Editorial extensions

If this is right

  • Accuracy on modal syllogisms is systematically lower for "must" than "may" across all ten open-weight models tested, with the gap reaching roughly 0.3 in soft accuracy.
  • A rejection bias toward "No" under necessity appears consistently across models, extending and refining the previously reported yes-response bias in LLMs.
  • Modus ponens is the easiest argument form for both LLMs and humans, while modus tollens is the hardest among valid forms; humans and LLMs share this preference order even though the effect sizes differ.
  • Input perplexity alone is a weak predictor (correlation ≈ −0.09) of logic accuracy, so benchmarks that rely on probability-based difficulty estimates miss a structured component of model behavior.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct testable extension would be to paraphrase the modal negations (e.g., "It is not certain that P" instead of "It's uncertain whether P") to see whether the necessity rejection bias is a semantic effect or a lexical artifact of the template wording.
  • The paper's modality findings suggest that planning systems built on LLMs should treat "must" assertions as unreliable for verification steps, since models tend to reject necessary conclusions even when they follow from premises.
  • The human–LLM divergence on necessity raises the possibility that instruction-tuning or RLHF, not pretraining alone, induces the rejection bias; this could be tested by comparing base and chat-tuned versions of the same model family.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces a controlled dataset of 24,000 natural-language yes/no items generated from propositional and modal syllogistic forms, evaluates ten open-weight LLMs with a probability-based soft-accuracy metric, fits mixed-effects models with modality, argument-form, and input-perplexity predictors, and compares LLM behavior with human responses on a subset of the items. The central claims are that logical form, especially modality, significantly predicts LLM performance beyond input perplexity; that LLMs exhibit an affirmation bias under possibility but a rejection bias under necessity; and that some argument-form preferences align with human data while the necessity rejection bias does not.

Significance. If the central claim holds, the paper makes a valuable empirical contribution: it provides a controlled, reusable testbed for propositional and modal reasoning, introduces a probability-based evaluation protocol, and offers systematic likelihood-ratio evidence that template-level logical properties matter over and above perplexity. The human behavioral comparison is also useful. The factorial synthesis, the explicit mixed-effects models, and the open-sourcing promise are strengths that make the results checkable. However, as detailed below, the modality-specific headline claim is underdetermined because the modality factor is perfectly collinear with a fixed lexical alternation in the templates, so the central interpretation is not yet established.

major comments (2)
  1. [§3.2/Table A1; §4.2.3/Eq. (5)] The modality factor in Eq. (5) is perfectly collinear with a fixed lexical alternation in the templates. Every necessity item uses 'certain'/'uncertain whether' for the positive/negated modal claims, while every possibility item uses 'possible'/'impossible'. The estimated effect (must 0.25 vs may 0.54 in Figure 6) therefore cannot be attributed to modal force rather than to the surface choice of 'uncertain' or to the negative polarity of the 'un-'/'im-' prefixes. In addition, 'It's uncertain whether P' is not semantically equivalent to ¬2P; it conventionally conveys both ¬2P and ¬2¬P, adding an extra premise that is absent from the 3 condition. Because the necessity rejection bias is the central behavioral claim (abstract, §4.2.3, §6), a paraphrase control that crosses modal meaning with at least two lexical realizations (for example, 'must/not necessarily' and 'certain/not certain' alongside 'may/not possibly' and 'possible/not possible') is required before the effect can be assigned to logical form. The Appendix C.1 necessitation result is suggestive, but it does not control the premise wording used in the main dataset.
  2. [§4.2.1 and §6] The paper's broader claim that 'logical forms should be considered as important factors' is supported for the argument-form dimension, since within a modality the four argument forms share a common modal lexical frame. However, the conclusion in §6 that 'the underlying logic forms play an important role in determining the performance' overstates what is identified for the modality dimension: the likelihood-ratio tests on Modality in Eqs. (4) and (5) test a factor that is aliased with template wording. The abstract and introduction should be re-scoped either to argument form plus a paraphrase-robust modality effect, or the modality-specific claims should be reported as depending on surface lexical choices until the control experiment is run.
minor comments (5)
  1. [§3.4] The phrase 'natural langauge template' contains a typo; it should be 'natural language template'.
  2. [§4.2.1] The description of Eq. (4) says that 'individual probability, coupled with a constant term, is modeled as a random effect,' which is confusing because the random-effect structure is a per-LLM intercept and slope for Perplexity; please rephrase to describe the (1 + Perplexity | LLM) specification explicitly.
  3. [§5 and Appendix A.2] The human experiment reports 710 responses but not the number of participants or the exclusion criteria; these details should be reported for reproducibility.
  4. [Table 2] Pairwise hypothesis-testing p-values are reported without a multiple-comparison correction; since all values are below 0.001 the conclusion is unlikely to change, but the correction or its absence should be stated.
  5. [Abstract and §4.2.1] The abstract's 'probability of input' is operationalized as input perplexity in Eq. (4); the manuscript should clarify the relation between the two at first use, since perplexity is a monotone function of average negative log-likelihood.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the logical-form factors are experimental manipulations, not fitted inputs; the regression analyses are descriptive and the central claims do not reduce to their inputs.

full rationale

The paper's central claims are empirical rather than derivational. The dataset is generated from 24 logical-form templates (Section 3.3; Table A1), with ground-truth answers assigned by the validity of the sequent, independently of model behavior. The logical form is therefore an experimentally manipulated independent variable, not a parameter fitted to the outcome. The mixed-effects models in Eqs. (4) and (5) fit modality, argument form, and perplexity to measured soft accuracy and to the relative probability of answering Yes; the estimated marginal means in Figures 3 and 6 are descriptive summaries of those fitted effects, not predictions forced by construction. A fitted factor is not a renamed prediction when the factor is part of the experimental design that defines the data. There is no load-bearing self-citation: the only author self-citation (Shi et al., 2022) is used to illustrate mixed evidence on probability versus execution-based evaluation and plays no role in justifying the logical-form claim, and no uniqueness theorem or ansatz is imported from prior work by the same authors. The potential confound flagged by the skeptical reader—that the modality factor in Eq. (5) is perfectly collinear with the fixed lexical realizations 'certain/uncertain' versus 'possible/impossible' in Table A1—is a genuine correctness and identifiability concern for the Section 4.2.3 rejection-bias interpretation, since no paraphrase control is run. However, this is not circularity: the paper does not define modality in terms of the measured outcome, and the claim is empirically testable with alternative surface realizations. The paper's own stated limitations (synthetic language, English-only coverage, limited human sample) are acknowledged scope restrictions and do not indicate circular reasoning. Overall, the derivation chain is self-contained: all reported effects are measured behavioral quantities, and no claim reduces by definition or by self-citation to its own inputs.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

No new theoretical entities are introduced; the dataset is a new empirical artifact, not a postulated entity. The free parameters are regression coefficients estimated from the data, which the paper's claims of significance depend on.

free parameters (3)
  • Modality fixed effects (must, may relative to propositional) = estimated in Eq. (4); not reported numerically
    Estimated from 24,000 LLM responses; used to claim modality differences (e.g., 'must' lower than 'may').
  • Argument form fixed effects (disjunctive, modus ponens, modus tollens) = estimated in Eq. (4)
    Used to claim modus ponens higher than modus tollens, etc.
  • Perplexity slope = rho = -0.09
    Negative correlation between perplexity and soft accuracy across 24,000 items.
assumptions (5)
  • domain assumption The natural-language templates in Table A1 faithfully and unambiguously express the underlying logical forms.
    Section 3.2 heuristic rules; the confound between logical form and surface wording (e.g., 'uncertain' vs 'impossible') is not controlled.
  • domain assumption Soft accuracy computed from next-token probabilities of Yes/No is a valid measure of LLM logical reasoning competence.
    Section 4.1, based on Hu and Levy (2023); presupposes that probability estimates reflect competence rather than decoding artifacts.
  • domain assumption The human Prolific sample (710 responses) is representative and adequately powered to detect modality effects.
    Section 5 and Appendix A.2; authors acknowledge limited sample size, and the GLMM finds no significant modality effect for humans, which may be a power issue.
  • standard math The mixed-effects model specifications (Eqs. 4, 5, 6) correctly capture per-model and per-participant variance structure.
    Standard statistical modeling; assumes conditional independence and appropriate random effects.
  • standard math Modal logic background, including Kripke semantics and the equivalences in Eqs. (2) and (3).
    Section 3.1; standard results, not derived in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Logical forms complement probability in understanding language model (and human) performance." pith.science (2026). https://pith.science/paper/C7YBFUM7

@misc{pith2026250209589,
  author       = {Pith},
  title        = {Pith review of: Logical forms complement probability in understanding language model (and human) performance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/C7YBFUM7}},
  note         = {Machine review of arXiv:2502.09589}
}
read the original abstract

With the increasing interest in using large language models (LLMs) for planning in natural language, understanding their behaviors becomes an important research question. This work conducts a systematic investigation of LLMs' ability to perform logical reasoning in natural language. We introduce a controlled dataset of hypothetical and disjunctive syllogisms in propositional and modal logic and use it as the testbed for understanding LLM performance. Our results lead to novel insights in predicting LLM behaviors: in addition to the probability of input (Gonen et al., 2023; McCoy et al., 2024), logical forms should be considered as important factors. In addition, we show similarities and discrepancies between the logical reasoning performances of humans and LLMs by collecting and comparing behavioral data from both.

Figures

Figures reproduced from arXiv: 2502.09589 by the authors.

Figure 1
Figure 1. Illustration of the fact that perplexity does [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The data synthesis pipeline: for each variable in logic forms (§ [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Estimated marginal means of logical form [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Illustration of per-model random effects on [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Correlation between mean perplexity and mean confidence score on each logic sequent. Each point [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Estimated marginal means of the factors in [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Estimated marginal means of logical form [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

53 extracted references · 25 canonical work pages

  1. [1]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    01.AI . 2024. https://doi.org/10.48550/arXiv.2403.04652 Yi: Open Foundation Models by 01. AI

  4. [4]

    AI@Meta. 2024. http://arxiv.org/abs/2407.21783 The Llama 3 Herd of Models

  5. [5]

    Roberta Ballarin. 2023. Modern origins of modal logic. In Edward N. Zalta and Uri Nodelman, editors, The Stanford Encyclopedia of Philosophy , F all 2023 edition. Metaphysics Research Lab, Stanford University

  6. [6]

    Leslie, and Uta Frith

    Simon Baron-Cohen , Alan M. Leslie, and Uta Frith. 1985. https://doi.org/10.1016/0010-0277(85)90022-8 Does the autistic child have a ``theory of mind'' ? Cognition, 21(1):37--46

  7. [7]

    Douglas Bates, Martin M \"a chler, Ben Bolker, and Steve Walker. 2015. https://doi.org/10.18637/jss.v067.i01 Fitting Linear Mixed-Effects Models Using lme4 . Journal of Statistical Software, 67(1)

  8. [8]

    Belem, Markelle Kelly, Mark Steyvers, Sameer Singh, and Padhraic Smyth

    Catarina G. Belem, Markelle Kelly, Mark Steyvers, Sameer Singh, and Padhraic Smyth. 2024. https://doi.org/10.48550/arXiv.2407.15814 Perceptions of Linguistic Uncertainty by Language Models and Humans

Show all 53 references
  1. [9]

    Peter Clark, Oyvind Tafjord, and Kyle Richardson. 2021. Transformers as soft reasoners over language. In IJCAI

  2. [10]

    Vittoria Dentella, Fritz G \"u nther, and Evelina Leivada. 2023. Systematic testing of three language models reveals low language accuracy, absence of response stability, and a yes-response bias. Proceedings of the National Academy of Sciences, 120(51):e2309583120

  3. [11]

    Tiwalayo Eisape, Michael Tessler, Ishita Dasgupta, Fei Sha, Sjoerd Steenkiste, and Tal Linzen. 2024. A systematic comparison of syllogistic reasoning in humans and language models. In NAACL

  4. [12]

    Clémentine Fourrier, Nathan Habib, Alina Lozovskaya, Konrad Szafer, and Thomas Wolf. 2024. Open llm leaderboard v2. https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard

  5. [13]

    Hila Gonen, Srini Iyer, Terra Blevins, Noah Smith, and Luke Zettlemoyer. 2023. Demystifying prompts in language models via perplexity estimation. In Findings of ACL: EMNLP

  6. [14]

    Christopher Hahn, Frederik Schmitt, Jens U Kreber, Markus Norman Rabe, and Bernd Finkbeiner. 2021. Teaching temporal logics to neural networks. In ICLR

  7. [15]

    Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano , Hannah Szab \'o , Ekaterina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, ...

  8. [16]

    Holliday, Matthew Mandelkern, and Cedegao E

    Wesley H. Holliday, Matthew Mandelkern, and Cedegao E. Zhang. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.222 Conditional and Modal Reasoning in Large Language Models . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages 3800...

  9. [17]

    Jennifer Hu and Roger Levy. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.306 Prompting is not a substitute for probability measurements in large language models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages 5040--5060,...

  10. [18]

    Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. 2022. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In ICML, pages 9118--9147. PMLR

  11. [19]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L \'e lio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Tho...

  12. [20]

    Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, L \'e lio Renard Lavaud, Lucile Saulnier, Marie-...

  13. [21]

    Philip Nicholas Johnson-Laird. 1983. Mental models: Towards a cognitive science of language, inference, and consciousness. 6. Harvard University Press

  14. [22]

    Henry A Kautz, Bart Selman, et al. 1992. Planning as satisfiability. In ECAI, volume 92, pages 359--363. Citeseer

  15. [23]

    Saul A. Kripke. 1959. https://doi.org/10.2307/2964568 A Completeness Theorem in Modal Logic . The Journal of Symbolic Logic, 24(1):1--14

  16. [24]

    Saul A. Kripke. 1963. https://doi.org/10.1002/malq.19630090502 Semantical Analysis of Modal Logic I Normal Modal Propositional Calculi . Mathematical Logic Quarterly, 9(5-6):67--96

  17. [25]

    Andrew K Lampinen, Ishita Dasgupta, Stephanie C Y Chan, Hannah R Sheahan, Antonia Creswell, Dharshan Kumaran, James L McClelland, and Felix Hill. 2024. https://doi.org/10.1093/pnasnexus/pgae233 Language models, like humans, show content effects on reasoning tasks . PNAS Nexus,...

  18. [26]

    Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, R \'e mi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al. 2022. Competition-level code generation with alphacode. Science, 378(6624):1092--1097

  19. [27]

    Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2020. https://doi.org/10.24963/ijcai.2020/501 LogiQA : A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning . In Proceedings of the Twenty-Ninth International Joint Conference on...

  20. [28]

    Ziqiao Ma, Jacob Sansom, Run Peng, and Joyce Chai. 2023. Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models . Findings of Empirical Methods in Natural Language Processing

  21. [29]

    Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D

    R. Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D. Hardy, and Thomas L. Griffiths. 2024. https://doi.org/10.1073/pnas.2322420121 Embers of autoregression show how large language models are shaped by the problem they are trained to solve . Proceedings of the National Academy ...

  22. [30]

    Clara Meister, Tiago Pimentel, Gian Wiher, and Ryan Cotterell. 2023. Locally typical sampling. Transactions of the Association for Computational Linguistics, 11:102--121

  23. [31]

    Microsoft. 2023. https://www.microsoft.com/en-us/research/blog/phi-2-the-surprising-power-of-small-language-models/ Phi-2: The surprising power of small language models . Microsoft Research Blog

  24. [32]

    Microsoft. 2024. http://arxiv.org/abs/2404.14219 Phi-3 Technical Report : A Highly Capable Language Model Locally on Your Phone

  25. [33]

    Kanishka Misra and Najoung Kim. 2024. Generating novel experimental hypotheses from language models: A case study on cross-dative generalization. arXiv preprint arXiv:2408.05086

  26. [34]

    Santiago Ontanon, Joshua Ainslie, Vaclav Cvicek, and Zachary Fisher. 2022. LogicInference : A new Datasaet for Teaching Logical Inference to seq2seq Models . In ICLR2022 Workshop on the Elements of Reasoning : Objects , Structure and Causality

  27. [35]

    OpenAI. 2024. Learning to reason with llms. https://openai.com/index/learning-to-reason-with-llms/

  28. [36]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 3...

  29. [37]

    Mihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura, Man Luo, Santosh Mashetty, Arindam Mitra, and Chitta Baral. 2024. https://aclanthology.org/2024.acl-long.739 L ogic B ench: Towards systematic evaluation of logical reasoning ability of large language models . In P...

  30. [38]

    David Premack and Guy Woodruff. 1978. https://doi.org/10.1017/S0140525X00076512 Does the chimpanzee have a theory of mind? Behavioral and Brain Sciences, 1(4):515--526

  31. [39]

    Neil Rabinowitz, Frank Perbet, Francis Song, Chiyuan Zhang, S. M. Ali Eslami, and Matthew Botvinick. 2018. Machine Theory of Mind . In Proceedings of the 35th International Conference on Machine Learning , pages 4218--4227. PMLR

  32. [40]

    Marco Ragni, Hannah Dames, Daniel Brand, and Nicolas Riesterer. 2019. When Does a Reasoner Respond : Nothing Follows ?: 41st Annual Meeting of the Cognitive Science Society . Proceedings of the 41st Annual Conference of the Cognitive Science Society, pages 2640--2645

  33. [41]

    Stephen W Raudenbush. 2002. Hierarchical linear models: Applications and data analysis methods. Advanced Quantitative Techniques in the Social Sciences Series/SAGE

  34. [42]

    Baptiste Rozi \`e re, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, J \'e r \'e my Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, ...

  35. [43]

    Abulhair Saparov and He He. 2022. Language Models Are Greedy Reasoners : A Systematic Formal Analysis of Chain-of-Thought . In The Eleventh International Conference on Learning Representations

  36. [44]

    Abulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi, Mehran Kazemi, Najoung Kim, and He He. 2023. Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD Examples . Advances in Neural Information Processing Systems, 36:3083--3105

  37. [45]

    Freda Shi, Daniel Fried, Marjan Ghazvininejad, Luke Zettlemoyer, and Sida I Wang. 2022. Natural language to code translation with execution. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing

  38. [46]

    Stuart M. Shieber. 1993. https://aclanthology.org/J93-1008 The problem of logical form equivalence . Computational Linguistics, 19(1):179--190

  39. [47]

    Damien Sileo and Antoine Lernould. 2023. https://aclanthology.org/2023.findings-emnlp.303 M ind G ames: Targeting theory of mind in large language models with dynamic epistemic modal logic . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 4570--4577

  40. [48]

    Oyvind Tafjord, Bhavana Dalvi, and Peter Clark. 2022. Entailer: Answering questions with faithful and truthful chains of reasoning. In EMNLP

  41. [49]

    Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530

  42. [50]

    Hugo Touvron, Louis Martin, and Kevin Stone. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models

  43. [51]

    Yuxuan Wan, Wenxuan Wang, Yiliu Yang, Youliang Yuan, Jen-tse Huang, Pinjia He, Wenxiang Jiao, and Michael Lyu. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.128 LogicAsker : Evaluating and Improving the Logical Reasoning Ability of Large Language Models . In Proceedings of...

  44. [52]

    Yimei Xiang. 2019. Two types of higher-order readings of wh-questions. Proceedings of the 22nd Amsterdam Colloquium

  45. [53]

    Shi Zong and Jimmy Lin. 2024. Categorical syllogisms revisited: A review of the logical reasoning abilities of llms for analyzing categorical syllogism. arXiv preprint arXiv:2406.18762

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.