REVIEW 3 major objections 3 minor 207 references
This paper claims that SAGE—splitting pragmatic modeling into LM proposers and evaluators over a symbolic task analysis—lets cognitive models scale to open-ended alternatives, but the LMs are reliable as proposers, not as formal evaluators.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 15:22 UTC pith:NOIWQBM2
load-bearing objection A careful, honest framework paper whose headline asymmetry (good proposers, shaky formal evaluators) is real and well-documented—more roadmap than finished solution, and worth engaging despite unquantified prompt tuning. the 3 major comments →
Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
SAGE models can be constructed for pragmatic production and interpretation by replacing manual alternative-specification with LM generation, while keeping the reasoning steps transparent. The paper's detailed module evaluations show that the LM proposers produce natural, contextually appropriate alternatives—rated by humans as comparable to human-written ones—whereas LM evaluators, when asked to make formal judgments such as literal semantic truth, differential complexity, or Gricean maxim flouting, are prompt-sensitive and often diverge from human judgments. The central discovery is therefore not just that the framework works end-to-end, but that the appropriate division of labor between ne
What carries the argument
The SAGE framework (ScAffolded Generative models for Explanation), built from three module types: proposers, evaluators, and selectors. Proposers use LMs to sample an open-ended space of candidate alternatives; evaluators assess those alternatives on dimensions like literal truth, complexity, or prior plausibility; selectors apply rule-based operations from an explicit task analysis. The central contribution is treating alternative-generation as a flexible LM subroutine while keeping the cognitive reasoning steps explicit and testable. The case studies instantiate the machinery in three task analyses: the Incremental Algorithm for referential expression generation, a markedness-blocking proc
Load-bearing premise
The framework's viability rests on the assumption that a single zero-shot LM call can reliably perform the evaluator roles the task analysis requires—deciding literal truth, comparing expression complexity, and detecting Gricean maxim flouting.
What would settle it
A benchmark study in which the evaluator modules are tested on held-out items without prompt tuning: if the semantic evaluator's accuracy on entailment pairs falls to chance, or if replacing the LM evaluators with random or fixed heuristics in the end-to-end SAGE pipeline does not reduce accuracy below the human-fit level, the claim that LM evaluators carry the formal judgment load would be falsified. The paper's own appendix already reports the semantic evaluator at 0.82 accuracy and near-universal maxim-flouting flags, so a broader replication of these failures on new stimuli would settle th
If this is right
- Manual specification of alternatives can be replaced by LM generation for a range of pragmatic tasks, opening models to open-ended contexts.
- End-to-end accuracy of SAGE models is not sufficient evidence of component adequacy; module-level evaluation against human judgments is necessary.
- The asymmetry between proposers and evaluators suggests practical guidance: use LMs for sampling alternatives and intuitive judgments; supply formal judgments from specialized components such as fine-tuned models, probability scoring, or symbolic methods.
- SAGE provides a concrete implementation route for verbal Gricean theory, enabling quantitative comparison of different assumption sets against human data.
Where Pith is reading between the lines
- A likely testable extension is replacing prompt-based evaluators with LM-internal probability scoring, such as conditional log-probabilities; if that improves evaluator reliability, the asymmetry would narrow.
- The observed evaluator weakness for abstract judgments suggests that neuro-symbolic cognitive models should keep formal reasoning outside the LM or fine-tune dedicated evaluators rather than rely on zero-shot prompts.
- The no-assumption model's comparable fit to the Gricean model hints that maxim-flouting detection may not be the driving component of implicature interpretation; a testable prediction is that omitting assumption evaluation yields similar or better fits on new implicature datasets.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SAGE (ScAffolded Generative models for Explanation), a neuro-symbolic framework for cognitive modeling of pragmatics. SAGE decomposes a pragmatic task into LM-based proposers (which generate open-ended utterance/interpretation alternatives), LM-based evaluators (which assess semantics, complexity, typicality, or maxim violation), and rule-based selectors that implement the symbolic task analysis. The framework is evaluated in three case studies: referential expression generation in a reference game, M-implicature interpretation from periphrastic causatives, and Gricean conversational implicature interpretation. The models are assessed with accuracy, ablations, baselines, human module ratings, and quantitative fit to human forced-choice data. The headline finding is an asymmetry: LM proposers generate viable alternatives, while LM evaluators perform well on intuitive judgments but are less reliable for formal/theoretical assessments, such as literal truth, differential complexity, and maxim flouting.
Significance. If the findings hold, the paper makes a useful methodological contribution: it demonstrates a concrete way to combine LMs with transparent symbolic task analyses for pragmatics, and it provides a systematic, honest assessment of which LM subtasks are currently viable. The paper is particularly strong in its evaluative practice: it uses freshly collected human data, externally grounded human module ratings, Bayesian model comparison, and it openly reports component failures and selection effects. The claimed proposer/evaluator asymmetry is a valuable empirical insight for the growing literature on neuro-symbolic cognitive models. The main risk is that the framework's promise of open-endedness depends on evaluator reliability, which the paper itself shows to be limited; the paper should therefore be read as a proof-of-concept with clear bottlenecks, not as a demonstration that all SAGE components already work.
major comments (3)
- [§3.3 / Appendix B.4] The markedness-blocking model discards runs in which no unblocked state-utterance pair is available. This introduces a selection effect on the reported M-implicature accuracy: the model is scored only on cases where the blocking mechanism successfully produces at least one unblocked interpretation. Since blocking is exactly the mechanism the model is intended to explain, the paper should report (i) the proportion of discarded runs, (ii) whether the discarded runs are systematically different (e.g., particular vignettes or utterance types), and (iii) accuracy when discarded runs are counted as failures. Without this, the 80% M-implicature accuracy is conditional on a favorable outcome of the proposed algorithm rather than an unconditional model prediction.
- [§4.2.1 / Appendix C.1.1] The assumption-evaluation module flags almost all maxims as violated, while humans are much more reluctant; in the no-assumptions ablation the model still achieves accuracy and human-data fit close to the full Gricean model. The paper acknowledges these facts, but their implications for the central claim are not sufficiently resolved. The Gricean AE model is credibly better than the no-assumptions model, but the differences are small, and the module-level evaluation shows that the assumption evaluator's output is not human-like. The authors should either provide a more direct attribution analysis (e.g., comparing models where the assumption evaluator is replaced by human violation judgments, or where the plausibility evaluator is ablated) or explicitly restrict the scope of the conclusion to the proposer-based component of SAGE.
- [§2.1 / Appendix A.1.2] The SemanticEvaluator prompt was optimized during development and the final module obtains only 0.82 accuracy on NLI-style and matched test sets. More importantly, the evaluation of the iterative model's final contrastivity is based on manual annotation by the authors, not on the module's own outputs. The paper notes that the model sometimes failed to recognize human-fully-contrastive utterances and iterated further. This makes it difficult to know how much of the IM's success is due to the semantic evaluator as opposed to the proposer and the manual evaluation procedure. Please report end-to-end contrastivity using the raw SemanticEvaluator outputs, and quantify how often the evaluator's errors changed the number of iterations or the selected utterance.
minor comments (3)
- [General] There are several typos and misspellings: 'pehnomena' (§2.4), 'interprepretation' (Appendix C heading), 'ConstrastivitySelector' (Algorithm 1), 'GTP-3.5-turbo' (§4.3, C.1.1), 'fomulate' (C.1.1), 'idiosynchracies' (C.3), and repeated 'the the' in several prompts. A careful proofread is recommended.
- [§2.2 / Appendix A.2.1] The single-pass model is described as an ablation of the iterative model, but its UtteranceProposer prompt differs (it does not constrain initial utterances to a single feature). This is a reasonable design choice, but the difference between the two models is not purely the iteration loop; the prompt change is a confound. This should be acknowledged explicitly.
- [§4.1 / Algorithm 4] The PlausibilityEvaluator uses empirical mutual information log P(a|u)/P(a) computed from LM token log-probabilities. This is an ad-hoc scoring rule and is not validated against human judgments in the paper. Please clarify its status as an assumption of the model, and ideally provide a small validation or ablation.
Circularity Check
No significant circularity: SAGE predictions are checked against external human data and prior theoretical targets, not against model-internal definitions.
full rationale
The paper's derivation chain is: task analysis defines proposer/evaluator/selector modules; LM modules generate and evaluate alternatives; rule-based selectors produce predictions; predictions are compared to external data. No step equates an input with an output by construction. In Case Study I, reference-game contrastivity is scored by manual annotation, with the paper explicitly stating: 'We use careful manual annotation of all simulation runs ... because the contrastivity calculation builds on the semantic evaluator, which may be challenging for LMs. This makes the evaluation more robust and less circular.' Thus the LM evaluator's output does not define the reported accuracy. In Case Study II, the M-implicature 'correct' answers come from Wilson & Katsos (2016) materials and theory; the MB model's algorithm operationalizes markedness blocking rather than reading off the target. In Case Study III, evaluations use freshly collected human forced-choice data and J. Hu et al. (2023) human data; the Gricean assumptions are borrowed from external theory, not fitted to the human responses. Component-level analyses use human naturalness ratings and benchmark NLI sets (SuperGLUE/SNLI), providing independent checks. The self-citations (e.g., Tsvilodub et al. 2024 for Case Study I, Tsvilodub, Hawkins, & Franke, 2025) are contextual and non-load-bearing: the relevant module details and evaluations are reproduced in the Appendices. Disclosed prompt sensitivities (e.g., DifferentialComplexityEvaluator 'was rather sensitive to details of the prompt') indicate a validity limitation, but no parameter is fitted to the outcome measure, so the predictions do not reduce to their inputs. No specific circular reduction can be exhibited.
Axiom & Free-Parameter Ledger
free parameters (6)
- SemanticEvaluator prompt formulation =
final prompt in Appendix A.1.2
- DifferentialComplexityEvaluator prompt =
final prompt in Appendix B.2
- AssumptionEvaluator prompt =
final prompt in Appendix C.1.1
- Sample sizes (n) =
4/8/10 utterances; 3 alternatives; 4 interpretations
- Max iterations in IM =
5
- LM backends and sampling parameters =
GPT-3.5-turbo tau=0.1; GPT-4o; Llama-3.1-8b-Instruct tau=0.8, topP=0.9, rep. penalty 1.8; text-davinci-003 for plausibil
axioms (5)
- domain assumption The Incremental Algorithm (Dale & Reiter 1995) is an appropriate task analysis of human referential expression generation
- domain assumption Markedness-blocking (per Jäger 2002 / bi-directional OT) explains I-/M-implicatures; marked expressions block typical interpretations
- domain assumption Gricean maxims, as decomposed into the sub-assumptions in Table 4, are the right assumptions for abductive implicature interpretation
- domain assumption LMs can approximate human intuitive commonsense knowledge (typicality, naturalness) under zero-shot prompting
- ad hoc to paper The plausibility of an assumption violation is proportional to empirical mutual information log P(a|u)/P(a) computed from LM token log-probabilities
invented entities (3)
-
SAGE framework (proposer/evaluator/selector modules)
independent evidence
-
Markedness-blocking (MB) model
independent evidence
-
Assumption-evaluation (AE) model
independent evidence
Cite this review
Pith. "Pith review of Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives." pith.science (2026). https://pith.science/paper/NOIWQBM2
@misc{pith2026260718443,
author = {Pith},
title = {Pith review of: Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives},
year = {2026},
howpublished = {\url{https://pith.science/paper/NOIWQBM2}},
note = {Machine review of arXiv:2607.18443}
}
read the original abstract
Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have chosen, or the alternative interpretations a listener might entertain. Formal and computational models of pragmatics must therefore specify the sets of alternatives that interlocutors reason over, which is often done through manual specification. Here we propose a framework, ScAffolded Generative models for Explanation (SAGE), that combines the explanatory transparency of cognitive models with the generative flexibility of language models (LMs). SAGE decomposes a pragmatic process into three kinds of modules: proposers, which use LMs to generate an open-ended space of candidate alternatives; evaluators, which assess those alternatives (e.g., their semantics, complexity, or typicality); and selectors, which implement the rule-based computational steps of a cognitively motivated task analysis. We assess SAGE in three case studies spanning pragmatic generation and interpretation-referential expression generation, manner (M-)implicatures, and Gricean conversational implicatures. SAGE models are evaluated critically using established methods from computational cognitive modeling, including ablations, baseline comparisons, and quantitative fit to human data. Across studies, SAGE models achieved high accuracy and often outperformed baselines, but component-level analyses reveal an asymmetry: LM proposers reliably generated alternatives well-suited to pragmatic modeling, whereas LM evaluators are better at providing intuitive judgements rather than judgements of theoretical or formal measures. We discuss the promise and the limitations of neuro-symbolic models as candidate explanatory accounts of human pragmatic language use.
Figures
Reference graph
Works this paper leans on
-
[1]
Haaf and Jeffrey N
Julia M. Haaf and Jeffrey N. Rouder , doi =. Some do and some don't? Accounting for variability of individual difference structures , volume =. Psychonomic Bulletin & Review , pages =
-
[2]
Kidd, Evan and Donnelly, Seamus and Christiansen, Morten H. , doi =. Individual Differences in Language Acquisition and Processing , volume =. Trends in Cognitive Sciences , number =
-
[3]
and West, Richard F
Stanovich, Keith E. and West, Richard F. , doi =. Individual differences in reasoning: Implications for the rationality debate? , volume =. Behavioral and Brain Sciences , number =
-
[4]
Some Notes on the Formal Properties of Bidirectional Optimality Theory , volume =
Gerhard J. Some Notes on the Formal Properties of Bidirectional Optimality Theory , volume =. Journal of Logic, Language and Information , number =
-
[5]
arXiv , author =:2410.20268 , primaryclass =
Centaur: a foundation model of human cognition , url =. arXiv , author =:2410.20268 , primaryclass =
-
[6]
Computational Brain & Behavior , pages =
Rutar, Danaja and Wolff, Erwin de and Rooij, Iris van and Kwisthout, Johan , doi =. Computational Brain & Behavior , pages =
-
[7]
The Stanford Encyclopedia of Philosophy , note =
Optimality-Theoretic and Game-Theoretic Approaches to Implicatures , year =. The Stanford Encyclopedia of Philosophy , note =
-
[8]
Language and strategic inference , year =
Prashant Parikh , school =. Language and strategic inference , year =
-
[9]
Horn , booktitle =
Laurence R. Horn , booktitle =. Towards a New Taxonomy for Pragmatic Inference:
-
[10]
Structurally-Defined Alternatives , volume =
Roni Katzir , doi =. Structurally-Defined Alternatives , volume =. Linguistics and Philosophy , number =
-
[11]
Meaning and Alternatives , url =
Gotzner, Nicole and Romoli, Jacopo , date-added =. Meaning and Alternatives , url =. Annual Review of Linguistics , number =. doi:10.1146/annurev-linguistics-031220-012013 , issn =
-
[12]
Game Theory and Pragmatics , year =
-
[13]
Quantity Implicatures, Exhaustive Interpretation, and Rational Conversation , volume =
Michael Franke , doi =. Quantity Implicatures, Exhaustive Interpretation, and Rational Conversation , volume =. Semantics & Pragmatics , keywords =
-
[14]
Conceptual alternatives: Competition in language and beyond , url =
Buccola, Brian and Kri. Conceptual alternatives: Competition in language and beyond , url =. doi:10.1007/s10988-021-09327-w , journal =
-
[15]
On the Characterization of Alternatives , volume =
Danny Fox and Roni Katzir , doi =. On the Characterization of Alternatives , volume =. Natural Language Semantics , pages =
-
[16]
The Role of Alternatives in Language , url =
Repp, Sophie and Spalek, Katharina , doi =. The Role of Alternatives in Language , url =. Frontiers in Communication , publisher =
-
[17]
Relevance: Communication and Cognition (2nd ed.) , year =
Dan Sperber and Deirdre Wilson , publisher =. Relevance: Communication and Cognition (2nd ed.) , year =
-
[18]
Quantity Implicatures , year =
Bart Geurts , date-added =. Quantity Implicatures , year =
-
[19]
Daniel Lassiter and Noah D. Goodman , date-added =. Adjectival vagueness in a Bayesian model of interpretation , volume =. doi:10.1007/s11229-015-0786-1 , journal =
-
[20]
Modeling atypicality inferences in pragmatic reasoning , year =
Kravtchenko, Ekaterina and Demberg, Vera , booktitle =. Modeling atypicality inferences in pragmatic reasoning , year =
-
[21]
Goodman , booktitle =
Leon Bergen and Roger Levy and Noah D. Goodman , booktitle =. That's what she (could have) said:
-
[22]
Optimality Theory and Pragmatics , year =
-
[23]
Hypothesis Only Baselines in Natural Language Inference , url =
Poliak, Adam and Naradowsky, Jason and Haldar, Aparajita and Rudinger, Rachel and Van Durme, Benjamin , booktitle =. Hypothesis Only Baselines in Natural Language Inference , url =. doi:10.18653/v1/S18-2023 , pages =
-
[24]
On the opportunities and risks of foundation models , year =
Bommasani, Rishi and Hudson, Drew A and Adeli, Ehsan and Altman, Russ and Arora, Simran and von Arx, Sydney and Bernstein, Michael S and Bohg, Jeannette and Bosselut, Antoine and Brunskill, Emma and others , journal =. On the opportunities and risks of foundation models , year =
-
[25]
Surface Form Competition: Why the Highest Probability Answer Isn
Holtzman, Ari and West, Peter and Shwartz, Vered and Choi, Yejin and Zettlemoyer, Luke , booktitle =. Surface Form Competition: Why the Highest Probability Answer Isn
-
[26]
Manner implicatures and how to spot them , volume =
Jessica Rett , journal =. Manner implicatures and how to spot them , volume =
-
[27]
Goodman , journal =
Leon Bergen and Roger Levy and Noah D. Goodman , journal =. Pragmatic Reasoning through Semantic Inference , volume =
-
[28]
Signal to Act:
Michael Franke , school =. Signal to Act:
-
[29]
Pragmatic Back-and-Forth Reasoning , year =
Michael Franke and Gerhard J. Pragmatic Back-and-Forth Reasoning , year =. Semantics, Pragmatics and the Case of Scalar Implicatures , chapter =
-
[30]
Some Aspects of Optimality in Natural Language Interpretation , volume =
Reinhard Blutner , journal =. Some Aspects of Optimality in Natural Language Interpretation , volume =
-
[31]
doi:10.1162/coli_a_00480 , journal =
Dimensions of Explanatory Value in NLP models , year =. doi:10.1162/coli_a_00480 , journal =
-
[32]
Prashant Parikh , booktitle =
-
[33]
Abduction, Belief and Context in Dialogue , year =
Harry Bunt and William Black , publisher =. Abduction, Belief and Context in Dialogue , year =
-
[34]
Hobbs and Mark Stickel and Paul Martin , journal =
Jerry R. Hobbs and Mark Stickel and Paul Martin , journal =. Interpretation as Abduction , volume =
-
[35]
interaction engine
Stephen C. Levinson , booktitle =. On the human "interaction engine" , year =
-
[36]
Advances in Neural Information Processing Systems , volume=
Training language models to follow instructions with human feedback , author=. Advances in Neural Information Processing Systems , volume=
-
[37]
arXiv preprint arXiv:2210.11416 , year=
Scaling instruction-finetuned language models , author=. arXiv preprint arXiv:2210.11416 , year=
-
[38]
Can AI language models replace human participants? , journal =
Danica Dillion and Niket Tandon and Yuling Gu and Kurt Gray , keywords =. Can AI language models replace human participants? , journal =. 2023 , issn =. doi:https://doi.org/10.1016/j.tics.2023.04.008 , url =
-
[39]
and Nye, Maxwell and Andreas, Jacob
Li, Belinda Z. and Nye, Maxwell and Andreas, Jacob. Implicit Representations of Meaning in Neural Language Models. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. doi:10.18653/v1/2021.acl-long.143
-
[40]
2023 , eprint=
Evaluating Pragmatic Abilities of Image Captioners on A3DS , author=. 2023 , eprint=
2023
-
[41]
Pre-proceedings of Trends in Experimental Pragmatics , pages=
In a manner of speaking: an empirical investigation of Manner Implicatures , author=. Pre-proceedings of Trends in Experimental Pragmatics , pages=
-
[42]
Speech acts , pages=
Logic and conversation , author=. Speech acts , pages=. 1975 , publisher=
1975
-
[43]
2000 , publisher=
Presumptive meanings: The theory of generalized conversational implicature , author=. 2000 , publisher=
2000
-
[44]
Computational Cognitive Modeling and Linguistic Theory , year =
Jakub Dotla. Computational Cognitive Modeling and Linguistic Theory , year =
-
[45]
Computational Linguistics , volume=
Computational generation of referring expressions: A survey , author=. Computational Linguistics , volume=. 2012 , publisher=
2012
-
[46]
Cognitive science , volume=
Computational interpretations of the Gricean maxims in the generation of referring expressions , author=. Cognitive science , volume=. 1995 , publisher=
1995
-
[47]
Journal of Artificial Intelligence Research , volume=
Survey of the state of the art in natural language generation: Core tasks, applications and evaluation , author=. Journal of Artificial Intelligence Research , volume=
-
[48]
1972 , publisher=
Human problem solving , author=. 1972 , publisher=
1972
-
[49]
Convention , publisher=
Lewis, David , journal=. Convention , publisher=
-
[50]
Science , volume=
Predicting pragmatic reasoning in language games , author=. Science , volume=. 2012 , publisher=
2012
-
[51]
Cognitive science , volume=
Characterizing the dynamics of learning in repeated reference games , author=. Cognitive science , volume=. 2020 , publisher=
2020
-
[52]
A game-theoretic approach to generating spatial descriptions , author=
-
[53]
population-level probabilistic modeling , author=
Reasoning in reference games: Individual-vs. population-level probabilistic modeling , author=. PloS one , volume=. 2016 , publisher=
2016
-
[54]
OpenAI blog , volume=
Language models are unsupervised multitask learners , author=. OpenAI blog , volume=
-
[55]
Language Models are Few-Shot Learners , volume =
Brown, Tom and Mann, Benjamin and Ryder, Nick and Subbiah, Melanie and Kaplan, Jared D and Dhariwal, Prafulla and Neelakantan, Arvind and Shyam, Pranav and Sastry, Girish and Askell, Amanda and Agarwal, Sandhini and Herbert-Voss, Ariel and Krueger, Gretchen and Henighan, Tom and Child, Rewon and Ramesh, Aditya and Ziegler, Daniel and Wu, Jeffrey and Winte...
-
[56]
arXiv preprint arXiv:2204.02329 , year=
Can language models learn from explanations in context? , author=. arXiv preprint arXiv:2204.02329 , year=
-
[57]
arXiv preprint arXiv:2303.12712 , year=
Sparks of artificial general intelligence: Early experiments with gpt-4 , author=. arXiv preprint arXiv:2303.12712 , year=
-
[58]
2023 , eprint=
GPT-4 Technical Report , author=. 2023 , eprint=
2023
-
[59]
arXiv preprint arXiv:2204.02311 , year=
Palm: Scaling language modeling with pathways , author=. arXiv preprint arXiv:2204.02311 , year=
-
[60]
arXiv preprint arXiv:2302.13971 , year=
Llama: Open and efficient foundation language models , author=. arXiv preprint arXiv:2302.13971 , year=
-
[61]
arXiv e-prints , pages=
The llama 3 herd of models , author=. arXiv e-prints , pages=
-
[62]
arXiv preprint arXiv:2201.11903 , year=
Chain of thought prompting elicits reasoning in large language models , author=. arXiv preprint arXiv:2201.11903 , year=
-
[63]
International Conference on Learning Representations (ICLR) , year=
React: Synergizing reasoning and acting in language models , author=. International Conference on Learning Representations (ICLR) , year=
-
[64]
arXiv preprint arXiv:2111.02080 , year=
An explanation of in-context learning as implicit bayesian inference , author=. arXiv preprint arXiv:2111.02080 , year=
-
[65]
arXiv preprint arXiv:2202.12837 , year=
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? , author=. arXiv preprint arXiv:2202.12837 , year=
-
[66]
arXiv preprint arXiv:2211.10435 , year=
PAL: Program-aided Language Models , author=. arXiv preprint arXiv:2211.10435 , year=
-
[67]
2023 , eprint=
From Word Models to World Models: Translating from Natural Language to the Probabilistic Language of Thought , author=. 2023 , eprint=
2023
-
[68]
arXiv preprint arXiv:2305.10601 , year=
Tree of thoughts: Deliberate problem solving with large language models , author=. arXiv preprint arXiv:2305.10601 , year=
-
[69]
arXiv preprint arXiv:2108.07258 , year=
On the opportunities and risks of foundation models , author=. arXiv preprint arXiv:2108.07258 , year=
-
[70]
2022 , eprint=
Language Model Cascades , author=. 2022 , eprint=
2022
-
[71]
2023 , eprint=
Faithful Chain-of-Thought Reasoning , author=. 2023 , eprint=
2023
-
[72]
2023 , eprint=
Toolformer: Language Models Can Teach Themselves to Use Tools , author=. 2023 , eprint=
2023
-
[73]
2022 , eprint=
Atlas: Few-shot Learning with Retrieval Augmented Language Models , author=. 2022 , eprint=
2022
-
[74]
A fine-grained comparison of pragmatic language understanding in humans and language models
Hu, Jennifer and Floyd, Sammy and Jouravlev, Olessia and Fedorenko, Evelina and Gibson, Edward. A fine-grained comparison of pragmatic language understanding in humans and language models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.230
-
[75]
Generated Knowledge Prompting for Commonsense Reasoning
Liu, Jiacheng and Liu, Alisa and Lu, Ximing and Welleck, Sean and West, Peter and Le Bras, Ronan and Choi, Yejin and Hajishirzi, Hannaneh. Generated Knowledge Prompting for Commonsense Reasoning. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. doi:10.18653/v1/2022.acl-long.225
-
[76]
2023 , eprint=
Invalid Logic, Equivalent Gains: The Bizarreness of Reasoning in Language Model Prompting , author=. 2023 , eprint=
2023
-
[77]
, author=
The language of generalization. , author=. Psychological review , volume=. 2019 , publisher=
2019
-
[78]
The handbook of pragmatics , pages=
Implicature , author=. The handbook of pragmatics , pages=. 2006 , publisher=
2006
-
[79]
, author=
Animal, dog, or dalmatian? Level of abstraction in nominal referring expressions. , author=. CogSci , year=
-
[80]
overinformative
When redundancy is useful: A Bayesian approach to “overinformative” referring expressions. , author=. Psychological Review , volume=. 2020 , publisher=
2020
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.