Pith. sign in

REVIEW 3 major objections 3 minor 207 references

This paper claims that SAGE—splitting pragmatic modeling into LM proposers and evaluators over a symbolic task analysis—lets cognitive models scale to open-ended alternatives, but the LMs are reliable as proposers, not as formal evaluators.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Combining language-model generation with rule-based selection reproduces several pragmatic phenomena, but the language models only worked reliably as idea generators, not as judges of formal linguistic properties.

T0 review reviewed 2026-08-01 challenge →

load-bearing objection A careful, honest framework paper whose headline asymmetry (good proposers, shaky formal evaluators) is real and well-documented—more roadmap than finished solution, and worth engaging despite unquantified prompt tuning. the 3 major comments →

arxiv 2607.18443 v1 pith:NOIWQBM2 submitted 2026-07-20 cs.CL

Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives

classification cs.CL
keywords pragmaticslanguage modelsneuro-symbolic modelingalternativesimplicaturereferential expression generationcognitive modelingSAGE framework
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that a neuro-symbolic framework called SAGE is a viable path to open-ended cognitive models of pragmatic language use. SAGE decomposes a pragmatic task into LM-based proposers that generate candidate expressions or interpretations, LM-based evaluators that assess them, and rule-based selectors that implement a cognitively motivated task analysis. Across three case studies—referential expression generation, manner implicatures, and Gricean conversational implicatures—the end-to-end models achieve high accuracy and often outperform simple LM baselines. The component-level message is an asymmetry: LM proposers generate alternatives that are well-suited to pragmatic modeling, while LM evaluators are better at intuitive judgments than at judgments of formal or theoretical measures. If the paper is right, the bottleneck for open-ended pragmatic modeling is no longer supplying alternatives by hand; it is building trustworthy evaluators.

Core claim

SAGE models can be constructed for pragmatic production and interpretation by replacing manual alternative-specification with LM generation, while keeping the reasoning steps transparent. The paper's detailed module evaluations show that the LM proposers produce natural, contextually appropriate alternatives—rated by humans as comparable to human-written ones—whereas LM evaluators, when asked to make formal judgments such as literal semantic truth, differential complexity, or Gricean maxim flouting, are prompt-sensitive and often diverge from human judgments. The central discovery is therefore not just that the framework works end-to-end, but that the appropriate division of labor between ne

What carries the argument

The SAGE framework (ScAffolded Generative models for Explanation), built from three module types: proposers, evaluators, and selectors. Proposers use LMs to sample an open-ended space of candidate alternatives; evaluators assess those alternatives on dimensions like literal truth, complexity, or prior plausibility; selectors apply rule-based operations from an explicit task analysis. The central contribution is treating alternative-generation as a flexible LM subroutine while keeping the cognitive reasoning steps explicit and testable. The case studies instantiate the machinery in three task analyses: the Incremental Algorithm for referential expression generation, a markedness-blocking proc

Load-bearing premise

The framework's viability rests on the assumption that a single zero-shot LM call can reliably perform the evaluator roles the task analysis requires—deciding literal truth, comparing expression complexity, and detecting Gricean maxim flouting.

What would settle it

A benchmark study in which the evaluator modules are tested on held-out items without prompt tuning: if the semantic evaluator's accuracy on entailment pairs falls to chance, or if replacing the LM evaluators with random or fixed heuristics in the end-to-end SAGE pipeline does not reduce accuracy below the human-fit level, the claim that LM evaluators carry the formal judgment load would be falsified. The paper's own appendix already reports the semantic evaluator at 0.82 accuracy and near-universal maxim-flouting flags, so a broader replication of these failures on new stimuli would settle th

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Manual specification of alternatives can be replaced by LM generation for a range of pragmatic tasks, opening models to open-ended contexts.
  • End-to-end accuracy of SAGE models is not sufficient evidence of component adequacy; module-level evaluation against human judgments is necessary.
  • The asymmetry between proposers and evaluators suggests practical guidance: use LMs for sampling alternatives and intuitive judgments; supply formal judgments from specialized components such as fine-tuned models, probability scoring, or symbolic methods.
  • SAGE provides a concrete implementation route for verbal Gricean theory, enabling quantitative comparison of different assumption sets against human data.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A likely testable extension is replacing prompt-based evaluators with LM-internal probability scoring, such as conditional log-probabilities; if that improves evaluator reliability, the asymmetry would narrow.
  • The observed evaluator weakness for abstract judgments suggests that neuro-symbolic cognitive models should keep formal reasoning outside the LM or fine-tune dedicated evaluators rather than rely on zero-shot prompts.
  • The no-assumption model's comparable fit to the Gricean model hints that maxim-flouting detection may not be the driving component of implicature interpretation; a testable prediction is that omitting assumption evaluation yields similar or better fits on new implicature datasets.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper introduces SAGE (ScAffolded Generative models for Explanation), a neuro-symbolic framework for cognitive modeling of pragmatics. SAGE decomposes a pragmatic task into LM-based proposers (which generate open-ended utterance/interpretation alternatives), LM-based evaluators (which assess semantics, complexity, typicality, or maxim violation), and rule-based selectors that implement the symbolic task analysis. The framework is evaluated in three case studies: referential expression generation in a reference game, M-implicature interpretation from periphrastic causatives, and Gricean conversational implicature interpretation. The models are assessed with accuracy, ablations, baselines, human module ratings, and quantitative fit to human forced-choice data. The headline finding is an asymmetry: LM proposers generate viable alternatives, while LM evaluators perform well on intuitive judgments but are less reliable for formal/theoretical assessments, such as literal truth, differential complexity, and maxim flouting.

Significance. If the findings hold, the paper makes a useful methodological contribution: it demonstrates a concrete way to combine LMs with transparent symbolic task analyses for pragmatics, and it provides a systematic, honest assessment of which LM subtasks are currently viable. The paper is particularly strong in its evaluative practice: it uses freshly collected human data, externally grounded human module ratings, Bayesian model comparison, and it openly reports component failures and selection effects. The claimed proposer/evaluator asymmetry is a valuable empirical insight for the growing literature on neuro-symbolic cognitive models. The main risk is that the framework's promise of open-endedness depends on evaluator reliability, which the paper itself shows to be limited; the paper should therefore be read as a proof-of-concept with clear bottlenecks, not as a demonstration that all SAGE components already work.

major comments (3)
  1. [§3.3 / Appendix B.4] The markedness-blocking model discards runs in which no unblocked state-utterance pair is available. This introduces a selection effect on the reported M-implicature accuracy: the model is scored only on cases where the blocking mechanism successfully produces at least one unblocked interpretation. Since blocking is exactly the mechanism the model is intended to explain, the paper should report (i) the proportion of discarded runs, (ii) whether the discarded runs are systematically different (e.g., particular vignettes or utterance types), and (iii) accuracy when discarded runs are counted as failures. Without this, the 80% M-implicature accuracy is conditional on a favorable outcome of the proposed algorithm rather than an unconditional model prediction.
  2. [§4.2.1 / Appendix C.1.1] The assumption-evaluation module flags almost all maxims as violated, while humans are much more reluctant; in the no-assumptions ablation the model still achieves accuracy and human-data fit close to the full Gricean model. The paper acknowledges these facts, but their implications for the central claim are not sufficiently resolved. The Gricean AE model is credibly better than the no-assumptions model, but the differences are small, and the module-level evaluation shows that the assumption evaluator's output is not human-like. The authors should either provide a more direct attribution analysis (e.g., comparing models where the assumption evaluator is replaced by human violation judgments, or where the plausibility evaluator is ablated) or explicitly restrict the scope of the conclusion to the proposer-based component of SAGE.
  3. [§2.1 / Appendix A.1.2] The SemanticEvaluator prompt was optimized during development and the final module obtains only 0.82 accuracy on NLI-style and matched test sets. More importantly, the evaluation of the iterative model's final contrastivity is based on manual annotation by the authors, not on the module's own outputs. The paper notes that the model sometimes failed to recognize human-fully-contrastive utterances and iterated further. This makes it difficult to know how much of the IM's success is due to the semantic evaluator as opposed to the proposer and the manual evaluation procedure. Please report end-to-end contrastivity using the raw SemanticEvaluator outputs, and quantify how often the evaluator's errors changed the number of iterations or the selected utterance.
minor comments (3)
  1. [General] There are several typos and misspellings: 'pehnomena' (§2.4), 'interprepretation' (Appendix C heading), 'ConstrastivitySelector' (Algorithm 1), 'GTP-3.5-turbo' (§4.3, C.1.1), 'fomulate' (C.1.1), 'idiosynchracies' (C.3), and repeated 'the the' in several prompts. A careful proofread is recommended.
  2. [§2.2 / Appendix A.2.1] The single-pass model is described as an ablation of the iterative model, but its UtteranceProposer prompt differs (it does not constrain initial utterances to a single feature). This is a reasonable design choice, but the difference between the two models is not purely the iteration loop; the prompt change is a confound. This should be acknowledged explicitly.
  3. [§4.1 / Algorithm 4] The PlausibilityEvaluator uses empirical mutual information log P(a|u)/P(a) computed from LM token log-probabilities. This is an ad-hoc scoring rule and is not validated against human judgments in the paper. Please clarify its status as an assumption of the model, and ideally provide a small validation or ablation.

Circularity Check

0 steps flagged

No significant circularity: SAGE predictions are checked against external human data and prior theoretical targets, not against model-internal definitions.

full rationale

The paper's derivation chain is: task analysis defines proposer/evaluator/selector modules; LM modules generate and evaluate alternatives; rule-based selectors produce predictions; predictions are compared to external data. No step equates an input with an output by construction. In Case Study I, reference-game contrastivity is scored by manual annotation, with the paper explicitly stating: 'We use careful manual annotation of all simulation runs ... because the contrastivity calculation builds on the semantic evaluator, which may be challenging for LMs. This makes the evaluation more robust and less circular.' Thus the LM evaluator's output does not define the reported accuracy. In Case Study II, the M-implicature 'correct' answers come from Wilson & Katsos (2016) materials and theory; the MB model's algorithm operationalizes markedness blocking rather than reading off the target. In Case Study III, evaluations use freshly collected human forced-choice data and J. Hu et al. (2023) human data; the Gricean assumptions are borrowed from external theory, not fitted to the human responses. Component-level analyses use human naturalness ratings and benchmark NLI sets (SuperGLUE/SNLI), providing independent checks. The self-citations (e.g., Tsvilodub et al. 2024 for Case Study I, Tsvilodub, Hawkins, & Franke, 2025) are contextual and non-load-bearing: the relevant module details and evaluations are reproduced in the Appendices. Disclosed prompt sensitivities (e.g., DifferentialComplexityEvaluator 'was rather sensitive to details of the prompt') indicate a validity limitation, but no parameter is fitted to the outcome measure, so the predictions do not reduce to their inputs. No specific circular reduction can be exhibited.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 3 invented entities

The paper borrows its task analyses (IA, bi-directional OT, Gricean maxims) from prior literature — legitimate scaffolding — but adds six hand-set modeling choices (prompt formulations fitted during development, sample sizes, iteration caps, backends). The most load-bearing premise, that zero-shot LMs can evaluate formal linguistic properties, is partially contradicted by the paper's own module evaluations, and the paper says so. No invented physical entities; the invented entities are model architectures with falsifiable output distributions.

free parameters (6)
  • SemanticEvaluator prompt formulation = final prompt in Appendix A.1.2
    Prompt was adjusted based on evaluation results during development; the choices among 'logical compatibility', 'true', 'contradictions', 'new information' phrasings are researcher degrees of freedom fitted to module accuracy.
  • DifferentialComplexityEvaluator prompt = final prompt in Appendix B.2
    Multiple prompt variants (chain-of-thought, few-shot, wh-question vs statement) were tested and the most robust chosen; this fitting affects whether utterance-state pairs are blocked, which determines M-implicature predictions.
  • AssumptionEvaluator prompt = final prompt in Appendix C.1.1
    Polarity of the main question ('doubt' vs 'is true') was tuned based on manual output inspection; the module's violation judgments drive the whole Gricean pipeline.
  • Sample sizes (n) = 4/8/10 utterances; 3 alternatives; 4 interpretations
    Numbers of LM samples per module were chosen by hand; the paper tests n=4 vs n=8 in case study I and finds no difference, but n is not justified for cases II-III.
  • Max iterations in IM = 5
    Hard cap on the Incremental Algorithm loop; affects which utterances reach the InfoMaxSelector.
  • LM backends and sampling parameters = GPT-3.5-turbo tau=0.1; GPT-4o; Llama-3.1-8b-Instruct tau=0.8, topP=0.9, rep. penalty 1.8; text-davinci-003 for plausibil
    Model choices and hyperparameters are selected for observed performance; the paper reports backbone-dependent results, showing that framework behavior is not backbone-invariant.
axioms (5)
  • domain assumption The Incremental Algorithm (Dale & Reiter 1995) is an appropriate task analysis of human referential expression generation
    Borrowed from prior literature; the paper explicitly notes it is 'not a strong contender for a cognitively plausible model' yet uses it as scaffolding (Section 2).
  • domain assumption Markedness-blocking (per Jäger 2002 / bi-directional OT) explains I-/M-implicatures; marked expressions block typical interpretations
    The theoretical premise of case study II; if false, the MB model's accuracy would not bear on human cognition (Section 3.1).
  • domain assumption Gricean maxims, as decomposed into the sub-assumptions in Table 4, are the right assumptions for abductive implicature interpretation
    The decomposition into 9 sub-maxims is the authors' operationalization; humans and the LM both showed weak condition-specific structure in violation judgments (Appendix C.1.1).
  • domain assumption LMs can approximate human intuitive commonsense knowledge (typicality, naturalness) under zero-shot prompting
    Central to the whole SAGE approach; partially supported by human module ratings, but only for proposers and intuitive evaluators, not formal evaluators (Sections 5.1, B.3).
  • ad hoc to paper The plausibility of an assumption violation is proportional to empirical mutual information log P(a|u)/P(a) computed from LM token log-probabilities
    A modeling choice for selecting one violated assumption; no independent validation that this measure matches human selection of the 'most plausible' flouted maxim (Section 4.1).
invented entities (3)
  • SAGE framework (proposer/evaluator/selector modules) independent evidence
    purpose: General decomposition for neuro-symbolic cognitive modeling of pragmatics
    The framework itself is the paper's main invention; its predictions are testable through the case-study pipelines, and the proposer/evaluator asymmetry is a falsifiable claim about LM capabilities.
  • Markedness-blocking (MB) model independent evidence
    purpose: Computational account of I-/M-implicatures via alternative generation and blocking
    Produces accuracy predictions on 15 vignettes that can be compared against human choices; tested against human data only indirectly (Section 3, Appendix B).
  • Assumption-evaluation (AE) model independent evidence
    purpose: Computational implementation of Gricean abductive implicature interpretation
    Predicts human forced-choice distributions; fit to human data is the falsifiable handle, though the no-assumptions ablation nearly matches it, weakening the entity's specificity (Section 4.2.1).

reviewed 2026-08-01 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives." pith.science (2026). https://pith.science/paper/NOIWQBM2

@misc{pith2026260718443,
  author       = {Pith},
  title        = {Pith review of: Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NOIWQBM2}},
  note         = {Machine review of arXiv:2607.18443}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have chosen, or the alternative interpretations a listener might entertain. Formal and computational models of pragmatics must therefore specify the sets of alternatives that interlocutors reason over, which is often done through manual specification. Here we propose a framework, ScAffolded Generative models for Explanation (SAGE), that combines the explanatory transparency of cognitive models with the generative flexibility of language models (LMs). SAGE decomposes a pragmatic process into three kinds of modules: proposers, which use LMs to generate an open-ended space of candidate alternatives; evaluators, which assess those alternatives (e.g., their semantics, complexity, or typicality); and selectors, which implement the rule-based computational steps of a cognitively motivated task analysis. We assess SAGE in three case studies spanning pragmatic generation and interpretation-referential expression generation, manner (M-)implicatures, and Gricean conversational implicatures. SAGE models are evaluated critically using established methods from computational cognitive modeling, including ablations, baseline comparisons, and quantitative fit to human data. Across studies, SAGE models achieved high accuracy and often outperformed baselines, but component-level analyses reveal an asymmetry: LM proposers reliably generated alternatives well-suited to pragmatic modeling, whereas LM evaluators are better at providing intuitive judgements rather than judgements of theoretical or formal measures. We discuss the promise and the limitations of neuro-symbolic models as candidate explanatory accounts of human pragmatic language use.

Figures

Figures reproduced from arXiv: 2607.18443 by Fausto Carcassi, Michael Franke, Polina Tsvilodub.

Figure 1
Figure 1. Figure 1: (A) Overview of our framework ScAffolded Generative models for Explanation (SAGE, in￾dicate here with the symbol ) for more open-ended cognitive modeling of human pragmatic language use. The framework identifies three kinds of modules that are scaffolded by a cognitively motivated explanatory task analysis: proposers, evaluators and selectors. We use language models to instan￾tiate neural modules ( ) that … view at source ↗
Figure 2
Figure 2. Figure 2: Side-by-side illustration of two models for contrastive utterance generation. In both models, [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: A: Reference game results: distribution over contrastivity values (y-axis) by number of distractors (x-axis) and number of utterances proposed (color). Error bars show bootstrapped 95%- CIs. B: Distribution over contrastivity values (y-axis) over increasing tree depth in the iterated model (extended utterance proposal and evaluation iterations; x-axis), by number of distractors (facets) and tree width (num… view at source ↗
Figure 4
Figure 4. Figure 4: A: Example of an experimental item for testing I- and M-implicature inferences. For each vignette, an unmarked and a marked utterance were presented with two paraphrases describing situations the speaker might have intended to convey, a typical situation and an atypical situation. We expect the typical situation to be chosen for the unmarked utterance (the I-implicature). We expect the atypical situation t… view at source ↗
Figure 5
Figure 5. Figure 5: A: Gricean assumptions used in the interpretation algorithm for implicatures. B: Flow chart of the Assumption-Evaluation (AE) model for general pragmatic interpretation. The assumption evaluator checks whether each of a list of assumptions holds for the trigger utterance, after which an evaluator select the most likely among the violated assumptions, if the set of assumptions is non-empty. An interpretatio… view at source ↗
Figure 6
Figure 6. Figure 6: Examples of an experimental item for testing implicatures possibly arising from violations of [PITH_FULL_IMAGE:figures/full_fig_p016_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Both plots show results across backbones and datasets. [PITH_FULL_IMAGE:figures/full_fig_p017_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Interpretation results in case study III. The average accuracy (i.e., proportion of correct [PITH_FULL_IMAGE:figures/full_fig_p020_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Density plots and summary statistics for the posterior distribution of differences in total [PITH_FULL_IMAGE:figures/full_fig_p021_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Human ratings (y-axis) of the utterance proposals against human-constructed reference [PITH_FULL_IMAGE:figures/full_fig_p039_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Results of human evaluation of the Gricean assumptions for case study III on samples [PITH_FULL_IMAGE:figures/full_fig_p046_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Most likely violations of assumptions identified by the PlausibilityEvaluator in the [PITH_FULL_IMAGE:figures/full_fig_p047_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Results of human evaluation of the InterpretationProposer samples from case study III. [PITH_FULL_IMAGE:figures/full_fig_p049_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Proportions of different responses in different trigger conditions of dataset 1 produced by [PITH_FULL_IMAGE:figures/full_fig_p050_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Interpretation results in case study III. The average accuracy (i.e., proportion of correct [PITH_FULL_IMAGE:figures/full_fig_p052_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Results from pragmatic interpretation tasks for both datasets. Y-axis shows accuracy [PITH_FULL_IMAGE:figures/full_fig_p057_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: Density plots and summary statistics for the posterior distribution of difference in (log) [PITH_FULL_IMAGE:figures/full_fig_p058_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: Density plots and summary statistics for the posterior distribution of differences in total [PITH_FULL_IMAGE:figures/full_fig_p059_18.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

207 extracted references · 9 canonical work pages

  1. [1]

    Haaf and Jeffrey N

    Julia M. Haaf and Jeffrey N. Rouder , doi =. Some do and some don't? Accounting for variability of individual difference structures , volume =. Psychonomic Bulletin & Review , pages =

  2. [2]

    Kidd, Evan and Donnelly, Seamus and Christiansen, Morten H. , doi =. Individual Differences in Language Acquisition and Processing , volume =. Trends in Cognitive Sciences , number =

  3. [3]

    and West, Richard F

    Stanovich, Keith E. and West, Richard F. , doi =. Individual differences in reasoning: Implications for the rationality debate? , volume =. Behavioral and Brain Sciences , number =

  4. [4]

    Some Notes on the Formal Properties of Bidirectional Optimality Theory , volume =

    Gerhard J. Some Notes on the Formal Properties of Bidirectional Optimality Theory , volume =. Journal of Logic, Language and Information , number =

  5. [5]

    arXiv , author =:2410.20268 , primaryclass =

    Centaur: a foundation model of human cognition , url =. arXiv , author =:2410.20268 , primaryclass =

  6. [6]

    Computational Brain & Behavior , pages =

    Rutar, Danaja and Wolff, Erwin de and Rooij, Iris van and Kwisthout, Johan , doi =. Computational Brain & Behavior , pages =

  7. [7]

    The Stanford Encyclopedia of Philosophy , note =

    Optimality-Theoretic and Game-Theoretic Approaches to Implicatures , year =. The Stanford Encyclopedia of Philosophy , note =

  8. [8]

    Language and strategic inference , year =

    Prashant Parikh , school =. Language and strategic inference , year =

  9. [9]

    Horn , booktitle =

    Laurence R. Horn , booktitle =. Towards a New Taxonomy for Pragmatic Inference:

  10. [10]

    Structurally-Defined Alternatives , volume =

    Roni Katzir , doi =. Structurally-Defined Alternatives , volume =. Linguistics and Philosophy , number =

  11. [11]

    Meaning and Alternatives , url =

    Gotzner, Nicole and Romoli, Jacopo , date-added =. Meaning and Alternatives , url =. Annual Review of Linguistics , number =. doi:10.1146/annurev-linguistics-031220-012013 , issn =

  12. [12]

    Game Theory and Pragmatics , year =

  13. [13]

    Quantity Implicatures, Exhaustive Interpretation, and Rational Conversation , volume =

    Michael Franke , doi =. Quantity Implicatures, Exhaustive Interpretation, and Rational Conversation , volume =. Semantics & Pragmatics , keywords =

  14. [14]

    Conceptual alternatives: Competition in language and beyond , url =

    Buccola, Brian and Kri. Conceptual alternatives: Competition in language and beyond , url =. doi:10.1007/s10988-021-09327-w , journal =

  15. [15]

    On the Characterization of Alternatives , volume =

    Danny Fox and Roni Katzir , doi =. On the Characterization of Alternatives , volume =. Natural Language Semantics , pages =

  16. [16]

    The Role of Alternatives in Language , url =

    Repp, Sophie and Spalek, Katharina , doi =. The Role of Alternatives in Language , url =. Frontiers in Communication , publisher =

  17. [17]

    Relevance: Communication and Cognition (2nd ed.) , year =

    Dan Sperber and Deirdre Wilson , publisher =. Relevance: Communication and Cognition (2nd ed.) , year =

  18. [18]

    Quantity Implicatures , year =

    Bart Geurts , date-added =. Quantity Implicatures , year =

  19. [19]

    Goodman , date-added =

    Daniel Lassiter and Noah D. Goodman , date-added =. Adjectival vagueness in a Bayesian model of interpretation , volume =. doi:10.1007/s11229-015-0786-1 , journal =

  20. [20]

    Modeling atypicality inferences in pragmatic reasoning , year =

    Kravtchenko, Ekaterina and Demberg, Vera , booktitle =. Modeling atypicality inferences in pragmatic reasoning , year =

  21. [21]

    Goodman , booktitle =

    Leon Bergen and Roger Levy and Noah D. Goodman , booktitle =. That's what she (could have) said:

  22. [22]

    Optimality Theory and Pragmatics , year =

  23. [23]

    Hypothesis Only Baselines in Natural Language Inference , url =

    Poliak, Adam and Naradowsky, Jason and Haldar, Aparajita and Rudinger, Rachel and Van Durme, Benjamin , booktitle =. Hypothesis Only Baselines in Natural Language Inference , url =. doi:10.18653/v1/S18-2023 , pages =

  24. [24]

    On the opportunities and risks of foundation models , year =

    Bommasani, Rishi and Hudson, Drew A and Adeli, Ehsan and Altman, Russ and Arora, Simran and von Arx, Sydney and Bernstein, Michael S and Bohg, Jeannette and Bosselut, Antoine and Brunskill, Emma and others , journal =. On the opportunities and risks of foundation models , year =

  25. [25]

    Surface Form Competition: Why the Highest Probability Answer Isn

    Holtzman, Ari and West, Peter and Shwartz, Vered and Choi, Yejin and Zettlemoyer, Luke , booktitle =. Surface Form Competition: Why the Highest Probability Answer Isn

  26. [26]

    Manner implicatures and how to spot them , volume =

    Jessica Rett , journal =. Manner implicatures and how to spot them , volume =

  27. [27]

    Goodman , journal =

    Leon Bergen and Roger Levy and Noah D. Goodman , journal =. Pragmatic Reasoning through Semantic Inference , volume =

  28. [28]

    Signal to Act:

    Michael Franke , school =. Signal to Act:

  29. [29]

    Pragmatic Back-and-Forth Reasoning , year =

    Michael Franke and Gerhard J. Pragmatic Back-and-Forth Reasoning , year =. Semantics, Pragmatics and the Case of Scalar Implicatures , chapter =

  30. [30]

    Some Aspects of Optimality in Natural Language Interpretation , volume =

    Reinhard Blutner , journal =. Some Aspects of Optimality in Natural Language Interpretation , volume =

  31. [31]

    doi:10.1162/coli_a_00480 , journal =

    Dimensions of Explanatory Value in NLP models , year =. doi:10.1162/coli_a_00480 , journal =

  32. [32]

    Prashant Parikh , booktitle =

  33. [33]

    Abduction, Belief and Context in Dialogue , year =

    Harry Bunt and William Black , publisher =. Abduction, Belief and Context in Dialogue , year =

  34. [34]

    Hobbs and Mark Stickel and Paul Martin , journal =

    Jerry R. Hobbs and Mark Stickel and Paul Martin , journal =. Interpretation as Abduction , volume =

  35. [35]

    interaction engine

    Stephen C. Levinson , booktitle =. On the human "interaction engine" , year =

  36. [36]

    Advances in Neural Information Processing Systems , volume=

    Training language models to follow instructions with human feedback , author=. Advances in Neural Information Processing Systems , volume=

  37. [37]

    arXiv preprint arXiv:2210.11416 , year=

    Scaling instruction-finetuned language models , author=. arXiv preprint arXiv:2210.11416 , year=

  38. [38]

    Can AI language models replace human participants? , journal =

    Danica Dillion and Niket Tandon and Yuling Gu and Kurt Gray , keywords =. Can AI language models replace human participants? , journal =. 2023 , issn =. doi:https://doi.org/10.1016/j.tics.2023.04.008 , url =

  39. [39]

    and Nye, Maxwell and Andreas, Jacob

    Li, Belinda Z. and Nye, Maxwell and Andreas, Jacob. Implicit Representations of Meaning in Neural Language Models. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. doi:10.18653/v1/2021.acl-long.143

  40. [40]

    2023 , eprint=

    Evaluating Pragmatic Abilities of Image Captioners on A3DS , author=. 2023 , eprint=

  41. [41]

    Pre-proceedings of Trends in Experimental Pragmatics , pages=

    In a manner of speaking: an empirical investigation of Manner Implicatures , author=. Pre-proceedings of Trends in Experimental Pragmatics , pages=

  42. [42]

    Speech acts , pages=

    Logic and conversation , author=. Speech acts , pages=. 1975 , publisher=

  43. [43]

    2000 , publisher=

    Presumptive meanings: The theory of generalized conversational implicature , author=. 2000 , publisher=

  44. [44]

    Computational Cognitive Modeling and Linguistic Theory , year =

    Jakub Dotla. Computational Cognitive Modeling and Linguistic Theory , year =

  45. [45]

    Computational Linguistics , volume=

    Computational generation of referring expressions: A survey , author=. Computational Linguistics , volume=. 2012 , publisher=

  46. [46]

    Cognitive science , volume=

    Computational interpretations of the Gricean maxims in the generation of referring expressions , author=. Cognitive science , volume=. 1995 , publisher=

  47. [47]

    Journal of Artificial Intelligence Research , volume=

    Survey of the state of the art in natural language generation: Core tasks, applications and evaluation , author=. Journal of Artificial Intelligence Research , volume=

  48. [48]

    1972 , publisher=

    Human problem solving , author=. 1972 , publisher=

  49. [49]

    Convention , publisher=

    Lewis, David , journal=. Convention , publisher=

  50. [50]

    Science , volume=

    Predicting pragmatic reasoning in language games , author=. Science , volume=. 2012 , publisher=

  51. [51]

    Cognitive science , volume=

    Characterizing the dynamics of learning in repeated reference games , author=. Cognitive science , volume=. 2020 , publisher=

  52. [52]

    A game-theoretic approach to generating spatial descriptions , author=

  53. [53]

    population-level probabilistic modeling , author=

    Reasoning in reference games: Individual-vs. population-level probabilistic modeling , author=. PloS one , volume=. 2016 , publisher=

  54. [54]

    OpenAI blog , volume=

    Language models are unsupervised multitask learners , author=. OpenAI blog , volume=

  55. [55]

    Language Models are Few-Shot Learners , volume =

    Brown, Tom and Mann, Benjamin and Ryder, Nick and Subbiah, Melanie and Kaplan, Jared D and Dhariwal, Prafulla and Neelakantan, Arvind and Shyam, Pranav and Sastry, Girish and Askell, Amanda and Agarwal, Sandhini and Herbert-Voss, Ariel and Krueger, Gretchen and Henighan, Tom and Child, Rewon and Ramesh, Aditya and Ziegler, Daniel and Wu, Jeffrey and Winte...

  56. [56]

    arXiv preprint arXiv:2204.02329 , year=

    Can language models learn from explanations in context? , author=. arXiv preprint arXiv:2204.02329 , year=

  57. [57]

    arXiv preprint arXiv:2303.12712 , year=

    Sparks of artificial general intelligence: Early experiments with gpt-4 , author=. arXiv preprint arXiv:2303.12712 , year=

  58. [58]

    2023 , eprint=

    GPT-4 Technical Report , author=. 2023 , eprint=

  59. [59]

    arXiv preprint arXiv:2204.02311 , year=

    Palm: Scaling language modeling with pathways , author=. arXiv preprint arXiv:2204.02311 , year=

  60. [60]

    arXiv preprint arXiv:2302.13971 , year=

    Llama: Open and efficient foundation language models , author=. arXiv preprint arXiv:2302.13971 , year=

  61. [61]

    arXiv e-prints , pages=

    The llama 3 herd of models , author=. arXiv e-prints , pages=

  62. [62]

    arXiv preprint arXiv:2201.11903 , year=

    Chain of thought prompting elicits reasoning in large language models , author=. arXiv preprint arXiv:2201.11903 , year=

  63. [63]

    International Conference on Learning Representations (ICLR) , year=

    React: Synergizing reasoning and acting in language models , author=. International Conference on Learning Representations (ICLR) , year=

  64. [64]

    arXiv preprint arXiv:2111.02080 , year=

    An explanation of in-context learning as implicit bayesian inference , author=. arXiv preprint arXiv:2111.02080 , year=

  65. [65]

    arXiv preprint arXiv:2202.12837 , year=

    Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? , author=. arXiv preprint arXiv:2202.12837 , year=

  66. [66]

    arXiv preprint arXiv:2211.10435 , year=

    PAL: Program-aided Language Models , author=. arXiv preprint arXiv:2211.10435 , year=

  67. [67]

    2023 , eprint=

    From Word Models to World Models: Translating from Natural Language to the Probabilistic Language of Thought , author=. 2023 , eprint=

  68. [68]

    arXiv preprint arXiv:2305.10601 , year=

    Tree of thoughts: Deliberate problem solving with large language models , author=. arXiv preprint arXiv:2305.10601 , year=

  69. [69]

    arXiv preprint arXiv:2108.07258 , year=

    On the opportunities and risks of foundation models , author=. arXiv preprint arXiv:2108.07258 , year=

  70. [70]

    2022 , eprint=

    Language Model Cascades , author=. 2022 , eprint=

  71. [71]

    2023 , eprint=

    Faithful Chain-of-Thought Reasoning , author=. 2023 , eprint=

  72. [72]

    2023 , eprint=

    Toolformer: Language Models Can Teach Themselves to Use Tools , author=. 2023 , eprint=

  73. [73]

    2022 , eprint=

    Atlas: Few-shot Learning with Retrieval Augmented Language Models , author=. 2022 , eprint=

  74. [74]

    A fine-grained comparison of pragmatic language understanding in humans and language models

    Hu, Jennifer and Floyd, Sammy and Jouravlev, Olessia and Fedorenko, Evelina and Gibson, Edward. A fine-grained comparison of pragmatic language understanding in humans and language models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.230

  75. [75]

    Generated Knowledge Prompting for Commonsense Reasoning

    Liu, Jiacheng and Liu, Alisa and Lu, Ximing and Welleck, Sean and West, Peter and Le Bras, Ronan and Choi, Yejin and Hajishirzi, Hannaneh. Generated Knowledge Prompting for Commonsense Reasoning. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. doi:10.18653/v1/2022.acl-long.225

  76. [76]

    2023 , eprint=

    Invalid Logic, Equivalent Gains: The Bizarreness of Reasoning in Language Model Prompting , author=. 2023 , eprint=

  77. [77]

    , author=

    The language of generalization. , author=. Psychological review , volume=. 2019 , publisher=

  78. [78]

    The handbook of pragmatics , pages=

    Implicature , author=. The handbook of pragmatics , pages=. 2006 , publisher=

  79. [79]

    , author=

    Animal, dog, or dalmatian? Level of abstraction in nominal referring expressions. , author=. CogSci , year=

  80. [80]

    overinformative

    When redundancy is useful: A Bayesian approach to “overinformative” referring expressions. , author=. Psychological Review , volume=. 2020 , publisher=

Showing first 80 references.

This paper was first reviewed by deepseek-v4-flash on August 1, 2026.