Pith. sign in

REVIEW 4 major objections 6 minor 54 references

Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues

T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Test shows only Claude 3.5 infers speaker ignorance from multiple cues.

desk verdict New empirical comparison with a real differential finding, but the pragmatic-competence claim outruns the task; needs a perceptual check and a direct test of speaker-inference. read the letter →

arxiv 2502.09120 v3 pith:6BSPMTN4 submitted 2025-02-13 cs.CL

classification cs.CL
keywords ignoranceimplicaturesvision-languagemodelspragmaticinferencemodifiednumeralscontextualcuesQuestionUnderDiscussionmultimodalreasoningacceptabilityjudgment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether vision-language models (VLMs) can go beyond literal meaning and infer that a speaker lacks precise knowledge, a pragmatic inference known as an ignorance implicature. GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet rated how appropriate utterances like 'at least four' or 'more than three' were for images that showed either every box open or two boxes closed. With only a visual cue, all three models leaned almost entirely on the lexical meaning of the modifier and ignored the scene. When a question-under-discussion (QUD) cue was added, GPT-4o and Gemini favored precise, literal readings and treated each cue independently, while Claude combined the visual and linguistic cues and produced ratings closer to human pragmatic judgments. The authors interpret Claude's cue integration as a sign of emerging pragmatic competence in multimodal models.

What carries the argument

The central testbed is the ignorance implicature of modified numerals, utterances such as 'at least four' or 'more than three' that convey the speaker does not know the exact number. The experiment combines three manipulated pieces: a visual cue (all boxes open versus two boxes closed), a linguistic cue (a how-many QUD versus a polar yes/no QUD), and the modifier type (bare, superlative, or comparative). The load-bearing mechanism is the nonlinear cue-combination threshold effect, adapted from anaphora-retrieval research: a single contextual cue may not push a model past the threshold for pragmatic interpretation, but two jointly available cues can. Claude's significant situation-by-QUD interaction is read as evidence that it crossed that threshold, while GPT-4o and Gemini's lack of interaction indicates the two cues remained unbound in their processing.

What would settle it

A direct perception check, asking each model how many boxes are open, how many are closed, and how many target objects are visible, would settle whether the approximate and precise conditions are genuinely distinct for these models. Alternatively, re-running the rating experiment with the image and modifier pairings reversed, holding the QUD constant, would show whether appropriateness ratings track the actual scene content or only the wording.

Watch

Extended reading notes

Core claim

In two rating experiments, the study establishes that the three tested VLMs do not process contextual cues uniformly. When judging image-text pairs, all models initially over-weight the modifier type, producing the order superlative > comparative > bare regardless of whether the exact count was visible. Adding a QUD that demands an exact number shifts GPT-4o and Gemini toward bare numerals, a preference for precision rather than for uncertainty-licensed readings. Claude behaves differently: its ratings show a significant interaction between the visual situation and the QUD, preferring bare numerals in precise scenes and modified numerals in approximate scenes under the how-many question, matching the human data the design is based on. The paper argues that this difference reflects not just cue weighting but whether the model's internal representation binds the two cues into a unified context, and it reads Claude's response as a threshold-like, nonlinear cue-combination effect.

Load-bearing premise

The entire comparison depends on the models actually seeing the scenes correctly, registering that some boxes are closed and that four objects are visible in the open boxes, because the study never verified visual perception with control questions.

Editorial extensions

If this is right

  • If Claude's cue integration is real, evaluating VLM pragmatic competence should include multi-cue settings, because single-cue tests can miss the ability to combine context.
  • GPT-4o and Gemini's failure to bind the two cues predicts systematic difficulty in multimodal dialogue where the intended meaning depends on jointly considering what is visible and what is asked.
  • The nonlinear threshold account implies that adding more contextual cues could produce abrupt rather than gradual improvements in pragmatic inference for models that can integrate them.
  • The 1-7 acceptability paradigm used here can serve as a reusable diagnostic for tracking pragmatic competence in future vision-language models.
  • The divergence among the three models suggests pragmatic behavior is not a uniform property of current VLMs but varies by architecture and training.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable implication the paper leaves implicit: if visual perception fidelity is verified, the observed model differences might partly reflect better visual grounding in Claude rather than better pragmatic reasoning, and the two can be separated by adding a perception control condition.
  • The same two-cue design could be applied to scalar implicatures (e.g., 'some' implying 'not all') to see whether cue integration is specific to modified numerals or generalizes across pragmatic phenomena.
  • Using inconsistent cue pairings (precise scene with a polar QUD, approximate scene with a how-many QUD) would sharpen the diagnosis of whether models bind cues or merely average their independent effects.
  • Probing model confidence or attention maps during the rating task could reveal whether the integration is a gradient phenomenon or a genuine threshold crossing.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper asks whether vision-language models (GPT-4o, Gemini 1.5 Pro, Claude 3.5 Sonnet) can perform pragmatic inference, specifically ignorance implicatures of modified numerals. In Experiment 1, models rate the appropriateness of bare, superlative, and comparative numeral sentences against images that are either precise (all boxes open) or approximate (two boxes closed). In Experiment 2, a Question Under Discussion cue (how-many vs. polar) is added. The authors report that in Experiment 1 all models rely primarily on the modifier, not the visual situation; in Experiment 2, GPT and Gemini prefer bare numerals under a how-many question while Claude shows a significant Situation-by-QUD interaction, which is interpreted as evidence that Claude integrates multiple contextual cues and may be approaching human-like pragmatic reasoning. The paper includes a public repository, a controlled stimulus design adapted from Cremers et al. (2022), and an explicit limitations section.

Significance. If the central claim were fully supported, the paper would be a useful contribution to the growing literature on pragmatic reasoning in multimodal models, because it applies an established psycholinguistic paradigm to VLMs and compares three state-of-the-art systems under systematically varied cues. The strengths include the two-cue factorial design, the grounding in prior human experiments, the transparency of the materials and code, and the candid acknowledgment of the absence of a direct human comparison. However, the significance currently depends on whether the task actually measures ignorance implicature and whether the statistical analysis supports the claimed Claude-specific integration; both points need substantial revision before the empirical contribution can be evaluated at face value.

major comments (4)
  1. [§3.1, §5.1] The task measures appropriateness, not the speaker's knowledge state. The Experiment 2 prompt is “Is the following answer to the question appropriate for the given image?”, and the approximate image contains four open apple boxes and two closed boxes. Because “at least four” is true regardless of what the closed boxes contain, while “four” is true only if the closed boxes contain no apples, a model's preference for the superlative under the how-many QUD is equally compatible with truth-conditional safety under the model's own uncertainty and with ignorance implicature. The design never varies speaker access independently of the model's access, so the central conclusion in Section 6 that Claude “may be moving closer to human-like pragmatic reasoning” is not uniquely supported. I would ask for an explicit speaker-knowledge manipulation or a speaker-confidence rating, plus a conservative rewrite of the ignorance-implicature claims.
  2. [§4.2, Tables 2–7] The statistical analyses are mislabeled. Section 4.2 calls the models “mixed-effects logistic regression” and cites Jaeger (2008), but the dependent variable is an integer 1–7 rating and all fixed-effect summaries in Tables 2–7 report t-values and linear coefficient estimates, not log-odds or z-values. If lmer with a Gaussian family was used, the models should be described as linear mixed models; if an ordinal or binomial model was intended, the reported coefficients and tests do not correspond to it. Please state the model family, link function, random-effect specification, and the method used for degrees of freedom, and correct the p-values if they rely on an inappropriate approximation.
  3. [§5.2, Table 7] The paper's central positive claim rests on a single interaction. In Section 5.2 and Table 7, Claude's salient contextual interaction is Situation:QUD (estimate -0.37, p < 0.01), but the text does not report the simple effects in the four Situation-by-QUD cells, and a significant interaction alone does not establish the ordering “bare > superlative in precise” and “superlative > bare in approximate” described in the text. The narrative also says that “only” this interaction was significant, yet Table 7 additionally shows QUD:Modifier-Comparative (p < 0.001) and Situation:QUD:Modifier-Comparative (p < 0.001). Please clarify the selection of effects and report contrasts or marginal means that directly test the claimed pattern.
  4. [§3.1] The visual cue may not be perceived as intended. Because VLMs have known counting and layout limitations, which the paper itself cites in Section 3.1, a model that fails to register the two closed boxes would treat the approximate condition as equivalent to the precise condition, making any Situation:QUD interaction an artifact of the linguistic prompt rather than cross-modal cue integration. I request a perception check (e.g., “How many boxes are closed?” and “How many apples are visible?”) for each model with accuracy reported, and a robustness analysis restricted to trials in which the model correctly identifies the closed boxes.
minor comments (6)
  1. [§4.2] The sentence “However, the similar pattern in the precise situation was unexpected” is confusing because the preceding sentence says the results aligned with Cremers et al.; please rewrite to state the human baseline and the observed deviation explicitly.
  2. [Abstract] The phrase “when only visual cues were provided” is inaccurate because the modifier text is always present; the intended contrast is the absence of a QUD cue, not the absence of linguistic input.
  3. [Tables 5–7] The header “Stdt” appears to be a typo for “Std”, and some entries are run together (e.g., “-3.67<0.001”); please format all tables consistently.
  4. [§4.2] Values such as “p < 0.01, 0.05” should be split into separate p-values for each coefficient, and significance stars or a separate column should be used to avoid ambiguity.
  5. [§6] The discussion of Parker's cue-combination scheme is speculative; the current data do not include a formal test of linear vs. nonlinear combination or a threshold effect, so this part should be labeled explicitly as a hypothesis for future work.
  6. [Limitations] The Limitations section already acknowledges that there is no direct human comparison and that model behavior may reflect statistical alignment with training data; this caveat should be carried into the Discussion and Conclusion, where the pragmatic-competence claim is currently stated without qualification.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the study reports direct behavioral measurements of VLM ratings against an external psycholinguistic benchmark, with no fitted parameter renamed as a prediction and no load-bearing self-citation.

full rationale

The paper's central claims are empirical observations of model behavior, not derivations from its own assumptions. Experiment 1 and Experiment 2 compare VLM appropriateness ratings across modifiers, situations, and QUDs, and the conclusions about Claude's cue integration are read directly from logged ratings and statistical interactions (e.g., the significant Situation:QUD interaction in Table 7). No parameter is fitted to a subset of the data and then used to predict a closely related quantity; the ratings are the data. The experimental design and prompt are adapted from external prior work (Cremers et al., 2022; Westera and Brasoveanu, 2014), and the interpretive reference to Parker (2019) is an analogy offered in the discussion, not an input to the analysis. The only self-citation is Cho and Kim (2024) in the related-work paragraph, which is not load-bearing for the present experiments. Concerns about construct validity, such as whether the appropriateness task measures speaker ignorance or merely logical safety, and about whether the models actually perceive the closed boxes, are substantive correctness risks but are not circularity: the paper does not define its target in terms of its measurements or justify its conclusions by citing its own prior conclusions. The limitation statement explicitly acknowledges that direct human comparison is absent, which further supports the view that the observed pattern is an empirical finding rather than a constructed equivalence.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim rests on domain assumptions about the validity of the experimental task and the perceptual fidelity of the models, not on fitted parameters or invented entities.

assumptions (3)
  • domain assumption The appropriateness rating on a 1-7 scale is a valid measure of ignorance implicature strength for VLMs.
    The paper assumes that the judgment task used for human participants transfers to models, based on Cremers et al. (2022).
  • domain assumption The visual and QUD manipulations are effective operationalizations of contextual precision and informativeness.
    The design relies on prior human evidence that open/closed boxes and howmany/polar questions modulate ignorance inferences.
  • domain assumption The VLMs correctly perceive the images, including the number of open boxes and the presence of closed boxes.
    No perceptual control was run; the authors rely on their stimulus design to be legible to VLMs, despite known counting limitations they cite.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues." pith.science (2026). https://pith.science/paper/6BSPMTN4

@misc{pith2026250209120,
  author       = {Pith},
  title        = {Pith review of: Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6BSPMTN4}},
  note         = {Machine review of arXiv:2502.09120}
}
read the original abstract

This study investigates whether vision-language models (VLMs) can perform pragmatic inference, focusing on ignorance implicatures, utterances that imply the speaker's lack of precise knowledge. To test this, we systematically manipulated contextual cues: the visually depicted situation (visual cue) and QUD-based linguistic prompts (linguistic cue). When only visual cues were provided, three state-of-the-art VLMs (GPT-4o, Gemini 1.5 Pro, and Claude 3.5 sonnet) produced interpretations largely based on the lexical meaning of the modified numerals. When linguistic cues were added to enhance contextual informativeness, Claude exhibited more human-like inference by integrating both types of contextual cues. In contrast, GPT and Gemini favored precise, literal interpretations. Although the influence of contextual cues increased, they treated each contextual cue independently and aligned them with semantic features rather than engaging in context-driven reasoning. These findings suggest that although the models differ in how they handle contextual cues, Claude's ability to combine multiple cues may signal emerging pragmatic competence in multimodal models.

Figures

Figures reproduced from arXiv: 2502.09120 by the authors.

Figure 1
Figure 1. Overview of the experimental procedure 3.2 Models and Procedure As VLMs for the experiment, we used GPT-4o (OpenAI, 2024), Gemini 1.5 Pro (Team et al., 2024), and Claude 3.5 sonnet (Anthropic, 2024). These models were selected due to their ability to process both image and text inputs simultaneously. They not only allow for the matching of images with text to determine their relationship but also provide the functio… view at source ↗
Figure 2
Figure 2. Result of Experiment1 — Mean scores for the [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 4
Figure 4. Modeling a threshold effect via linear and [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figures from the paper (1 more)
Figure 3
Figure 3. Figure 3: Result of Experiment2 — Mean scores for the [PITH_FULL_IMAGE:figures/full_fig_p007_3.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

54 extracted references · 34 canonical work pages

  1. [1]

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716--23736

  2. [2]

    Anthropic. 2024. https://www.anthropic.com/news/claude-3-5-sonnet Claude 3.5 Sonnet . Anthropic

  3. [3]

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 2425--2433

  4. [4]

    R Harald Baayen. 2008. Analyzing linguistic data: A practical introduction to statistics using R. Cambridge university press

  5. [5]

    R Harald Baayen, Douglas J Davidson, and Douglas M Bates. 2008. Mixed-effects modeling with crossed random effects for subjects and items. Journal of memory and language, 59(4):390--412

  6. [6]

    Douglas Bates, Martin Maechler, Ben Bolker, Steven Walker, Rune Haubo Bojesen Christensen, Henrik Singmann, Bin Dai, Gabor Grothendieck, Peter Green, and M Ben Bolker. 2015. Package ‘lme4’. convergence, 12(1):2

  7. [7]

    Daniel B \"u ring. 2008. The least at least can do. In West Coast Conference on Formal Linguistics (WCCFL), volume 26, pages 114--120. Citeseer

  8. [8]

    Francesca Capuano and Barbara Kaup. 2024. Pragmatic reasoning in gpt models: Replication of a subtle negation effect. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 46

Show all 54 references
  1. [9]

    Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597--1607. PmLR

  2. [10]

    Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. 2022. Pali: A jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794

  3. [11]

    Ye-eun Cho and Seong mook Kim. 2024. Pragmatic inference of scalar implicature by llms. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), pages 10--20

  4. [12]

    Alex Clark et al. 2015. Pillow (pil fork) documentation. readthedocs

  5. [13]

    Herbert H Clark. 1996. Using language. Cambridge University Press

  6. [14]

    Elizabeth Coppock and Thomas Brochhagen. 2013 a . Diagnosing truth, interactive sincerity, and depictive sincerity. In Semantics and Linguistic Theory, pages 358--375

  7. [15]

    Elizabeth Coppock and Thomas Brochhagen. 2013 b . Raising and resolving issues with scalar modifiers. Semantics and Pragmatics, 6:3--1

  8. [16]

    Alexandre Cremers, Liz Coppock, Jakub Dotla c il, and Floris Roelofsen. 2022. Ignorance implicatures of modified numerals. Linguistics and Philosophy, 45(3):683--740

  9. [17]

    Chris Cummins. 2013. Modelling implicatures from modified numerals. Lingua, 132:103--114

  10. [18]

    Chris Cummins and Napoleon Katsos. 2010. Comparative and superlative quantifiers: Pragmatic effects of comparison type. Journal of Semantics, 27(3):271--305

  11. [19]

    Chris Cummins, Uli Sauerland, and Stephanie Solt. 2012. Granularity and scalar implicature in numerical expressions. Linguistics and Philosophy, 35:135--169

  12. [20]

    Bart Geurts, Napoleon Katsos, Chris Cummins, Jonas Moons, and Leo Noordman. 2010. Scalar quantifiers: Logic, acquisition, and processing. Language and cognitive processes, 25(1):130--148

  13. [21]

    Bart Geurts and Rick Nouwen. 2007. 'at least'et al.: the semantics of scalar modifiers. Language, pages 533--559

  14. [22]

    H Paul Grice. 1975. Logic and Conversation. Brill

  15. [23]

    Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9729--9738

  16. [24]

    Jennifer Hu, Roger Levy, Judith Degen, and Sebastian Schuster. 2023. Expectations over unspoken alternatives predict pragmatic inferences. Transactions of the Association for Computational Linguistics, 11:885--901

  17. [25]

    Jennifer Hu, Roger Levy, and Sebastian Schuster. 2022. Predicting scalar diversity with context-driven uncertainty over alternatives. In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pages 68--74

  18. [26]

    T Florian Jaeger. 2008. Categorical data analysis: Away from anovas (transformation or not) and towards logit mixed models. Journal of memory and language, 59(4):434--446

  19. [27]

    T Florian Jaeger, Emily M Bender, and Jennifer E Arnold. 2011. Corpus-based research on language production: Information density and reducible subject relatives. Language from a cognitive perspective: grammar, usage and processing. Studies in honor of Tom Wasow, pages 161--198

  20. [28]

    Adam Kendon. 2004. Gesture: Visible action as utterance. Cambridge University Press

  21. [29]

    Julia Kruk, Jonah Lubin, Karan Sikka, Xiao Lin, Dan Jurafsky, and Ajay Divakaran. 2019. Integrating text and image: Determining multimodal document intent in instagram posts. arXiv preprint arXiv:1904.09073

  22. [30]

    Stephen C Levinson. 2000. Presumptive meanings: The theory of generalized conversational implicature. MIT press

  23. [31]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730--19742. PMLR

  24. [32]

    Benjamin Lipkin, Lionel Wong, Gabriel Grand, and Joshua B Tenenbaum. 2023. Evaluating statistical language models as pragmatic reasoners. arXiv preprint arXiv:2305.01020

  25. [33]

    Weizhe Liu, Mathieu Salzmann, and Pascal Fua. 2019. Context-aware crowd counting. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5099--5108

  26. [34]

    Jean-Claude Martin, Patrizia Paggio, P Kuenlein, Rainer Stiefelhagen, and Fabio Pianesi. 2007. Multimodal corpora for modelling human multimodal behaviour. Special issue of the International Journal of Language Resources and Evaluation, 41(3-4)

  27. [35]

    Clemens Mayr and Marie-Christine Meyer. 2014. More than at least. In Slides presented at the Two days at least workshop, Utrecht

  28. [36]

    David McNeill. 2008. Gesture and thought. In Gesture and thought. University of Chicago press

  29. [37]

    Rick Nouwen. 2010. Two kinds of modified numerals. Semantics and Pragmatics, 3:3--1

  30. [38]

    OpenAI. 2024. https://openai.com/index/hello-gpt-4o Hello GPT-4o . OpenAI

  31. [39]

    Roni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada, Inbar Mosseri, Michal Irani, and Tali Dekel. 2023. Teaching clip to count to ten. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3170--3180

  32. [40]

    Dan Parker. 2019. Cue combinatorics in memory retrieval for anaphora. Cognitive science, 43(3):e12715

  33. [41]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...

  34. [42]

    Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri, and Anh Totti Nguyen. 2024. Vision language models are blind. In Proceedings of the Asian Conference on Computer Vision, pages 18--34

  35. [43]

    Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. In International conference on machine learning, pages 8821--8831. Pmlr

  36. [44]

    Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28

  37. [45]

    Karan Sikka, Lucas Van Bramer, and Ajay Divakaran. 2019. Deep unified multimodal embeddings for understanding both content and users in social media networks. arXiv preprint arXiv:1905.07075

  38. [46]

    Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530

  39. [47]

    R Core Team. 2023. https://www.R-project.org/ R: A Language and Environment for Statistical Computing . R Foundation for Statistical Computing, Vienna, Austria

  40. [48]

    Polina Tsvilodub, Paul Marty, Sonia Ramotowska, Jacopo Romoli, and Michael Franke. 2024. Experimental pragmatics with machines: Testing llm predictions for the inferences of plain and embedded disjunctions. arXiv preprint arXiv:2405.05776

  41. [49]

    Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015. Show and tell: A neural image caption generator. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3156--3164

  42. [50]

    Matthijs Westera and Adrian Brasoveanu. 2014. Ignorance in context: The interaction of modified numerals and quds. In Semantics and Linguistic Theory, pages 414--431

  43. [51]

    Deirdre Wilson and Dan Sperber. 1995. Relevance theory. Oxford:Blackwell

  44. [52]

    Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. 2022. Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917

  45. [53]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  46. [54]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.