REVIEW 4 major objections 6 minor 54 references
Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues
T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Test shows only Claude 3.5 infers speaker ignorance from multiple cues.
desk verdict New empirical comparison with a real differential finding, but the pragmatic-competence claim outruns the task; needs a perceptual check and a direct test of speaker-inference. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central testbed is the ignorance implicature of modified numerals, utterances such as 'at least four' or 'more than three' that convey the speaker does not know the exact number. The experiment combines three manipulated pieces: a visual cue (all boxes open versus two boxes closed), a linguistic cue (a how-many QUD versus a polar yes/no QUD), and the modifier type (bare, superlative, or comparative). The load-bearing mechanism is the nonlinear cue-combination threshold effect, adapted from anaphora-retrieval research: a single contextual cue may not push a model past the threshold for pragmatic interpretation, but two jointly available cues can. Claude's significant situation-by-QUD interaction is read as evidence that it crossed that threshold, while GPT-4o and Gemini's lack of interaction indicates the two cues remained unbound in their processing.
What would settle it
A direct perception check, asking each model how many boxes are open, how many are closed, and how many target objects are visible, would settle whether the approximate and precise conditions are genuinely distinct for these models. Alternatively, re-running the rating experiment with the image and modifier pairings reversed, holding the QUD constant, would show whether appropriateness ratings track the actual scene content or only the wording.
Extended reading notes
Core claim
In two rating experiments, the study establishes that the three tested VLMs do not process contextual cues uniformly. When judging image-text pairs, all models initially over-weight the modifier type, producing the order superlative > comparative > bare regardless of whether the exact count was visible. Adding a QUD that demands an exact number shifts GPT-4o and Gemini toward bare numerals, a preference for precision rather than for uncertainty-licensed readings. Claude behaves differently: its ratings show a significant interaction between the visual situation and the QUD, preferring bare numerals in precise scenes and modified numerals in approximate scenes under the how-many question, matching the human data the design is based on. The paper argues that this difference reflects not just cue weighting but whether the model's internal representation binds the two cues into a unified context, and it reads Claude's response as a threshold-like, nonlinear cue-combination effect.
Load-bearing premise
The entire comparison depends on the models actually seeing the scenes correctly, registering that some boxes are closed and that four objects are visible in the open boxes, because the study never verified visual perception with control questions.
Editorial extensions
If this is right
- If Claude's cue integration is real, evaluating VLM pragmatic competence should include multi-cue settings, because single-cue tests can miss the ability to combine context.
- GPT-4o and Gemini's failure to bind the two cues predicts systematic difficulty in multimodal dialogue where the intended meaning depends on jointly considering what is visible and what is asked.
- The nonlinear threshold account implies that adding more contextual cues could produce abrupt rather than gradual improvements in pragmatic inference for models that can integrate them.
- The 1-7 acceptability paradigm used here can serve as a reusable diagnostic for tracking pragmatic competence in future vision-language models.
- The divergence among the three models suggests pragmatic behavior is not a uniform property of current VLMs but varies by architecture and training.
Reading between the lines
- A testable implication the paper leaves implicit: if visual perception fidelity is verified, the observed model differences might partly reflect better visual grounding in Claude rather than better pragmatic reasoning, and the two can be separated by adding a perception control condition.
- The same two-cue design could be applied to scalar implicatures (e.g., 'some' implying 'not all') to see whether cue integration is specific to modified numerals or generalizes across pragmatic phenomena.
- Using inconsistent cue pairings (precise scene with a polar QUD, approximate scene with a how-many QUD) would sharpen the diagnosis of whether models bind cues or merely average their independent effects.
- Probing model confidence or attention maps during the rating task could reveal whether the integration is a gradient phenomenon or a genuine threshold crossing.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether vision-language models (GPT-4o, Gemini 1.5 Pro, Claude 3.5 Sonnet) can perform pragmatic inference, specifically ignorance implicatures of modified numerals. In Experiment 1, models rate the appropriateness of bare, superlative, and comparative numeral sentences against images that are either precise (all boxes open) or approximate (two boxes closed). In Experiment 2, a Question Under Discussion cue (how-many vs. polar) is added. The authors report that in Experiment 1 all models rely primarily on the modifier, not the visual situation; in Experiment 2, GPT and Gemini prefer bare numerals under a how-many question while Claude shows a significant Situation-by-QUD interaction, which is interpreted as evidence that Claude integrates multiple contextual cues and may be approaching human-like pragmatic reasoning. The paper includes a public repository, a controlled stimulus design adapted from Cremers et al. (2022), and an explicit limitations section.
Significance. If the central claim were fully supported, the paper would be a useful contribution to the growing literature on pragmatic reasoning in multimodal models, because it applies an established psycholinguistic paradigm to VLMs and compares three state-of-the-art systems under systematically varied cues. The strengths include the two-cue factorial design, the grounding in prior human experiments, the transparency of the materials and code, and the candid acknowledgment of the absence of a direct human comparison. However, the significance currently depends on whether the task actually measures ignorance implicature and whether the statistical analysis supports the claimed Claude-specific integration; both points need substantial revision before the empirical contribution can be evaluated at face value.
major comments (4)
- [§3.1, §5.1] The task measures appropriateness, not the speaker's knowledge state. The Experiment 2 prompt is “Is the following answer to the question appropriate for the given image?”, and the approximate image contains four open apple boxes and two closed boxes. Because “at least four” is true regardless of what the closed boxes contain, while “four” is true only if the closed boxes contain no apples, a model's preference for the superlative under the how-many QUD is equally compatible with truth-conditional safety under the model's own uncertainty and with ignorance implicature. The design never varies speaker access independently of the model's access, so the central conclusion in Section 6 that Claude “may be moving closer to human-like pragmatic reasoning” is not uniquely supported. I would ask for an explicit speaker-knowledge manipulation or a speaker-confidence rating, plus a conservative rewrite of the ignorance-implicature claims.
- [§4.2, Tables 2–7] The statistical analyses are mislabeled. Section 4.2 calls the models “mixed-effects logistic regression” and cites Jaeger (2008), but the dependent variable is an integer 1–7 rating and all fixed-effect summaries in Tables 2–7 report t-values and linear coefficient estimates, not log-odds or z-values. If lmer with a Gaussian family was used, the models should be described as linear mixed models; if an ordinal or binomial model was intended, the reported coefficients and tests do not correspond to it. Please state the model family, link function, random-effect specification, and the method used for degrees of freedom, and correct the p-values if they rely on an inappropriate approximation.
- [§5.2, Table 7] The paper's central positive claim rests on a single interaction. In Section 5.2 and Table 7, Claude's salient contextual interaction is Situation:QUD (estimate -0.37, p < 0.01), but the text does not report the simple effects in the four Situation-by-QUD cells, and a significant interaction alone does not establish the ordering “bare > superlative in precise” and “superlative > bare in approximate” described in the text. The narrative also says that “only” this interaction was significant, yet Table 7 additionally shows QUD:Modifier-Comparative (p < 0.001) and Situation:QUD:Modifier-Comparative (p < 0.001). Please clarify the selection of effects and report contrasts or marginal means that directly test the claimed pattern.
- [§3.1] The visual cue may not be perceived as intended. Because VLMs have known counting and layout limitations, which the paper itself cites in Section 3.1, a model that fails to register the two closed boxes would treat the approximate condition as equivalent to the precise condition, making any Situation:QUD interaction an artifact of the linguistic prompt rather than cross-modal cue integration. I request a perception check (e.g., “How many boxes are closed?” and “How many apples are visible?”) for each model with accuracy reported, and a robustness analysis restricted to trials in which the model correctly identifies the closed boxes.
minor comments (6)
- [§4.2] The sentence “However, the similar pattern in the precise situation was unexpected” is confusing because the preceding sentence says the results aligned with Cremers et al.; please rewrite to state the human baseline and the observed deviation explicitly.
- [Abstract] The phrase “when only visual cues were provided” is inaccurate because the modifier text is always present; the intended contrast is the absence of a QUD cue, not the absence of linguistic input.
- [Tables 5–7] The header “Stdt” appears to be a typo for “Std”, and some entries are run together (e.g., “-3.67<0.001”); please format all tables consistently.
- [§4.2] Values such as “p < 0.01, 0.05” should be split into separate p-values for each coefficient, and significance stars or a separate column should be used to avoid ambiguity.
- [§6] The discussion of Parker's cue-combination scheme is speculative; the current data do not include a formal test of linear vs. nonlinear combination or a threshold effect, so this part should be labeled explicitly as a hypothesis for future work.
- [Limitations] The Limitations section already acknowledges that there is no direct human comparison and that model behavior may reflect statistical alignment with training data; this caveat should be carried into the Discussion and Conclusion, where the pragmatic-competence claim is currently stated without qualification.
Circularity Check
No circularity: the study reports direct behavioral measurements of VLM ratings against an external psycholinguistic benchmark, with no fitted parameter renamed as a prediction and no load-bearing self-citation.
full rationale
The paper's central claims are empirical observations of model behavior, not derivations from its own assumptions. Experiment 1 and Experiment 2 compare VLM appropriateness ratings across modifiers, situations, and QUDs, and the conclusions about Claude's cue integration are read directly from logged ratings and statistical interactions (e.g., the significant Situation:QUD interaction in Table 7). No parameter is fitted to a subset of the data and then used to predict a closely related quantity; the ratings are the data. The experimental design and prompt are adapted from external prior work (Cremers et al., 2022; Westera and Brasoveanu, 2014), and the interpretive reference to Parker (2019) is an analogy offered in the discussion, not an input to the analysis. The only self-citation is Cho and Kim (2024) in the related-work paragraph, which is not load-bearing for the present experiments. Concerns about construct validity, such as whether the appropriateness task measures speaker ignorance or merely logical safety, and about whether the models actually perceive the closed boxes, are substantive correctness risks but are not circularity: the paper does not define its target in terms of its measurements or justify its conclusions by citing its own prior conclusions. The limitation statement explicitly acknowledges that direct human comparison is absent, which further supports the view that the observed pattern is an empirical finding rather than a constructed equivalence.
Assumptions & free parameters
assumptions (3)
- domain assumption The appropriateness rating on a 1-7 scale is a valid measure of ignorance implicature strength for VLMs.
- domain assumption The visual and QUD manipulations are effective operationalizations of contextual precision and informativeness.
- domain assumption The VLMs correctly perceive the images, including the number of open boxes and the presence of closed boxes.
Cite this review
Pith. "Pith review of Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues." pith.science (2026). https://pith.science/paper/6BSPMTN4
@misc{pith2026250209120,
author = {Pith},
title = {Pith review of: Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues},
year = {2026},
howpublished = {\url{https://pith.science/paper/6BSPMTN4}},
note = {Machine review of arXiv:2502.09120}
}
read the original abstract
This study investigates whether vision-language models (VLMs) can perform pragmatic inference, focusing on ignorance implicatures, utterances that imply the speaker's lack of precise knowledge. To test this, we systematically manipulated contextual cues: the visually depicted situation (visual cue) and QUD-based linguistic prompts (linguistic cue). When only visual cues were provided, three state-of-the-art VLMs (GPT-4o, Gemini 1.5 Pro, and Claude 3.5 sonnet) produced interpretations largely based on the lexical meaning of the modified numerals. When linguistic cues were added to enhance contextual informativeness, Claude exhibited more human-like inference by integrating both types of contextual cues. In contrast, GPT and Gemini favored precise, literal interpretations. Although the influence of contextual cues increased, they treated each contextual cue independently and aligned them with semantic features rather than engaging in context-driven reasoning. These findings suggest that although the models differ in how they handle contextual cues, Claude's ability to combine multiple cues may signal emerging pragmatic competence in multimodal models.
Figures
Reference graph
Works this paper leans on
-
[1]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716--23736
2022
-
[2]
Anthropic. 2024. https://www.anthropic.com/news/claude-3-5-sonnet Claude 3.5 Sonnet . Anthropic
work page 2024
-
[3]
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 2425--2433
2015
-
[4]
R Harald Baayen. 2008. Analyzing linguistic data: A practical introduction to statistics using R. Cambridge university press
work page 2008
-
[5]
R Harald Baayen, Douglas J Davidson, and Douglas M Bates. 2008. Mixed-effects modeling with crossed random effects for subjects and items. Journal of memory and language, 59(4):390--412
work page 2008
-
[6]
Douglas Bates, Martin Maechler, Ben Bolker, Steven Walker, Rune Haubo Bojesen Christensen, Henrik Singmann, Bin Dai, Gabor Grothendieck, Peter Green, and M Ben Bolker. 2015. Package ‘lme4’. convergence, 12(1):2
work page 2015
-
[7]
Daniel B \"u ring. 2008. The least at least can do. In West Coast Conference on Formal Linguistics (WCCFL), volume 26, pages 114--120. Citeseer
work page 2008
-
[8]
Francesca Capuano and Barbara Kaup. 2024. Pragmatic reasoning in gpt models: Replication of a subtle negation effect. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 46
work page 2024
Show all 54 references
-
[9]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597--1607. PmLR
2020
-
[10]
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. 2022. Pali: A jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794
2022 arXiv
-
[11]
Ye-eun Cho and Seong mook Kim. 2024. Pragmatic inference of scalar implicature by llms. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), pages 10--20
2024
-
[12]
Alex Clark et al. 2015. Pillow (pil fork) documentation. readthedocs
2015
-
[13]
Herbert H Clark. 1996. Using language. Cambridge University Press
1996
-
[14]
Elizabeth Coppock and Thomas Brochhagen. 2013 a . Diagnosing truth, interactive sincerity, and depictive sincerity. In Semantics and Linguistic Theory, pages 358--375
2013
-
[15]
Elizabeth Coppock and Thomas Brochhagen. 2013 b . Raising and resolving issues with scalar modifiers. Semantics and Pragmatics, 6:3--1
2013
-
[16]
Alexandre Cremers, Liz Coppock, Jakub Dotla c il, and Floris Roelofsen. 2022. Ignorance implicatures of modified numerals. Linguistics and Philosophy, 45(3):683--740
2022
-
[17]
Chris Cummins. 2013. Modelling implicatures from modified numerals. Lingua, 132:103--114
2013
-
[18]
Chris Cummins and Napoleon Katsos. 2010. Comparative and superlative quantifiers: Pragmatic effects of comparison type. Journal of Semantics, 27(3):271--305
2010
-
[19]
Chris Cummins, Uli Sauerland, and Stephanie Solt. 2012. Granularity and scalar implicature in numerical expressions. Linguistics and Philosophy, 35:135--169
2012
-
[20]
Bart Geurts, Napoleon Katsos, Chris Cummins, Jonas Moons, and Leo Noordman. 2010. Scalar quantifiers: Logic, acquisition, and processing. Language and cognitive processes, 25(1):130--148
2010
-
[21]
Bart Geurts and Rick Nouwen. 2007. 'at least'et al.: the semantics of scalar modifiers. Language, pages 533--559
2007
-
[22]
H Paul Grice. 1975. Logic and Conversation. Brill
1975
-
[23]
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9729--9738
2020
-
[24]
Jennifer Hu, Roger Levy, Judith Degen, and Sebastian Schuster. 2023. Expectations over unspoken alternatives predict pragmatic inferences. Transactions of the Association for Computational Linguistics, 11:885--901
2023
-
[25]
Jennifer Hu, Roger Levy, and Sebastian Schuster. 2022. Predicting scalar diversity with context-driven uncertainty over alternatives. In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pages 68--74
2022
-
[26]
T Florian Jaeger. 2008. Categorical data analysis: Away from anovas (transformation or not) and towards logit mixed models. Journal of memory and language, 59(4):434--446
2008
-
[27]
T Florian Jaeger, Emily M Bender, and Jennifer E Arnold. 2011. Corpus-based research on language production: Information density and reducible subject relatives. Language from a cognitive perspective: grammar, usage and processing. Studies in honor of Tom Wasow, pages 161--198
2011
-
[28]
Adam Kendon. 2004. Gesture: Visible action as utterance. Cambridge University Press
2004
-
[29]
Julia Kruk, Jonah Lubin, Karan Sikka, Xiao Lin, Dan Jurafsky, and Ajay Divakaran. 2019. Integrating text and image: Determining multimodal document intent in instagram posts. arXiv preprint arXiv:1904.09073
2019 arXiv
-
[30]
Stephen C Levinson. 2000. Presumptive meanings: The theory of generalized conversational implicature. MIT press
2000
-
[31]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730--19742. PMLR
2023
-
[32]
Benjamin Lipkin, Lionel Wong, Gabriel Grand, and Joshua B Tenenbaum. 2023. Evaluating statistical language models as pragmatic reasoners. arXiv preprint arXiv:2305.01020
2023 arXiv
-
[33]
Weizhe Liu, Mathieu Salzmann, and Pascal Fua. 2019. Context-aware crowd counting. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5099--5108
2019
-
[34]
Jean-Claude Martin, Patrizia Paggio, P Kuenlein, Rainer Stiefelhagen, and Fabio Pianesi. 2007. Multimodal corpora for modelling human multimodal behaviour. Special issue of the International Journal of Language Resources and Evaluation, 41(3-4)
2007
-
[35]
Clemens Mayr and Marie-Christine Meyer. 2014. More than at least. In Slides presented at the Two days at least workshop, Utrecht
2014
-
[36]
David McNeill. 2008. Gesture and thought. In Gesture and thought. University of Chicago press
2008
-
[37]
Rick Nouwen. 2010. Two kinds of modified numerals. Semantics and Pragmatics, 3:3--1
2010
-
[38]
OpenAI. 2024. https://openai.com/index/hello-gpt-4o Hello GPT-4o . OpenAI
2024
-
[39]
Roni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada, Inbar Mosseri, Michal Irani, and Tali Dekel. 2023. Teaching clip to count to ten. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3170--3180
2023
-
[40]
Dan Parker. 2019. Cue combinatorics in memory retrieval for anaphora. Cognitive science, 43(3):e12715
2019
-
[41]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[42]
Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri, and Anh Totti Nguyen. 2024. Vision language models are blind. In Proceedings of the Asian Conference on Computer Vision, pages 18--34
2024
-
[43]
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. In International conference on machine learning, pages 8821--8831. Pmlr
2021
-
[44]
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28
2015
-
[45]
Karan Sikka, Lucas Van Bramer, and Ajay Divakaran. 2019. Deep unified multimodal embeddings for understanding both content and users in social media networks. arXiv preprint arXiv:1905.07075
2019 arXiv
-
[46]
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530
2024 arXiv
-
[47]
R Core Team. 2023. https://www.R-project.org/ R: A Language and Environment for Statistical Computing . R Foundation for Statistical Computing, Vienna, Austria
2023
-
[48]
Polina Tsvilodub, Paul Marty, Sonia Ramotowska, Jacopo Romoli, and Michael Franke. 2024. Experimental pragmatics with machines: Testing llm predictions for the inferences of plain and embedded disjunctions. arXiv preprint arXiv:2405.05776
2024 arXiv
-
[49]
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015. Show and tell: A neural image caption generator. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3156--3164
2015
-
[50]
Matthijs Westera and Adrian Brasoveanu. 2014. Ignorance in context: The interaction of modified numerals and quds. In Semantics and Linguistic Theory, pages 414--431
2014
-
[51]
Deirdre Wilson and Dan Sperber. 1995. Relevance theory. Oxford:Blackwell
1995
-
[52]
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. 2022. Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917
2022 arXiv
-
[53]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[54]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.