{"id":"fe5376cd-22df-4330-8d9a-26acd1cb42c5","arxiv_id":"2502.09120","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a controlled image-text task, Claude 3.5 integrated visual and question cues to make pragmatic speaker-ignorance inferences, while GPT-4o and Gemini 1.5 Pro relied more on literal meanings.","lead":"Can AI models tell when a speaker is unsure of the exact number? This paper tested three vision-language models on pictures of boxes with hidden and visible apples. One model, Claude, combined visual and question cues in a more human-like way, while GPT and Gemini leaned on literal word meanings.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The task asks about appropriateness, not the speaker's knowledge; since 'at least four' is always true in the approximate condition, Claude's cue integration may reflect truth-conditional safety rather than ignorance implicature.","rationale":"The reader's weakest assumption, perceptual fidelity, is a legitimate concern, but the existing data provide indirect evidence against it: significant situation×modifier interactions appear in Experiment 1 for all three models (e.g., Claude: p<0.001), suggesting the models do respond differently to the two images. The more load-bearing issue is construct validity: the dependent variable (appropriateness) cannot distinguish genuine pragmatic inference about the speaker's ignorance from simple truth-conditional safety under the model's own uncertainty. This matters because the paper's headline claim is about 'emerging pragmatic competence' and 'inferring speaker's ignorance.' A model could produce the exact same ratings by choosing the utterance that is guaranteed true ('at least four') versus the one that is exactly true but possibly false ('four') in the approximate condition. The design never manipulates the speaker's epistemic access independently of the model's view; the model and speaker share the same information. The limitations section acknowledges the absence of a direct human baseline but does not acknowledge this confound. My proposed asymmetric-knowledge condition would settle whether the effect is speaker-specific. If it fails, the central claim should be tempered to 'context-sensitive semantic evaluation' rather than 'pragmatic inference.' The statistical mislabeling and threshold-effect speculation are secondary; the construct-validity issue is the primary reason the verdict should remain conditional.","tokens_in":11825,"tokens_out":9777,"duration_ms":93675,"concrete_test":"Add a third 'asymmetric-knowledge' condition to Experiment 2: the same approximate image, but with text stating 'The speaker cannot see the two closed boxes, but you (the model) know they are empty.' Ask the same appropriateness question with 'There are four apples' and 'There are at least four apples.' If Claude rates 'four' as less appropriate than 'at least four' in this condition, the original effect reflects inference about the speaker's ignorance. If Claude rates 'four' as highly appropriate (matching the precise condition), the original pattern is explained by the model's own uncertainty rather than speaker-focused pragmatic reasoning.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—that Claude's integration of visual and linguistic cues signals emerging pragmatic competence—depends on the construct validity of the task as a measure of speaker ignorance. In Experiment 2, the prompt (Section 5.1) asks: 'Is the following answer to the question appropriate for the given image?' It never asks about the speaker's knowledge state. In the approximate condition, the model sees four open boxes with apples and two closed boxes; it does not know the contents of the closed boxes, and neither does the speaker. Thus 'There are four apples' is possibly false, while 'There are at least four apples' is guaranteed true. A model could prefer 'at least four' simply because it is the logically safest statement given the model's own uncertainty, without any inference about the speaker's epistemic state. The design never varies the speaker's access independently of the model's access; the model and speaker always share the same view. Therefore, the observed pattern—Claude preferring superlative over bare in the approximate/howmany condition—is equally consistent with (a) genuine ignorance implicature, (b) a rational choice under uncertainty, or (c) a lexical preference for hedged modifiers. The paper's conclusion (Section 6) that Claude 'may be moving closer to human-like pragmatic reasoning' overreaches. This is not merely a question of interpretation: the title asks whether VLMs can infer the speaker's ignorance, but the experiment does not isolate that target.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks whether vision-language models (GPT-4o, Gemini 1.5 Pro, Claude 3.5 Sonnet) can perform pragmatic inference, specifically ignorance implicatures of modified numerals. In Experiment 1, models rate the appropriateness of bare, superlative, and comparative numeral sentences against images that are either precise (all boxes open) or approximate (two boxes closed). In Experiment 2, a Question Under Discussion cue (how-many vs. polar) is added. The authors report that in Experiment 1 all models rely primarily on the modifier, not the visual situation; in Experiment 2, GPT and Gemini prefer bare numerals under a how-many question while Claude shows a significant Situation-by-QUD interaction, which is interpreted as evidence that Claude integrates multiple contextual cues and may be approaching human-like pragmatic reasoning. The paper includes a public repository, a controlled stimulus design adapted from Cremers et al. (2022), and an explicit limitations section.","tokens_in":12080,"tokens_out":6643,"duration_ms":72742,"significance":"If the central claim were fully supported, the paper would be a useful contribution to the growing literature on pragmatic reasoning in multimodal models, because it applies an established psycholinguistic paradigm to VLMs and compares three state-of-the-art systems under systematically varied cues. The strengths include the two-cue factorial design, the grounding in prior human experiments, the transparency of the materials and code, and the candid acknowledgment of the absence of a direct human comparison. However, the significance currently depends on whether the task actually measures ignorance implicature and whether the statistical analysis supports the claimed Claude-specific integration; both points need substantial revision before the empirical contribution can be evaluated at face value.","major_comments":[{"comment":"The task measures appropriateness, not the speaker's knowledge state. The Experiment 2 prompt is “Is the following answer to the question appropriate for the given image?”, and the approximate image contains four open apple boxes and two closed boxes. Because “at least four” is true regardless of what the closed boxes contain, while “four” is true only if the closed boxes contain no apples, a model's preference for the superlative under the how-many QUD is equally compatible with truth-conditional safety under the model's own uncertainty and with ignorance implicature. The design never varies speaker access independently of the model's access, so the central conclusion in Section 6 that Claude “may be moving closer to human-like pragmatic reasoning” is not uniquely supported. I would ask for an explicit speaker-knowledge manipulation or a speaker-confidence rating, plus a conservative rewrite of the ignorance-implicature claims.","section":"§3.1, §5.1"},{"comment":"The statistical analyses are mislabeled. Section 4.2 calls the models “mixed-effects logistic regression” and cites Jaeger (2008), but the dependent variable is an integer 1–7 rating and all fixed-effect summaries in Tables 2–7 report t-values and linear coefficient estimates, not log-odds or z-values. If lmer with a Gaussian family was used, the models should be described as linear mixed models; if an ordinal or binomial model was intended, the reported coefficients and tests do not correspond to it. Please state the model family, link function, random-effect specification, and the method used for degrees of freedom, and correct the p-values if they rely on an inappropriate approximation.","section":"§4.2, Tables 2–7"},{"comment":"The paper's central positive claim rests on a single interaction. In Section 5.2 and Table 7, Claude's salient contextual interaction is Situation:QUD (estimate -0.37, p < 0.01), but the text does not report the simple effects in the four Situation-by-QUD cells, and a significant interaction alone does not establish the ordering “bare > superlative in precise” and “superlative > bare in approximate” described in the text. The narrative also says that “only” this interaction was significant, yet Table 7 additionally shows QUD:Modifier-Comparative (p < 0.001) and Situation:QUD:Modifier-Comparative (p < 0.001). Please clarify the selection of effects and report contrasts or marginal means that directly test the claimed pattern.","section":"§5.2, Table 7"},{"comment":"The visual cue may not be perceived as intended. Because VLMs have known counting and layout limitations, which the paper itself cites in Section 3.1, a model that fails to register the two closed boxes would treat the approximate condition as equivalent to the precise condition, making any Situation:QUD interaction an artifact of the linguistic prompt rather than cross-modal cue integration. I request a perception check (e.g., “How many boxes are closed?” and “How many apples are visible?”) for each model with accuracy reported, and a robustness analysis restricted to trials in which the model correctly identifies the closed boxes.","section":"§3.1"}],"minor_comments":[{"comment":"The sentence “However, the similar pattern in the precise situation was unexpected” is confusing because the preceding sentence says the results aligned with Cremers et al.; please rewrite to state the human baseline and the observed deviation explicitly.","section":"§4.2"},{"comment":"The phrase “when only visual cues were provided” is inaccurate because the modifier text is always present; the intended contrast is the absence of a QUD cue, not the absence of linguistic input.","section":"Abstract"},{"comment":"The header “Stdt” appears to be a typo for “Std”, and some entries are run together (e.g., “-3.67<0.001”); please format all tables consistently.","section":"Tables 5–7"},{"comment":"Values such as “p < 0.01, 0.05” should be split into separate p-values for each coefficient, and significance stars or a separate column should be used to avoid ambiguity.","section":"§4.2"},{"comment":"The discussion of Parker's cue-combination scheme is speculative; the current data do not include a formal test of linear vs. nonlinear combination or a threshold effect, so this part should be labeled explicitly as a hypothesis for future work.","section":"§6"},{"comment":"The Limitations section already acknowledges that there is no direct human comparison and that model behavior may reflect statistical alignment with training data; this caveat should be carried into the Discussion and Conclusion, where the pragmatic-competence claim is currently stated without qualification.","section":"Limitations"}],"recommendation":"major_revision","confidential_remarks":"The central difficulty is construct validity, not execution: the dataset and design are transparent and reusable, but as currently framed the title and conclusion claim more than the task can support. I would encourage the editor to invite a revision that either adds a speaker-knowledge manipulation or reframes the paper as a study of cue integration in VLM appropriateness judgments. The perception check is a prerequisite in either case."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look: this is the first systematic comparison of three proprietary VLMs on ignorance implicatures with both visual and linguistic cue manipulations, and it reports a clear differential pattern: Claude integrates the cues, GPT and Gemini do not. The empirical core is real and the materials and code are public. But the title and discussion ask whether VLMs can infer the speaker's ignorance, and the experiments never actually isolate that. The task is an appropriateness judgment, not a knowledge judgment, and in the crucial approximate condition \"at least four\" is always true. So Claude's preference for the superlative there could just be the safe answer under uncertainty, not a genuine inference about the speaker's epistemic state. That is a fixable interpretation problem, not a broken data pattern.\n\nWhat is new and good: adapting Cremers et al.'s design to VLMs is a legitimate move, the results are clearly presented, and the finding that GPT and Gemini lean on lexical semantics while Claude shifts with context is new and worth knowing. The authors also explicitly acknowledge the lack of a human baseline and the limited model pool, which is honest.\n\nThe soft spots are real but proportionate. First, the statistics are mislabeled: mixed-effects logistic regression implies a binary outcome, but the dependent variable is a 1-7 rating and the tables report t-values. This is likely a linear mixed model and should be described as such. Second, the central Claude claim rests on a single interaction (Situation:QUD, p<0.01) with no omnibus test; I would want to see that the pattern is robust across items and that the threshold-effect language is not just post hoc. Third, and more substantively, there is no perceptual check. The authors cite Paiss et al. on VLM counting limitations, yet they never verify that the models actually register the two closed boxes. Without that, the approximate condition may be visually inert.\n\nThe stress-test concern about construct validity holds up. Because the model and the speaker share the same view, the design cannot distinguish between the model reasoning about the speaker's ignorance and the model giving the logically safest answer given its own uncertainty. The conclusion that Claude \"may be moving closer to human-like pragmatic reasoning\" overreaches. The pattern is still interesting, but it needs a condition that varies the speaker's access independently, or a direct question about the speaker's knowledge, before that claim is warranted.\n\nThis paper deserves a serious referee, not a desk reject. The empirical comparison is new and will likely be cited. It would need a major revision: fix the statistical label, add a perceptual control, and either directly test or substantially soften the pragmatic-competence interpretation. I would send it out.","headline":"New empirical comparison with a real differential finding, but the pragmatic-competence claim outruns the task; needs a perceptual check and a direct test of speaker-inference.","tokens_in":12585,"tokens_out":3613,"would_cite":true,"duration_ms":37257,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Test shows only Claude 3.5 infers speaker ignorance from multiple cues.","keywords":["ignorance implicatures","vision-language models","pragmatic inference","modified numerals","contextual cues","Question Under Discussion","multimodal reasoning","acceptability judgment"],"falsifier":"A direct perception check, asking each model how many boxes are open, how many are closed, and how many target objects are visible, would settle whether the approximate and precise conditions are genuinely distinct for these models. Alternatively, re-running the rating experiment with the image and modifier pairings reversed, holding the QUD constant, would show whether appropriateness ratings track the actual scene content or only the wording.","tokens_in":11612,"feed_emoji":"🧠","tokens_out":8091,"duration_ms":69762,"temperature":0.7,"pith_summary":"The paper asks whether vision-language models (VLMs) can go beyond literal meaning and infer that a speaker lacks precise knowledge, a pragmatic inference known as an ignorance implicature. GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet rated how appropriate utterances like 'at least four' or 'more than three' were for images that showed either every box open or two boxes closed. With only a visual cue, all three models leaned almost entirely on the lexical meaning of the modifier and ignored the scene. When a question-under-discussion (QUD) cue was added, GPT-4o and Gemini favored precise, literal readings and treated each cue independently, while Claude combined the visual and linguistic cues and produced ratings closer to human pragmatic judgments. The authors interpret Claude's cue integration as a sign of emerging pragmatic competence in multimodal models.","feed_headline":"Test shows only Claude 3.5 infers speaker ignorance from multiple cues","feed_subtitle":"GPT-4o and Gemini treated the cues separately and favored literal precision; Claude integrated them.","key_machinery":"The central testbed is the ignorance implicature of modified numerals, utterances such as 'at least four' or 'more than three' that convey the speaker does not know the exact number. The experiment combines three manipulated pieces: a visual cue (all boxes open versus two boxes closed), a linguistic cue (a how-many QUD versus a polar yes/no QUD), and the modifier type (bare, superlative, or comparative). The load-bearing mechanism is the nonlinear cue-combination threshold effect, adapted from anaphora-retrieval research: a single contextual cue may not push a model past the threshold for pragmatic interpretation, but two jointly available cues can. Claude's significant situation-by-QUD interaction is read as evidence that it crossed that threshold, while GPT-4o and Gemini's lack of interaction indicates the two cues remained unbound in their processing.","core_discovery":"In two rating experiments, the study establishes that the three tested VLMs do not process contextual cues uniformly. When judging image-text pairs, all models initially over-weight the modifier type, producing the order superlative > comparative > bare regardless of whether the exact count was visible. Adding a QUD that demands an exact number shifts GPT-4o and Gemini toward bare numerals, a preference for precision rather than for uncertainty-licensed readings. Claude behaves differently: its ratings show a significant interaction between the visual situation and the QUD, preferring bare numerals in precise scenes and modified numerals in approximate scenes under the how-many question, matching the human data the design is based on. The paper argues that this difference reflects not just cue weighting but whether the model's internal representation binds the two cues into a unified context, and it reads Claude's response as a threshold-like, nonlinear cue-combination effect.","pith_inferences":["A testable implication the paper leaves implicit: if visual perception fidelity is verified, the observed model differences might partly reflect better visual grounding in Claude rather than better pragmatic reasoning, and the two can be separated by adding a perception control condition.","The same two-cue design could be applied to scalar implicatures (e.g., 'some' implying 'not all') to see whether cue integration is specific to modified numerals or generalizes across pragmatic phenomena.","Using inconsistent cue pairings (precise scene with a polar QUD, approximate scene with a how-many QUD) would sharpen the diagnosis of whether models bind cues or merely average their independent effects.","Probing model confidence or attention maps during the rating task could reveal whether the integration is a gradient phenomenon or a genuine threshold crossing."],"forward_implications":["If Claude's cue integration is real, evaluating VLM pragmatic competence should include multi-cue settings, because single-cue tests can miss the ability to combine context.","GPT-4o and Gemini's failure to bind the two cues predicts systematic difficulty in multimodal dialogue where the intended meaning depends on jointly considering what is visible and what is asked.","The nonlinear threshold account implies that adding more contextual cues could produce abrupt rather than gradual improvements in pragmatic inference for models that can integrate them.","The 1-7 acceptability paradigm used here can serve as a reusable diagnostic for tracking pragmatic competence in future vision-language models.","The divergence among the three models suggests pragmatic behavior is not a uniform property of current VLMs but varies by architecture and training."],"supporting_citations":[{"why":"Supplies the human acceptability-judgment paradigm, the precise/approximate visual manipulation, and the QUD conditions that the study adapts and compares against.","marker":"Cremers et al. (2022)"},{"why":"Established that how-many vs polar QUDs modulate ignorance implicatures in humans; the basis for the linguistic cue.","marker":"Westera and Brasoveanu (2014)"},{"why":"Documents VLMs' counting limitations, motivating the cap of four target objects and the structured box layout to keep visual perception feasible.","marker":"Paiss et al. (2023)"},{"why":"Provides the nonlinear cue-combination threshold account that the paper uses to interpret Claude's integration of the two contextual cues.","marker":"Parker (2019)"},{"why":"The Maxim of Quantity is the theoretical foundation for defining ignorance implicatures as pragmatic inferences.","marker":"Grice (1975)"},{"why":"Distinguishes two kinds of modified numerals and grounds the superlative > comparative > bare hierarchy in the ratings.","marker":"Nouwen (2010)"},{"why":"Argues that superlative modifiers have more complex semantics, explaining why 'at least' triggers ignorance more consistently.","marker":"Geurts and Nouwen (2007)"},{"why":"Analyzes scalar modifiers such as at least/more than and their role in raising issues of speaker ignorance.","marker":"Coppock and Brochhagen (2013b)"}],"fun_headline_variants":["Claude 3.5 reads speaker ignorance, GPT and Gemini stay literal","Only Claude combines cues to infer speaker uncertainty","Vision-language models fail at pragmatic inference, except Claude","Claude 3.5 shows emerging pragmatic edge over GPT and Gemini","Speaker ignorance: Claude gets the hint, rivals don't"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire comparison depends on the models actually seeing the scenes correctly, registering that some boxes are closed and that four objects are visible in the open boxes, because the study never verified visual perception with control questions.","fun_headline_variants_meta":{"raw":{"variants":["Claude 3.5 reads speaker ignorance, GPT and Gemini stay literal","Only Claude combines cues to infer speaker uncertainty","Vision-language models fail at pragmatic inference, except Claude","Claude 3.5 shows emerging pragmatic edge over GPT and Gemini","Speaker ignorance: Claude gets the hint, rivals don't"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000414,"raw_usage":{"total_tokens":2118,"prompt_tokens":904,"completion_tokens":1214,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":1130}},"tokens_in":520,"tokens_out":1214,"duration_ms":9385,"temperature":1.0,"reasoning_tokens":1130,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T22:33:34.035254+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct perception check, asking each model how many boxes are open, how many are closed, and how many target objects are visible, would settle whether the approximate and precise conditions are genuinely distinct for these models. Alternatively, re-running the rating experiment with the image and modifier pairings reversed, holding the QUD constant, would show whether appropriateness ratings track the actual scene content or only the wording.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the human acceptability-judgment paradigm, the precise/approximate visual manipulation, and the QUD conditions that the study adapts and compares against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Established that how-many vs polar QUDs modulate ignorance implicatures in humans; the basis for the linguistic cue."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the nonlinear cue-combination threshold account that the paper uses to interpret Claude's integration of the two contextual cues."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Maxim of Quantity is the theoretical foundation for defining ignorance implicatures as pragmatic inferences."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Distinguishes two kinds of modified numerals and grounds the superlative > comparative > bare hierarchy in the ratings."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Argues that superlative modifiers have more complex semantics, explaining why 'at least' triggers ignorance more consistently."}],"review_version":1}