{"id":"3fa85a0d-0df4-45ab-a6b2-85fabb79b865","arxiv_id":"2411.14137","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new 1,677-item benchmark shows that vision-language models struggle to resolve ambiguous indirect expressions even when they are given the visual context that makes the intent clear.","lead":"VAGUE is a new benchmark that tests whether AI vision-language models can use an image to figure out what a speaker really means when they make an indirect, ambiguous remark. The authors find that today's models improve when given images but still lag far behind human performance, and their main error is taking words literally instead of reasoning about the scene.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed human–model gap and the 'perceive but not reason' conclusion rest on single-annotator labels and unvalidated SU distractors; a multi-annotator replication on the 400-item subset would settle whether either leg holds.","rationale":"The reader's weakest assumption was the single-annotator human upper bound. I agree that this is the most fragile quantitative anchor, but I would sharpen the concern: the same annotation gap also undermines the failure-mode analysis, which is what converts 'models score lower' into 'models perceive but do not reason.' The SU category is generated and filtered jointly with the indirect expression, so without independent evidence that SU options are unambiguous errors, the dominance of SU errors could be a distractor-difficulty artifact rather than a reasoning deficit. The proposed multi-annotator study directly tests both the gold-label reliability and the SU-distractor validity. If the concern lands, the paper would need to add such validation before its central interpretation is accepted; if it does not land, the current conditional verdict is appropriate. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":29998,"tokens_out":4777,"duration_ms":50996,"concrete_test":"Recruit at least five independent annotators, blinded to the gold labels, to answer the same 400-item subset and also rate each of the four options on a 1–5 acceptability scale given the image and utterance. Compute (i) majority-vote human accuracy against the gold labels, (ii) pairwise inter-annotator agreement (e.g., Krippendorff's alpha), and (iii) the proportion of SU distractors rated acceptable (score ≥ 3). If alpha is below about 0.6, or if SU acceptability is comparable to the correct option's acceptability, the paper's central 'perceive but not reason' conclusion is not supported. If majority accuracy stays near 94% and SU acceptability is low, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Sec. 5.2, Table 2) is that current models 'perceive images but do not effectively reason with them,' supported by a 94% vs. ~72% human–model gap and by Superficial Understanding (SU) being the dominant error type. Both legs depend on the gold labels being the unique, unambiguous intended interpretation and on the three distractor types being genuinely wrong and comparably plausible. The only validation is a single student researcher annotating the 400-item subset (Appendix E), with no inter-annotator agreement and no independent check that SU options are actually incorrect rather than merely less preferred. Because SU options are generated by GPT-4o in the same prompt as the indirect expression (Fig. J10) and were selected together during human filtering (Sec. 4.2.2), the SU category could be systematically more tempting to models without being semantically invalid. If a broader rater pool frequently disagrees with the gold label, or rates SU distractors as acceptable interpretations, then both the size of the human–model gap and the claim that SU errors indicate a reasoning deficit rather than distractor difficulty would be unsupported. This is the load-bearing point: the benchmark's headline interpretation is only as strong as its label and distractor validation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces VAGUE, a 1,677-item benchmark for multimodal intention disambiguation (MID). Each item consists of an image, an ambiguous indirect utterance, and four multiple-choice interpretations, with distractors categorized as Fake Scene Understanding (FS), Superficial Understanding (SU), and Nonexistent Entity (NE). The dataset is built from VCR and Ego4D images with GPT-4o-generated text that is then human-rated and filtered. The authors evaluate LMs, Socratic Models, and VLMs, reporting that visual cues improve accuracy but that all models fall well below a human accuracy of 94% on a 400-item subset. Failure analysis indicates that SU errors are the most common error type, leading the authors to conclude that current models perceive image content but do not effectively reason with it. The paper also reports Chain-of-Thought experiments on proprietary and open models.","tokens_in":30267,"tokens_out":5103,"duration_ms":52667,"significance":"If the benchmark's validity assumptions hold, VAGUE is a useful resource for evaluating multimodal pragmatic reasoning and theory-of-mind-adjacent inference, and the explicit distractor taxonomy provides a diagnosable error analysis that goes beyond simple accuracy reporting. The paper evaluates a broad set of open and proprietary models, reports both multiple-choice and free-form results, and releases code and data, which are concrete strengths. However, the headline quantitative claims rest on two load-bearing validation legs: a human performance estimate from a single annotator, and the semantic validity of the generated distractors, particularly the SU category. Both legs are currently under-supported, and the internal inconsistency in the human row of Table 2 compounds the concern. The core idea is valuable, but the evidence as presented does not yet fully establish the strength of the stated human-model gap or the 'perceive but not reason' interpretation.","major_comments":[{"comment":"The human upper bound of 94% is based on a single student researcher annotating a 400-item subset, with no inter-annotator agreement and no independent second rater. Because the headline result is the large human-model gap, this is load-bearing: a broader rater pool could yield different answers and change the gap. In addition, the human row in Table 2 is internally inconsistent: the reported Correct count is 374, the incorrect counts sum to 12+4+8=24, giving 398 items rather than 400, and 374/400 is 93.5%, not 94.0%. The authors should report the exact denominator, reconcile the counts, and provide multi-annotator agreement statistics on at least a representative subsample.","section":"Sec. 5.3, Appendix E, Table 2"},{"comment":"The central failure-mode claim that SU is the dominant error and indicates a reasoning deficit presupposes that SU distractors are semantically invalid and roughly as plausible as the correct answer. However, SU options are generated by GPT-4o in the same prompt as the indirect expression (Fig. J10) and are selected together with that expression during human filtering (Sec. 4.2.2). No human validation is reported for whether SU options are actually wrong or merely less preferred, and the same holds for FS and NE distractors. The authors should run an annotator study that rates each distractor as valid/invalid and compares perceived plausibility against the correct option; without this, the SU-error analysis cannot distinguish a reasoning deficit from a distractor-difficulty artifact.","section":"Sec. 4.2.2, Sec. 4.2.3, Fig. J10, Sec. 5.2"},{"comment":"The claim that 'Superficial Understanding is the most common error type' is made from raw error counts without comparison to chance or to option-level baselines. Because there are three incorrect options, a random chooser would select each incorrect type one-third of the time conditional on being wrong; some models' SU rates are close to or below that baseline in the full tables, e.g., Qwen2.5-VL-72B VLM on VCR has 159 SU errors out of 295 total errors (53.9%), but this is not compared with a chance baseline or with the relative prevalence of each distractor type in the dataset. The authors should report per-option choice rates, chance-adjusted error proportions, and, ideally, position- and length-controlled analyses to support the failure-mode interpretation.","section":"Sec. 5.2, Fig. 4"},{"comment":"The OCR validation for person indicator tags is not sufficient to support the claim that the task only requires 'basic OCR' that models handle reliably. The test uses only three COCO images, all with the same correct answer (red shirts), and does not use any VAGUE images, person tags in varied fonts/overlays, or distractor questions. If person indicators are sometimes unreadable or ambiguous in the actual benchmark, model errors could be attributed to grounding failures rather than intent reasoning. A small but representative OCR evaluation on the actual VAGUE images, or an explicit analysis of how many errors involve the wrong person, would address this concern.","section":"Appendix B.4, Sec. 4.1.3"}],"minor_comments":[{"comment":"InternVL-3 (38B) SM accuracy on VAGUE-VCR is reported as 47.2 in Table 1 but 47.6 in Table J5; the numbers should be reconciled.","section":"Table 1 vs Appendix J5"},{"comment":"The appendix heading reads 'Free-From Answering' and should be 'Free-Form Answering'.","section":"Appendix D title"},{"comment":"Reference [5] lists the author name as 'Zhawnen Chen'; this appears to be a typo and should be corrected.","section":"References"},{"comment":"The dataset includes an 'ordering' field for MCQ options, but the paper does not state whether option order was randomized per model or per human participant; position bias could affect the failure-mode analysis, so the protocol should be described.","section":"Fig. J25 / dataset structure"},{"comment":"The increase values in Table 1 are reported as differences from LM accuracy, but no confidence intervals or significance tests are provided; given that some differences are small (e.g., Ovis2 VLM vs SM), a statement about run-to-run variability or deterministic decoding would help.","section":"Sec. 5.1"}],"recommendation":"major_revision","confidential_remarks":"The benchmark is well-motivated and the release of data and code is a practical contribution. My main concern is that the two load-bearing validation pillars — the human upper bound and the distractor validity — are currently supported by very thin evidence. Both are fixable within the manuscript's scope through additional human annotation and a clearer reporting of the human evaluation protocol. I would also encourage the authors to add chance baselines for the error-type analysis and to correct the Table 2 arithmetic before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"VAGUE is a genuinely new benchmark, and the core finding — that VLMs underuse visual context when inferring speaker intent, with literal/superficial readings the dominant error — is plausible and worth testing. The task formulation (Multimodal Intention Disambiguation), the distractor taxonomy (FS/SU/NE), and the LM/SM/VLM controlled comparison are the actual contributions. The construction is unusually careful: prompts are in the appendix, human rating criteria are explicit, and the authors acknowledge the obvious failure modes of their own pipeline. The breadth of models (twelve) and the consistent ordering — more visual cues help, up to a point — give the main table real weight.\n\nThe soft spots are real but not fatal. The human baseline is one annotator on 400 items. That is thin for a benchmark whose headline number is a human–model gap. Even if a broader rater pool drops human accuracy by ten points, the gap with the best open models would survive, but the exact 94% figure and the SU-error interpretation would need re-validation. The SU distractors are generated together with the indirect expressions and filtered by the same raters; there is no independent check that they are genuinely wrong rather than merely less preferred. That matters because 'SU is the dominant error' is the evidence for the 'perceive but not reason' claim. I'd like a multi-annotator agreement study on the 400-item subset, plus a distractor-plausibility rating. The OCR sanity check (three COCO images, all models perfect) is close to useless, but it is peripheral. No error bars is a minor complaint; benchmarks often run once.\n\nThe circularity concern the skeptics raise — GPT-4o generated the data, so GPT-4o should win — does not land. GPT-4o scores 65% on VCR, below Qwen2.5-VL-72B and InternVL, so the generator is not gaming the benchmark. The limitations section is honest about cultural and annotation-tool noise.\n\nBottom line: this is a solid benchmark paper with one load-bearing soft spot. It deserves a serious referee, and the fix is clear. I'd send it out, and I'd want the revision to include a second annotator pass or at least a clear acknowledgment of the single-annotator human bound. For the field, VAGUE is a useful testbed for pragmatic reasoning in VLMs, and the failure-mode analysis gives developers something to aim at.","headline":"A genuinely new benchmark with a plausible central finding, but the human–model gap and the 'perceive but not reason' claim rest on a single annotator and unvalidated distractors — fixable, and worth refereeing.","tokens_in":30798,"tokens_out":1966,"would_cite":true,"duration_ms":20213,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces VAGUE, a benchmark claiming that vision-language models can perceive visual details yet fail to reason from them to infer a speaker's true intent.","keywords":["multimodal intention disambiguation","visual context","ambiguous expressions","benchmark","vision-language models","theory of mind","counterfactual choices","human evaluation"],"falsifier":"Have three or more independent raters answer the same 400-item subset and compute inter-annotator agreement; if human accuracy drops toward the model range or agreement is low, then the reported 20-point gap and the 'perception without reasoning' interpretation would not survive.","tokens_in":29790,"feed_emoji":"🖼️","tokens_out":4187,"duration_ms":38334,"temperature":0.7,"pith_summary":"This paper introduces VAGUE, a benchmark of 1,677 images paired with indirect, ambiguous utterances and four-choice interpretations, to test whether multimodal AI can infer a speaker's true intent from visual context alone. The central claim is that current vision-language models can use visual information but cannot reason with it: accuracy rises steadily as more visual detail is supplied, yet the best model still lands near 72 percent while humans reach 94 percent. The paper further claims that the dominant failure mode is 'superficial understanding'—models pick the literal reading of the utterance instead of the intention the image disambiguates. A sympathetic reader would take this as quantitative evidence that perceiving a scene and understanding what a person wants done in it are distinct capabilities, with the latter still largely missing.","feed_headline":"AI sees the image but misses the speaker's point","feed_subtitle":"New benchmark of 1,677 ambiguous scenes shows best models trail human intent reading by ~20 points","key_machinery":"The load-bearing mechanism is the benchmark's construction pipeline rather than a mathematical identity. Each item starts with a direct, solvable request grounded in a physical object present in the image, from which an indirect expression is generated that hides both the action and the object. Three interpretable counterfactual choices are engineered per item: a Fake Scene interpretation from an imagined different image, a Superficial Understanding from literal reading, and a Nonexistent Entity that swaps in an object absent from the scene. This taxonomy turns every wrong answer into a diagnostic, letting the paper attribute model errors to perception failures versus reasoning failures. Human filtering based on explicit criteria of relevance, solvability, consistency, and ambiguity is what makes the answers visually dependent, so the task cannot be solved from text priors alone.","core_discovery":"The paper's core discovery is a measurable gap between visual perception and multimodal reasoning. On VAGUE, each item presents an image from the speaker's viewpoint, an indirect expression (for example, 'Hey person1, spot the difference, this parking's a bit too special, isn't it?'), and four interpretations designed so that only the visual context makes one correct. The authors report that all twelve evaluated models improve when given captions and improve further with raw images, showing they do extract visual cues; nevertheless, on a filtered 400-item subset humans answer correctly 94 percent of the time while the strongest model, Qwen2.5-VL-Instruct (72B), reaches 72.3 percent. Error analysis attributes most mistakes to Superficial Understanding choices, which follow the literal wording, rather than to misreading the scene (Fake Scene) or hallucinating an object (Nonexistent Entity). The paper concludes that models perceive image content but fail to integrate it into intent inference.","pith_inferences":["If the single-annotator human ceiling of 94 percent holds up under multi-rater testing, then the VAGUE-style format could be adapted as a live probe for embodied assistants: a system that fails here would likely also mishandle indirect user requests in real scenes.","The paper's own observation that proprietary models do better with captions than with raw images suggests a testable extension: supplying models with more detailed, structured scene descriptions might shrink the gap more than any reasoning prompt alone.","Because all utterances were drafted by a single generative model and filtered by English-speaking annotators, the ambiguity types are likely skewed toward Western sarcasm and idioms; a cross-lingual version could reveal whether the perception-reasoning gap is language-dependent.","A direct probe of the paper's interpretation would be to feed models the correct answer's reasoning but with the image removed; if performance then collapses, the visual dependency claim is confirmed, whereas if it stays high, some items may be solvable from text priors."],"forward_implications":["Adding visual cues—from no image, to a short caption, to the raw image—consistently raises accuracy across nearly all evaluated models, confirming that the task is genuinely multimodal.","The roughly 20-point gap between the best model and human performance on the same 400 items implies that intent disambiguation is not yet solved by scaling or instruction tuning alone.","Because Superficial Understanding is the most frequent error, progress on VAGUE would come less from better object recognition than from deeper pragmatic reasoning over what the speaker is asking for.","The counterfactual design lets each wrong answer be classified as Fake Scene, Superficial Understanding, or Nonexistent Entity, which is how the paper identifies the dominant failure mode.","Chain-of-thought prompting helps proprietary models only when they see the raw image, suggesting that explicit reasoning can partially compensate for missing visual grounding."],"supporting_citations":[{"why":"Supplies the 1,144 staged, complex scene images from Visual Commonsense Reasoning along with their existing person annotations.","marker":"[46]"},{"why":"Supplies the 533 natural, egocentric frames from Ego4D that anchor the real-world subset of the benchmark.","marker":"[41]"},{"why":"GPT-4o generates the direct expressions, indirect expressions, and multiple-choice candidates before human filtering.","marker":"[31]"},{"why":"The Recognize Anything tagging model filters candidate images by the number of physical objects present.","marker":"[49]"},{"why":"YOLOv11 detects and labels persons in Ego4D frames to create the person indicators used in prompts.","marker":"[18]"},{"why":"Defines the Socratic Models setting, the caption-based intermediate baseline between text-only and full-image inputs.","marker":"[47]"}],"fun_headline_variants":["AI sees images but fails to infer intent, VAGUE shows","VAGUE benchmark: AI looks but doesn't reason with visuals","Multimodal AI can't read the room, new benchmark finds","Visual cues don't help AI grasp meaning, VAGUE reveals","AI sees the scene but misses the speaker's intent"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that humans solve this task at 94 percent and that the correct answers are visually dependent rests on a single student annotator scoring a filtered 400-item subset, with no second annotator measuring agreement.","fun_headline_variants_meta":{"raw":{"variants":["AI sees images but fails to infer intent, VAGUE shows","VAGUE benchmark: AI looks but doesn't reason with visuals","Multimodal AI can't read the room, new benchmark finds","Visual cues don't help AI grasp meaning, VAGUE reveals","AI sees the scene but misses the speaker's intent"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000161,"raw_usage":{"total_tokens":1230,"prompt_tokens":931,"completion_tokens":299,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":211}},"tokens_in":547,"tokens_out":299,"duration_ms":3638,"temperature":1.0,"reasoning_tokens":211,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:30:03.036730+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have three or more independent raters answer the same 400-item subset and compute inter-annotator agreement; if human accuracy drops toward the model range or agreement is low, then the reported 20-point gap and the 'perception without reasoning' interpretation would not survive.","supporting_citations":[{"cited_title":"From recognition to cognition: Visual commonsense reason- ing","cited_arxiv_id":null,"evidence_quote":"Supplies the 1,144 staged, complex scene images from Visual Commonsense Reasoning along with their existing person annotations."},{"cited_title":"Ego4d: Around the world in 3,000 hours of egocentric video","cited_arxiv_id":null,"evidence_quote":"Supplies the 533 natural, egocentric frames from Ego4D that anchor the real-world subset of the benchmark."},{"cited_title":"Bad” column while prioritizing those in the “Good","cited_arxiv_id":null,"evidence_quote":"The Recognize Anything tagging model filters candidate images by the number of physical objects present."},{"cited_title":"Yolov11: An overview of the key architectural enhancements, 2024","cited_arxiv_id":null,"evidence_quote":"YOLOv11 detects and labels persons in Ego4D frames to create the person indicators used in prompts."},{"cited_title":"Socratic models: Composing zero-shot multimodal reasoning with language","cited_arxiv_id":null,"evidence_quote":"Defines the Socratic Models setting, the caption-based intermediate baseline between text-only and full-image inputs."}],"review_version":1}