Pith. sign in

REVIEW 6 cited by

Vision-Language Models Do Not Understand Negation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.09425 v2 pith:66DQDQ2R submitted 2025-01-16 cs.CV cs.CL

classification cs.CVcs.CL
keywords negationmodelsnegatedcaptionsunderstandvision-languagevlmsapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Many practical vision-language applications require models that understand negation, e.g., when using natural language to retrieve images which contain certain objects but not others. Despite advancements in vision-language models (VLMs) through large-scale training, their ability to comprehend negation remains underexplored. This study addresses the question: how well do current VLMs understand negation? We introduce NegBench, a new benchmark designed to evaluate negation understanding across 18 task variations and $79$k examples spanning image, video, and medical datasets. The benchmark consists of two core tasks designed to evaluate negation understanding in diverse multimodal settings: Retrieval with Negation and Multiple Choice Questions with Negated Captions. Our evaluation reveals that modern VLMs struggle significantly with negation, often performing at chance level. To address these shortcomings, we explore a data-centric approach wherein we finetune CLIP models on large-scale synthetic datasets containing millions of negated captions. We show that this approach can result in a 10% increase in recall on negated queries and a 28% boost in accuracy on multiple-choice questions with negated captions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs

    cs.CV 2026-05 conditional novelty 7.0 of 10

    Medical VLMs frequently select negated options that contradict visible chest X-ray findings, achieving only ~30% accuracy on direct presence probes, but a post-hoc consistency verifier raises accuracy above 95%.

  2. Disparities In Negation Understanding Across Languages In Vision-Language Models

    cs.CL 2026-04 unverdicted novelty 7.0 of 10

    VLMs exhibit affirmation bias that varies by language, with a new multilingual benchmark showing CLIP at or below chance on non-Latin scripts, MultiCLIP most uniform, and SpaceVLM corrections effective unevenly across...

  3. Uneven Evolution of Cognition Across Generations of Generative AI Models

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    Generative AI models show strong verbal comprehension and working memory but near-floor perceptual reasoning, with abstract reasoning improving faster in language than visual formats across generations.

  4. Trace Mutation in Human-LLM Dialogue: The Transcript as Forensic and Mitigation Surface

    cs.HC 2026-03 unverdicted novelty 6.0 of 10

    Trace mutations are a class of context failures in LLM conversations consisting of utterance effacement and genitive dissociation that distort the shared record while resisting ordinary repair.

  5. Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs

    cs.CV 2025-09 conditional novelty 5.0 of 10

    Tree-based reasoning prompts consistently underperform standard zero-shot prompting for VLM image classification on GTSRB and CIFAR-10 across three models.

  6. Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution

    cs.AI 2026-07 reject novelty 4.0 of 10

    Negation is reported as a non-separable signal in standard VLM embeddings, cross-modal attention recovers up to +7% F1, and the paper claims visual negation depends on textual context.

Pith tools