Pith. sign in

REVIEW 12 cited by

Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.05338 v2 pith:AG6XBWED submitted 2023-10-09 cs.CV cs.CL

classification cs.CVcs.CL
keywords modelsobjecthallucinationvisualnegativenegpnopeobjects
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Object hallucination poses a significant challenge in vision-language (VL) models, often leading to the generation of nonsensical or unfaithful responses with non-existent objects. However, the absence of a general measurement for evaluating object hallucination in VL models has hindered our understanding and ability to mitigate this issue. In this work, we present NOPE (Negative Object Presence Evaluation), a novel benchmark designed to assess object hallucination in VL models through visual question answering (VQA). We propose a cost-effective and scalable approach utilizing large language models to generate 29.5k synthetic negative pronoun (NegP) data of high quality for NOPE. We extensively investigate the performance of 10 state-of-the-art VL models in discerning the non-existence of objects in visual questions, where the ground truth answers are denoted as NegP (e.g., "none"). Additionally, we evaluate their standard performance on visual questions on 9 other VQA datasets. Through our experiments, we demonstrate that no VL model is immune to the vulnerability of object hallucination, as all models achieve accuracy below 10\% on NegP. Furthermore, we uncover that lexically diverse visual questions, question types with large scopes, and scene-relevant objects capitalize the risk of object hallucination in VL models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs

    cs.CV 2026-07 conditional novelty 6.0 of 10

    C-PTQ weights quantization error by per-channel Fisher information of the task loss, improving low-bit accuracy of multimodal LLMs by small margins over existing channel-wise scaling methods.

  2. The 3D Mirage: Probing and Taming 3D Hallucinations

    cs.CV 2025-12 reject novelty 6.0 of 10

    Depth models hallucinate 3D bumps on flat illusion images when context is cropped; the paper adds a benchmark, two scores, and a LoRA fine-tune that reduces the artifact on the same dataset.

  3. Evaluating Hallucination in Large Vision-Language Models based on Context-Aware Object Similarities

    cs.CV 2025-01 conditional novelty 6.0 of 10

    CAOS is a framework that uses word-embedding similarities and an ensemble of vision-language models to detect and explain object hallucinations in image captioning.

  4. VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification

    cs.CV 2025-01 conditional novelty 6.0 of 10

    VASparse combines visual-aware token pruning, embedding-based visual contrastive decoding, and an attention-sink penalty to reduce visual hallucinations in LVLMs without extra training.

  5. FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A modular query-code-update pipeline reduces measurement hallucinations in chest X-ray reports by replacing model-generated numbers with measurements from specialized vision tools.

  6. DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models

    cs.CV 2024-11 reject novelty 6.0 of 10

    Cross-modal attention maps contain a signal that a small trained network can use to flag hallucinated LVLM outputs, according to tests on POPE, AMBER, and COCO captions.

  7. Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions

    cs.CV 2025-01 reject novelty 5.0 of 10

    The paper introduces a face-attribute hallucination benchmark and MIAVLM, a model that fuses multiview generated images and negative-instruction training, reporting higher balanced accuracy than zero-shot baselines.

  8. Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A study of InstructBLIP and mPLUG-Owl2 finds that scene words like grass and tree co-occur with hallucinated objects, and a two-step foreground/background prompt lowers hallucination scores.

  9. Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models

    cs.CV 2025-07 reject novelty 4.0 of 10

    Expert-CFG combines entropy-based uncertainty selection with classifier-free guidance over expert-highlighted text to refine MedVLM outputs, reporting gains on VQA-RAD, SLAKE, and PathVQA.

  10. HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A progressive two-stage knowledge distillation framework (HKD4VLM) reports first-place F1 scores of 98.2% and 98.4% on multimodal hallucination and factuality detection, but its ablation lacks a directly fine-tuned baseline.

  11. PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model

    cs.CV 2025-01 reject novelty 4.0 of 10

    PAINT reduces hallucination in LLaVA-1.5 by selectively amplifying attention to ViT-defined local and summary tokens, but the gains are selected on the test set and lack error bars.

  12. Empowering Multimodal LLMs with External Tools: A Comprehensive Survey

    cs.CV 2025-08 unverdicted novelty 2.0 of 10

    A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.

Pith tools