REVIEW 12 cited by
Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Object hallucination poses a significant challenge in vision-language (VL) models, often leading to the generation of nonsensical or unfaithful responses with non-existent objects. However, the absence of a general measurement for evaluating object hallucination in VL models has hindered our understanding and ability to mitigate this issue. In this work, we present NOPE (Negative Object Presence Evaluation), a novel benchmark designed to assess object hallucination in VL models through visual question answering (VQA). We propose a cost-effective and scalable approach utilizing large language models to generate 29.5k synthetic negative pronoun (NegP) data of high quality for NOPE. We extensively investigate the performance of 10 state-of-the-art VL models in discerning the non-existence of objects in visual questions, where the ground truth answers are denoted as NegP (e.g., "none"). Additionally, we evaluate their standard performance on visual questions on 9 other VQA datasets. Through our experiments, we demonstrate that no VL model is immune to the vulnerability of object hallucination, as all models achieve accuracy below 10\% on NegP. Furthermore, we uncover that lexically diverse visual questions, question types with large scopes, and scene-relevant objects capitalize the risk of object hallucination in VL models.
Forward citations
Cited by 12 Pith papers
-
C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs
C-PTQ weights quantization error by per-channel Fisher information of the task loss, improving low-bit accuracy of multimodal LLMs by small margins over existing channel-wise scaling methods.
-
The 3D Mirage: Probing and Taming 3D Hallucinations
Depth models hallucinate 3D bumps on flat illusion images when context is cropped; the paper adds a benchmark, two scores, and a LoRA fine-tune that reduces the artifact on the same dataset.
-
Evaluating Hallucination in Large Vision-Language Models based on Context-Aware Object Similarities
CAOS is a framework that uses word-embedding similarities and an ensemble of vision-language models to detect and explain object hallucinations in image captioning.
-
VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification
VASparse combines visual-aware token pruning, embedding-based visual contrastive decoding, and an attention-sink penalty to reduce visual hallucinations in LVLMs without extra training.
-
FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Models
A modular query-code-update pipeline reduces measurement hallucinations in chest X-ray reports by replacing model-generated numbers with measurements from specialized vision tools.
-
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
Cross-modal attention maps contain a signal that a small trained network can use to flag hallucinated LVLM outputs, according to tests on POPE, AMBER, and COCO captions.
-
Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions
The paper introduces a face-attribute hallucination benchmark and MIAVLM, a model that fuses multiview generated images and negative-instruction training, reporting higher balanced accuracy than zero-shot baselines.
-
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
A study of InstructBLIP and mPLUG-Owl2 finds that scene words like grass and tree co-occur with hallucinated objects, and a two-step foreground/background prompt lowers hallucination scores.
-
Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models
Expert-CFG combines entropy-based uncertainty selection with classifier-free guidance over expert-highlighted text to refine MedVLM outputs, reporting gains on VQA-RAD, SLAKE, and PathVQA.
-
HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
A progressive two-stage knowledge distillation framework (HKD4VLM) reports first-place F1 scores of 98.2% and 98.4% on multimodal hallucination and factuality detection, but its ablation lacks a directly fine-tuned baseline.
-
PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model
PAINT reduces hallucination in LLaVA-1.5 by selectively amplifying attention to ViT-defined local and summary tokens, but the gains are selected on the test set and lack error bars.
-
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.
Discussion (0). Continue with ORCID to comment.