REVIEW 18 cited by
Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
While GPT-4V(ision) impressively models both visual and textual information simultaneously, it's hallucination behavior has not been systematically assessed. To bridge this gap, we introduce a new benchmark, namely, the Bias and Interference Challenges in Visual Language Models (Bingo). This benchmark is designed to evaluate and shed light on the two common types of hallucinations in visual language models: bias and interference. Here, bias refers to the model's tendency to hallucinate certain types of responses, possibly due to imbalance in its training data. Interference pertains to scenarios where the judgment of GPT-4V(ision) can be disrupted due to how the text prompt is phrased or how the input image is presented. We identify a notable regional bias, whereby GPT-4V(ision) is better at interpreting Western images or images with English writing compared to images from other countries or containing text in other languages. Moreover, GPT-4V(ision) is vulnerable to leading questions and is often confused when interpreting multiple images together. Popular mitigation approaches, such as self-correction and chain-of-thought reasoning, are not effective in resolving these challenges. We also identified similar biases and interference vulnerabilities with LLaVA and Bard. Our results characterize the hallucination challenges in GPT-4V(ision) and state-of-the-art visual-language models, and highlight the need for new solutions. The Bingo benchmark is available at https://github.com/gzcch/Bingo.
Forward citations
Cited by 18 Pith papers
-
Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions
Context-guided multi-turn interviews reveal more VLM hallucinations than static benchmarks, and those hallucinations increase with conversational history and false-premise questions.
-
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
Reasoning models trade visual grounding for language-based inference, and this paper measures that trade-off with a new metric and benchmark.
-
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
SPARC selectively and progressively reinforces attention to relevant image tokens during decoding, improving both precision and recall in detailed image captioning compared to baselines and prior hallucination-mitigat...
-
MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding
MolSight integrates a Molecular Topology Module and Molecular Grounding Module into VLMs to enhance molecular image understanding and claims to outperform prior models on chemical visual tasks.
-
Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models
Multimodal reasoning models hallucinate at high-entropy cognitive bifurcation points due to loss of visual semantic anchoring, and the V-STAR training paradigm with HVAR rewards and FRM reflection mitigates this by re...
-
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
SEAM measures VLM reasoning consistency across modalities using paired semantically equivalent textual and visual notations, and finds systematic vision-language imbalance.
-
From Charts to Fair Narratives: Uncovering and Mitigating Geo-Economic Biases in Chart-to-Text
Vision-language models produce more positive chart summaries for high-income countries than for middle- or low-income countries, and a simple positive prompt only partly fixes the bias.
-
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
ScaleCap is an inference-time captioning pipeline that enriches captions with heuristic questions and filters hallucinations via offline contrastive sentence rating, yielding a 450K dataset that improves LVLM pretrain...
-
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
A new benchmark with edited-image question pairs shows that multimodal reasoning models lose accuracy when visual cues change, suggesting their reasoning is often not faithfully tied to the image.
-
Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
RVCD uses YOLO detections and retrieved single-concept AI images to adjust LVLM logits at decode time, cutting CHAIR hallucination rates by roughly half versus prior contrastive decoding baselines.
-
MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization
MMedPO weights preference-optimization training samples by clinical relevance scores, combining hallucinated text answers and locally noised lesion images, and reports improved medical VQA and report generation metrics.
-
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
X3-OPD improves audio-grounded reasoning by training the audio student on its own rollouts with token-level teacher feedback, using a three-tier paired text-audio corpus.
-
MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
Grouping instruction-tuning datasets by redundancy, uniqueness, or synergy of text-image interaction improves vision-language model accuracy over single-task and unselective multi-task tuning.
-
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
Multimodal LLMs frequently accept hallucinated emotion claims on a new adversarial benchmark, with the worst failures on image, audio, and video perception rather than on textbook emotion knowledge.
-
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
A comprehensive survey that frames multimodal understanding and generation as next token prediction and proposes a five-part taxonomy.
-
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
A broad survey that organizes MLLM evaluation benchmarks into capability categories, explains benchmark construction and scoring methods, and identifies gaps in current evaluation practice.
-
Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents
A survey proposing a source-and-impact taxonomy (input, model, combined; security, privacy, ethics) for threats to LLM-based agents, with feature analysis and four case studies.
-
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
A survey that organizes VQA methods from feature extraction through MLLM reasoning, datasets, and metrics, without introducing new experimental results.
Discussion (0). Continue with ORCID to comment.