The paper finds the LLM is mostly faithful given good captions, the CLIP vision encoder contributes perception errors, and the projector preserves visual information but aligns it poorly with text.
Lawrence Zitnick, and Devi Parikh
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
The paper finds the LLM is mostly faithful given good captions, the CLIP vision encoder contributes perception errors, and the projector preserves visual information but aligns it poorly with text.