On LLaVA and InstructBLIP, updating 0.003 percent of parameters at test time with rewards from a trained CLIP evaluator reduces object hallucination rates by about 15 to 17 percent in single-run AMBER results.
Chatgpt (mar 14 version),
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.MM 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Mitigating Image Captioning Hallucinations in Vision-Language Models
On LLaVA and InstructBLIP, updating 0.003 percent of parameters at test time with rewards from a trained CLIP evaluator reduces object hallucination rates by about 15 to 17 percent in single-run AMBER results.