A judge-free self-improvement pipeline for multimodal LLMs that generates hallucinated caption pairs with a controlled decoding ratio, filters and swaps them with CLIP scores, and trains with DPO, reporting reduced hallucination on Object HalBench and a new IC dataset.
Multi-modal hal- lucination control by visual information grounding
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach
A judge-free self-improvement pipeline for multimodal LLMs that generates hallucinated caption pairs with a controlled decoding ratio, filters and swaps them with CLIP scores, and trains with DPO, reporting reduced hallucination on Object HalBench and a new IC dataset.