Rationale-augmented instruction tuning plus self-critic DPO reduces visual hallucination and improves multimodal reasoning in LVLMs across several benchmarks.
An aug- mented benchmark dataset for geometric question answer- ing through dual parallel text encoding
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning
Rationale-augmented instruction tuning plus self-critic DPO reduces visual hallucination and improves multimodal reasoning in LVLMs across several benchmarks.