At each decoding step, TTH blends the LVLM's logits for candidate object tokens with CLIP image-text similarity scores, weighted by the model's uncertainty, to suppress hallucinated objects.
arXiv preprint arXiv:2602.01047 (2026)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Test-Time Hallucination Control in Large Vision-Language Models
At each decoding step, TTH blends the LVLM's logits for candidate object tokens with CLIP image-text similarity scores, weighted by the model's uncertainty, to suppress hallucinated objects.