Embedding textual instructions directly into images raises Qwen2.5-VL's POPE accuracy from 80.2 to 84.3 percent while pushing LLaVA-1.5 and InstructBLIP to near-random accuracy.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
Embedding textual instructions directly into images raises Qwen2.5-VL's POPE accuracy from 80.2 to 84.3 percent while pushing LLaVA-1.5 and InstructBLIP to near-random accuracy.