Modeling sparse visual state updates instead of full images cuts generated visual tokens ~55.6% and improves interleaved multimodal reasoning over full-image ULMMs.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
CONDITIONAL 2representative citing papers
Attenuating FFN outputs at mid-to-late transformer layers reduces language-prior dominance and thereby mitigates object hallucinations in LVLMs while preserving efficiency.
citing papers explorer
-
DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models
Modeling sparse visual state updates instead of full images cuts generated visual tokens ~55.6% and improves interleaved multimodal reasoning over full-image ULMMs.
-
FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models
Attenuating FFN outputs at mid-to-late transformer layers reduces language-prior dominance and thereby mitigates object hallucinations in LVLMs while preserving efficiency.