Video MLLMs mostly ignore motion in pixel-level visual grounding; a new motion-centric benchmark shows large performance drops.
Visual instruction tuning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
Video MLLMs mostly ignore motion in pixel-level visual grounding; a new motion-centric benchmark shows large performance drops.