Simple cluster-based averaging of visual tokens is competitive with, and sometimes better than, attention-based token selection on LLaVA and VILA benchmarks, though the advantage over prior methods is mixed.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Token Sequence Compression for Efficient Multimodal Computing
Simple cluster-based averaging of visual tokens is competitive with, and sometimes better than, attention-based token selection on LLaVA and VILA benchmarks, though the advantage over prior methods is mixed.