Simple cluster-based averaging of visual tokens is competitive with, and sometimes better than, attention-based token selection on LLaVA and VILA benchmarks, though the advantage over prior methods is mixed.
Llava-onevision: Easy visual task transfer
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Token Sequence Compression for Efficient Multimodal Computing
Simple cluster-based averaging of visual tokens is competitive with, and sometimes better than, attention-based token selection on LLaVA and VILA benchmarks, though the advantage over prior methods is mixed.