FocusLLaVA compresses visual tokens to 39% using a vision-guided region sampler plus a text-guided attention sampler, beating its LLaVA-NeXT baseline on 9 of 10 benchmarks.
Introducing our multimodal models, 2023
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression
FocusLLaVA compresses visual tokens to 39% using a vision-guided region sampler plus a text-guided attention sampler, beating its LLaVA-NeXT baseline on 9 of 10 benchmarks.