A small video model with a group resampler reduces video input to a few hundred tokens and beats several 7B models on common video benchmarks.
On attention redundancy: A comprehensive study
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler
A small video model with a group resampler reduces video input to a few hundred tokens and beats several 7B models on common video benchmarks.