DeCafNet reduces long-video temporal grounding cost by up to 47 percent while improving accuracy, using a lightweight sidekick encoder to select salient clips for a heavy expert encoder.
Is space-time attention all you need for video understanding? InProceedings of the International Conference on Machine Learning (ICML), 2021
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos
DeCafNet reduces long-video temporal grounding cost by up to 47 percent while improving accuracy, using a lightweight sidekick encoder to select salient clips for a heavy expert encoder.