Sali4Vid improves dense video captioning by reweighting video features with timestamp-derived sigmoid importance and adaptively retrieving captions per semantic segment, achieving new SOTA on YouCook2 and ViTT.
Dense-Captioning Events in Videos: SYSU Submission to ActivityNet Challenge 2020
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
abstract
This technical report presents a brief description of our submission to the dense video captioning task of ActivityNet Challenge 2020. Our approach follows a two-stage pipeline: first, we extract a set of temporal event proposals; then we propose a multi-event captioning model to capture the event-level temporal relationships and effectively fuse the multi-modal information. Our approach achieves a 9.28 METEOR score on the test set.
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
Sali4Vid improves dense video captioning by reweighting video features with timestamp-derived sigmoid importance and adaptively retrieving captions per semantic segment, achieving new SOTA on YouCook2 and ViTT.