A complementary masking approach trains a mask generator with positive and negative masked captioning losses to implicitly align event locations with captions under weak supervision.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
A complementary masking approach trains a mask generator with positive and negative masked captioning losses to implicitly align event locations with captions under weak supervision.