A cross-modal knowledge transfer network performs unsupervised temporal sentence grounding by adapting appearance knowledge from Image-Noun pairs and action knowledge from Video-Verb pairs using a copy-paste refinement.
InEuropean Conference on Com- puter Vision, pages 447–463
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CV 2verdicts
UNVERDICTED 2representative citing papers
A multi-scale and cross-scale contrastive learning framework uses intra-encoder stage features and a new sampling process to link short-range and long-range video moments for temporal grounding.
citing papers explorer
-
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding
A cross-modal knowledge transfer network performs unsupervised temporal sentence grounding by adapting appearance knowledge from Image-Noun pairs and action knowledge from Video-Verb pairs using a copy-paste refinement.
-
Multi-Scale Contrastive Learning for Video Temporal Grounding
A multi-scale and cross-scale contrastive learning framework uses intra-encoder stage features and a new sampling process to link short-range and long-range video moments for temporal grounding.