Grounding-MD is a grounded video-language pre-training framework that achieves state-of-the-art zero-shot and supervised moment detection via structured prompts and early-late cross-modal fusion.
Boundary content graph neural network for temporal action proposal generation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection
Grounding-MD is a grounded video-language pre-training framework that achieves state-of-the-art zero-shot and supervised moment detection via structured prompts and early-late cross-modal fusion.