WSV trains a zero-shot video captioner on synthetic video latents generated from text, then uses a prompter plus GPT-2 at inference, reaching 52.0 BLEU@4 and 95.7 CIDEr on MSVD without seeing real video during training.
Text-only training for image captioning using noise-injected CLIP,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning
WSV trains a zero-shot video captioner on synthetic video latents generated from text, then uses a prompter plus GPT-2 at inference, reaching 52.0 BLEU@4 and 95.7 CIDEr on MSVD without seeing real video during training.