Frame-level captions with per-frame cross-attention and parallel multi-window denoising reduce semantic confusion in long multi-scene video generation in the authors' internal evaluation.
Panda-70m: Captioning 70m videos with multiple cross-modality teachers
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
Frame-level captions with per-frame cross-attention and parallel multi-window denoising reduce semantic confusion in long multi-scene video generation in the authors' internal evaluation.