CI-VID provides 341,550 interleaved text-video sequences with individual and transition captions and shows initial evidence that fine-tuning on them improves coherent multi-scene video generation.
Video generation models as world simulators, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CI-VID: A Coherent Interleaved Text-Video Dataset
CI-VID provides 341,550 interleaved text-video sequences with individual and transition captions and shows initial evidence that fine-tuning on them improves coherent multi-scene video generation.