CI-VID provides 341,550 interleaved text-video sequences with individual and transition captions and shows initial evidence that fine-tuning on them improves coherent multi-scene video generation.
Sharegpt4v: Improving large multi-modal models with better captions
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
dataset 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
CI-VID: A Coherent Interleaved Text-Video Dataset
CI-VID provides 341,550 interleaved text-video sequences with individual and transition captions and shows initial evidence that fine-tuning on them improves coherent multi-scene video generation.