Pith. sign in

Sharegpt4v: Improving large multi-modal models with better captions

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

citation-role summary

dataset 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

dataset 1

polarities

use dataset 1

representative citing papers

CI-VID: A Coherent Interleaved Text-Video Dataset

cs.CV · 2025-07-02 · conditional · novelty 6.0

CI-VID provides 341,550 interleaved text-video sequences with individual and transition captions and shows initial evidence that fine-tuning on them improves coherent multi-scene video generation.

citing papers explorer

Showing 1 of 1 citing paper.

  • CI-VID: A Coherent Interleaved Text-Video Dataset cs.CV · 2025-07-02 · conditional · none · ref 6

    CI-VID provides 341,550 interleaved text-video sequences with individual and transition captions and shows initial evidence that fine-tuning on them improves coherent multi-scene video generation.