CAPE-T2V fine-tunes a prompt enhancer on captioner-generated targets, then uses that same enhancer to write both the video model's fine-tuning captions and the inference-time prompt rewrites, reducing the distribution gap between training and inference conditioning.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation
CAPE-T2V fine-tunes a prompt enhancer on captioner-generated targets, then uses that same enhancer to write both the video model's fine-tuning captions and the inference-time prompt rewrites, reducing the distribution gap between training and inference conditioning.