DiVE generates multi-view driving videos conditioned on text, boxes, road sketches, and camera poses, reporting SOTA FID/FVD/KPM on nuScenes plus a 2.62x faster accelerated variant.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DiVE: Efficient Multi-View Driving Scenes Generation Based on Video Diffusion Transformer
DiVE generates multi-view driving videos conditioned on text, boxes, road sketches, and camera poses, reporting SOTA FID/FVD/KPM on nuScenes plus a 2.62x faster accelerated variant.