Videos can be reconstructed and generated from temporally compressed, frozen semantic features, and the compressed space supports faster and better class-conditional video generation than conventional video VAE latents.
The latent perturbation is used only as a reconstruction-training augmentation; evaluation, latent-statistics estimation, and latent video generation all use clean latents
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
V-RAE: Rethinking Video Latent Spaces for Generation
Videos can be reconstructed and generated from temporally compressed, frozen semantic features, and the compressed space supports faster and better class-conditional video generation than conventional video VAE latents.