Pith. sign in

REVIEW 9 cited by

CV-VAE: A Compatible Video VAE for Latent Generative Video Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.20279 v2 pith:7U5M4AHJ submitted 2024-05-30 cs.CV cs.AIeess.IV

classification cs.CVcs.AIeess.IV
keywords videomodelslatentspacecompatibilitycv-vaediffusion-basedframes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Spatio-temporal compression of videos, utilizing networks such as Variational Autoencoders (VAE), plays a crucial role in OpenAI's SORA and numerous other video generative models. For instance, many LLM-like video models learn the distribution of discrete tokens derived from 3D VAEs within the VQVAE framework, while most diffusion-based video models capture the distribution of continuous latent extracted by 2D VAEs without quantization. The temporal compression is simply realized by uniform frame sampling which results in unsmooth motion between consecutive frames. Currently, there lacks of a commonly used continuous video (3D) VAE for latent diffusion-based video models in the research community. Moreover, since current diffusion-based approaches are often implemented using pre-trained text-to-image (T2I) models, directly training a video VAE without considering the compatibility with existing T2I models will result in a latent space gap between them, which will take huge computational resources for training to bridge the gap even with the T2I models as initialization. To address this issue, we propose a method for training a video VAE of latent video models, namely CV-VAE, whose latent space is compatible with that of a given image VAE, e.g., image VAE of Stable Diffusion (SD). The compatibility is achieved by the proposed novel latent space regularization, which involves formulating a regularization loss using the image VAE. Benefiting from the latent space compatibility, video models can be trained seamlessly from pre-trained T2I or video models in a truly spatio-temporally compressed latent space, rather than simply sampling video frames at equal intervals. With our CV-VAE, existing video models can generate four times more frames with minimal finetuning. Extensive experiments are conducted to demonstrate the effectiveness of the proposed video VAE.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A hierarchical motion autoencoder with a conditional diffusion decoder reconstructs 16-frame videos from latents as small as 0.07% of the input size while maintaining competitive PSNR and perceptual scores.

  2. Interspatial Attention for Efficient 4D Human Video Generation

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A new symmetric 3D-to-2D attention mechanism with relative positional encodings, plus a motion-tuned video VAE, improves controllable 4D human video generation.

  3. Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Can3Tok tokenizes scene-level 3D Gaussian splats into canonical latent tokens with normalization and saliency filtering, enabling reconstruction and text/image-to-3D generation.

  4. Large Motion Video Autoencoding with Cross-modal Video VAE

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A two-stage video autoencoder with temporal-aware spatial compression and a separate motion compressor reports state-of-the-art reconstruction quality on WebVid, Inter4K, and a large-motion test set.

  5. Mimir: Improving Video Diffusion Models for Precise Text Understanding

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Mimir fuses T5 encoder tokens with Phi-3.5 decoder-only LLM tokens using zero-conv, normalization, and four learnable stabilizer tokens, improving text-to-video semantic fidelity.

  6. WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model

    cs.CV 2024-11 conditional novelty 6.0 of 10

    WF-VAE uses multi-level Haar wavelets to route low-frequency video content around a smaller backbone, cutting compute and memory, and adds a lossless Causal Cache for block-wise inference.

  7. HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    HumanGenesis is a claimed state-of-the-art framework that couples 3D Gaussian reconstruction, LLM-based critique, pose guidance, and diffusion-based harmonization for synthetic human video generation; unverified in th...

  8. VanGogh: A Unified Multimodal Diffusion-based Framework for Video Colorization

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A single diffusion model colorizes videos from text, exemplar images, and hints, reporting better temporal consistency and color fidelity than prior single-condition methods.

  9. LaVin-DiT: Large Vision Diffusion Transformer

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A single diffusion transformer with a spatial-temporal VAE and in-context example pairs unifies over 20 vision tasks, with strong scores on several image benchmarks and uneven evidence on video tasks.

Pith tools