Pith. sign in

REVIEW 4 cited by

CV-VAE: A Compatible Video VAE for Latent Generative Video Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.20279 v2 pith:7U5M4AHJ submitted 2024-05-30 cs.CV cs.AIeess.IV

classification cs.CVcs.AIeess.IV
keywords videomodelslatentspacecompatibilitycv-vaediffusion-basedframes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Spatio-temporal compression of videos, utilizing networks such as Variational Autoencoders (VAE), plays a crucial role in OpenAI's SORA and numerous other video generative models. For instance, many LLM-like video models learn the distribution of discrete tokens derived from 3D VAEs within the VQVAE framework, while most diffusion-based video models capture the distribution of continuous latent extracted by 2D VAEs without quantization. The temporal compression is simply realized by uniform frame sampling which results in unsmooth motion between consecutive frames. Currently, there lacks of a commonly used continuous video (3D) VAE for latent diffusion-based video models in the research community. Moreover, since current diffusion-based approaches are often implemented using pre-trained text-to-image (T2I) models, directly training a video VAE without considering the compatibility with existing T2I models will result in a latent space gap between them, which will take huge computational resources for training to bridge the gap even with the T2I models as initialization. To address this issue, we propose a method for training a video VAE of latent video models, namely CV-VAE, whose latent space is compatible with that of a given image VAE, e.g., image VAE of Stable Diffusion (SD). The compatibility is achieved by the proposed novel latent space regularization, which involves formulating a regularization loss using the image VAE. Benefiting from the latent space compatibility, video models can be trained seamlessly from pre-trained T2I or video models in a truly spatio-temporally compressed latent space, rather than simply sampling video frames at equal intervals. With our CV-VAE, existing video models can generate four times more frames with minimal finetuning. Extensive experiments are conducted to demonstrate the effectiveness of the proposed video VAE.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A hierarchical motion autoencoder with a conditional diffusion decoder reconstructs 16-frame videos from latents as small as 0.07% of the input size while maintaining competitive PSNR and perceptual scores.

  2. Interspatial Attention for Efficient 4D Human Video Generation

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A new symmetric 3D-to-2D attention mechanism with relative positional encodings, plus a motion-tuned video VAE, improves controllable 4D human video generation.

  3. Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Can3Tok tokenizes scene-level 3D Gaussian splats into canonical latent tokens with normalization and saliency filtering, enabling reconstruction and text/image-to-3D generation.

  4. HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    HumanGenesis is a claimed state-of-the-art framework that couples 3D Gaussian reconstruction, LLM-based critique, pose guidance, and diffusion-based harmonization for synthetic human video generation; unverified in th...

Pith tools