Pith. sign in

REVIEW 11 cited by

OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.01199 v2 pith:XNN2I2L2 submitted 2024-09-02 cs.CV eess.IV

OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model

classification cs.CV eess.IV
keywords videocompressionod-vaevideosreconstructionlatentlvdmsdiffusion
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Variational Autoencoder (VAE), compressing videos into latent representations, is a crucial preceding component of Latent Video Diffusion Models (LVDMs). With the same reconstruction quality, the more sufficient the VAE's compression for videos is, the more efficient the LVDMs are. However, most LVDMs utilize 2D image VAE, whose compression for videos is only in the spatial dimension and often ignored in the temporal dimension. How to conduct temporal compression for videos in a VAE to obtain more concise latent representations while promising accurate reconstruction is seldom explored. To fill this gap, we propose an omni-dimension compression VAE, named OD-VAE, which can temporally and spatially compress videos. Although OD-VAE's more sufficient compression brings a great challenge to video reconstruction, it can still achieve high reconstructed accuracy by our fine design. To obtain a better trade-off between video reconstruction quality and compression speed, four variants of OD-VAE are introduced and analyzed. In addition, a novel tail initialization is designed to train OD-VAE more efficiently, and a novel inference strategy is proposed to enable OD-VAE to handle videos of arbitrary length with limited GPU memory. Comprehensive experiments on video reconstruction and LVDM-based video generation demonstrate the effectiveness and efficiency of our proposed methods.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Ultra-Fast Neural Video Compression

    cs.CV 2026-06 unverdicted novelty 7.0

    DCVC-UF uses chunk-based joint encoding and parallel frame-specific decoding to deliver ultra-fast neural video compression while claiming new state-of-the-art rate-distortion performance.

  2. Efficient Video Diffusion Models: Advancements and Challenges

    cs.CV 2026-04 unverdicted novelty 7.0

    A survey that groups efficient video diffusion methods into four paradigms—step distillation, efficient attention, model compression, and cache/trajectory optimization—and outlines open challenges for practical use.

  3. Modality-Aware and Anatomical Vector-Quantized Autoencoding for Multimodal Brain MRI

    cs.CV 2026-04 unverdicted novelty 7.0

    NeuroQuant is a modality-aware 3D VQ-VAE that uses dual-stream encoding, a shared anatomical codebook, and FiLM to achieve superior multi-modal brain MRI reconstruction.

  4. ChopGrad: Pixel-Wise Losses for Latent Video Diffusion via Truncated Backpropagation

    cs.CV 2026-03 unverdicted novelty 7.0

    ChopGrad truncates backpropagation to local frame windows in video diffusion models, reducing memory from linear in frame count to constant while enabling pixel-wise loss fine-tuning.

  5. TivTok: Broadcasting Time-Invariant Tokens for Scalable Video Tokenization

    cs.CV 2026-06 unverdicted novelty 6.0

    TivTok factorizes video clips into reusable time-invariant tokens and frame-specific time-variant tokens via Scope-Induced Factorization and Invariant Broadcasting, achieving 2.91x better compression for 128-frame vid...

  6. Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability

    cs.CV 2025-12 conditional novelty 6.0

    A video VAE whose latents are biased toward low frequencies and a few dominant channel modes improves text-to-video diffusion convergence and reward scores.

  7. Task-Oriented Communication for Human Action Understanding via Edge-Cloud Co-Inference

    eess.SP 2026-05 unverdicted novelty 5.0

    TOAU compresses human motion videos to 9 bits per frame with pose estimation and VQ-VAE, then aligns the tokens to a vision-language model via a lightweight projector, achieving 1% transmission payload and 20% latency...

  8. Video Generation with Predictive Latents

    cs.CV 2026-05 unverdicted novelty 5.0

    PV-VAE improves video latent spaces for generation by unifying reconstruction with future-frame prediction, reporting 52% faster convergence and 34.42 FVD gain over Wan2.2 VAE on UCF101.

  9. Integrating Anatomical Priors into a Causal Diffusion Model

    cs.CV 2025-09 conditional novelty 5.0

    A mask-guided causal diffusion model generates 3D brain MRI counterfactuals whose cortical volume measurements match known alcohol-use-disorder effects, but those effects are inserted via fitted masks rather than discovered.

  10. HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics

    cs.CV 2025-08 unverdicted novelty 5.0

    HumanGenesis is a claimed state-of-the-art framework that couples 3D Gaussian reconstruction, LLM-based critique, pose guidance, and diffusion-based harmonization for synthetic human video generation; unverified in th...

  11. HunyuanVideo: A Systematic Framework For Large Video Generative Models

    cs.CV 2024-12 unverdicted novelty 5.0

    HunyuanVideo presents a 13B-parameter open-source video generative model with integrated data, architecture, training, and inference systems whose professional evaluations show it outperforming prior SOTA models inclu...