Pith. sign in

REVIEW 2 cited by

Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.08818 v2 pith:V2JSTAML submitted 2023-04-18 cs.CV cs.LG

classification cs.CVcs.LG
keywords diffusionmodelvideoimagelatentldmsmodelsresolution
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Latent Diffusion Models (LDMs) enable high-quality image synthesis while avoiding excessive compute demands by training a diffusion model in a compressed lower-dimensional latent space. Here, we apply the LDM paradigm to high-resolution video generation, a particularly resource-intensive task. We first pre-train an LDM on images only; then, we turn the image generator into a video generator by introducing a temporal dimension to the latent space diffusion model and fine-tuning on encoded image sequences, i.e., videos. Similarly, we temporally align diffusion model upsamplers, turning them into temporally consistent video super resolution models. We focus on two relevant real-world applications: Simulation of in-the-wild driving data and creative content creation with text-to-video modeling. In particular, we validate our Video LDM on real driving videos of resolution 512 x 1024, achieving state-of-the-art performance. Furthermore, our approach can easily leverage off-the-shelf pre-trained image LDMs, as we only need to train a temporal alignment model in that case. Doing so, we turn the publicly available, state-of-the-art text-to-image LDM Stable Diffusion into an efficient and expressive text-to-video model with resolution up to 1280 x 2048. We show that the temporal layers trained in this way generalize to different fine-tuned text-to-image LDMs. Utilizing this property, we show the first results for personalized text-to-video generation, opening exciting directions for future content creation. Project page: https://research.nvidia.com/labs/toronto-ai/VideoLDM/

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 17 citations worldwide. Full citation record

  1. RDPO: Real Data Preference Optimization for Physics Consistency Video Generation

    cs.CV 2025-06 conditional novelty 8.0 of 10

    RDPO builds preference pairs by reverse-sampling real video latents with a pre-trained generator, then fine-tunes with Flow-DPO, improving physics consistency metrics on two video models.

  2. Align Your Structures: Generating Trajectories with Structure Pretraining for Molecular Dynamics

    cs.LG 2026-04 conditional novelty 6.5 of 10

    Structure-pretrained diffusion plus an equivariant temporal interpolator generates chemically realistic MD trajectories on small molecules, tetrapeptides, and proteins by separating spatial and temporal learning.

Pith tools