Pith. sign in

REVIEW 2 cited by

Video Diffusion Models: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.03150 v2 pith:QCADFRMX submitted 2024-05-06 cs.CV cs.LG

classification cs.CVcs.LG
keywords modelssurveyvideodiffusionapplicationsarchitecturaldevelopmentsevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion generative models have recently become a powerful technique for creating and modifying high-quality, coherent video content. This survey provides a comprehensive overview of the critical components of diffusion models for video generation, including their applications, architectural design, and temporal dynamics modeling. The paper begins by discussing the core principles and mathematical formulations, then explores various architectural choices and methods for maintaining temporal consistency. A taxonomy of applications is presented, categorizing models based on input modalities such as text prompts, images, videos, and audio signals. Advancements in text-to-video generation are discussed to illustrate the state-of-the-art capabilities and limitations of current approaches. Additionally, the survey summarizes recent developments in training and evaluation practices, including the use of diverse video and image datasets and the adoption of various evaluation metrics to assess model performance. The survey concludes with an examination of ongoing challenges, such as generating longer videos and managing computational costs, and offers insights into potential future directions for the field. By consolidating the latest research and developments, this survey aims to serve as a valuable resource for researchers and practitioners working with video diffusion models. Website: https://github.com/ndrwmlnk/Awesome-Video-Diffusion-Models

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dynamic data generation and dynamic portfolio selection: an application of a score-based diffusion model

    q-fin.PM 2025-07 reject novelty 6.0 of 10

    An adaptive score-based diffusion model generates sequential market scenarios with adapted-Wasserstein error bounds, and a policy-gradient agent trained on these scenarios outperforms several portfolio benchmarks.

  2. TrajFlow: Multi-modal Motion Prediction via Flow Matching

    cs.CV 2025-06 conditional novelty 6.0 of 10

    TrajFlow uses flow matching with a multi-query transformer to predict multiple trajectories in one pass and a Plackett-Luce ranking loss to improve confidence scores, reporting small SOTA gains on WOMD.

Pith tools