Pith. sign in

REVIEW 4 cited by

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.06782 v2 pith:LSNEJEC4 submitted 2025-02-10 cs.CV

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT

classification cs.CV
keywords lumina-videonext-ditvideogenerationtrainingachievesdataefficiency
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Recent advancements have established Diffusion Transformers (DiTs) as a dominant framework in generative modeling. Building on this success, Lumina-Next achieves exceptional performance in the generation of photorealistic images with Next-DiT. However, its potential for video generation remains largely untapped, with significant challenges in modeling the spatiotemporal complexity inherent to video data. To address this, we introduce Lumina-Video, a framework that leverages the strengths of Next-DiT while introducing tailored solutions for video synthesis. Lumina-Video incorporates a Multi-scale Next-DiT architecture, which jointly learns multiple patchifications to enhance both efficiency and flexibility. By incorporating the motion score as an explicit condition, Lumina-Video also enables direct control of generated videos' dynamic degree. Combined with a progressive training scheme with increasingly higher resolution and FPS, and a multi-source training scheme with mixed natural and synthetic data, Lumina-Video achieves remarkable aesthetic quality and motion smoothness at high training and inference efficiency. We additionally propose Lumina-V2A, a video-to-audio model based on Next-DiT, to create synchronized sounds for generated videos. Codes are released at https://www.github.com/Alpha-VLLM/Lumina-Video.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer

    cs.CV 2025-11 unverdicted novelty 7.0

    One-to-All Animation enables alignment-free character animation and image pose transfer via self-supervised outpainting reformulation, reference extraction, hybrid fusion attention, identity-robust pose control, and t...

  2. GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

    cs.CV 2026-03 conditional novelty 6.0

    Feature-space Gaussian Splat Feature Adapter (GS-Adapter) grounds camera-controlled video diffusion in 3D Gaussians, improving geometric consistency and controllability over SEVA and CameraCtrl without retraining geom...

  3. Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance

    cs.LG 2025-08 conditional novelty 4.0

    Blending a base diffusion model with its RL-finetuned version at sampling time lets users dial alignment strength, with the blend weight corresponding to the KL-regularization coefficient beta/w.

  4. Evolution of Video Generative Foundations

    cs.CV 2026-04 unverdicted novelty 2.0

    This survey traces video generation technology from GANs to diffusion models and then to autoregressive and multimodal approaches while analyzing principles, strengths, and future trends.