REVIEW 10 cited by
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce MarDini, a new family of video diffusion models that integrate the advantages of masked auto-regression (MAR) into a unified diffusion model (DM) framework. Here, MAR handles temporal planning, while DM focuses on spatial generation in an asymmetric network design: i) a MAR-based planning model containing most of the parameters generates planning signals for each masked frame using low-resolution input; ii) a lightweight generation model uses these signals to produce high-resolution frames via diffusion de-noising. MarDini's MAR enables video generation conditioned on any number of masked frames at any frame positions: a single model can handle video interpolation (e.g., masking middle frames), image-to-video generation (e.g., masking from the second frame onward), and video expansion (e.g., masking half the frames). The efficient design allocates most of the computational resources to the low-resolution planning model, making computationally expensive but important spatio-temporal attention feasible at scale. MarDini sets a new state-of-the-art for video interpolation; meanwhile, within few inference steps, it efficiently generates videos on par with those of much more expensive advanced image-to-video models.
Forward citations
Cited by 10 Pith papers
-
Stream Forcing: Constructing Unified Training Trajectory for Robust Streaming Video Generation
Stream Forcing constructs a curriculum training trajectory in a Logit-normal parameterized sampling space to reconcile training coverage with inference consistency for streaming video diffusion, reporting FVD improvements.
-
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
Resampling Forcing trains autoregressive video diffusion models on self-resampled degraded histories with a causal mask, achieving stable long-horizon generation without a teacher or discriminator.
-
Controllable Coupled Image Generation via Diffusion Models
A cross-attention control method that couples backgrounds across multiple generated images by blending LLM-extracted background and entity prompts with a time-varying weight optimized for background similarity and tex...
-
Capturing Conditional Dependence via Auto-regressive Diffusion Models
Auto-regressive diffusion models provably control conditional-distribution sampling error with only a factor-K increase in inference cost, unlike vanilla diffusion where conditional error can blow up despite small joi...
-
Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression
HMA is a masked autoregressive transformer that predicts future video and actions across many robot embodiments, running up to 15x faster than prior diffusion-based video simulators while matching or improving visual ...
-
CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models
A token-pruning cache method cuts diffusion model computation by roughly half while keeping image quality, using noise magnitude, spatial clustering, and selection balance.
-
Ingredients: Blending Custom Photos with Video Diffusion Transformers
Ingredients adds a mask-supervised identity router to a video diffusion transformer, enabling multi-person videos from a few reference photos without per-identity fine-tuning.
-
Towards Precise Scaling Laws for Video Diffusion Transformers
Video diffusion transformers follow scaling laws, and tuning learning rate and batch size per model and data size makes those laws precise enough to predict loss and optimal model size at larger scales.
-
Video Diffusion Transformers are In-Context Learners
Concatenating multiple video clips into one input and fine-tuning a LoRA adapter lets a pretrained video diffusion transformer produce consistent multi-scene videos from a single prompt.
-
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
A speech-to-speech dialogue model that predicts mel-spectrograms with flow matching, jointly trained with discrete text tokens, achieves lower WER than a discrete speech token baseline.
Discussion (0). Continue with ORCID to comment.