Pith. sign in

REVIEW 12 cited by

VDT: General-purpose Video Diffusion Transformers via Mask Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13311 v2 pith:5B6D2SV5 submitted 2023-05-22 cs.CV

classification cs.CV
keywords videomaskmodelinggenerationmechanismspatial-temporaltransformersconditioning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work introduces Video Diffusion Transformer (VDT), which pioneers the use of transformers in diffusion-based video generation. It features transformer blocks with modularized temporal and spatial attention modules to leverage the rich spatial-temporal representation inherited in transformers. We also propose a unified spatial-temporal mask modeling mechanism, seamlessly integrated with the model, to cater to diverse video generation scenarios. VDT offers several appealing benefits. 1) It excels at capturing temporal dependencies to produce temporally consistent video frames and even simulate the physics and dynamics of 3D objects over time. 2) It facilitates flexible conditioning information, \eg, simple concatenation in the token space, effectively unifying different token lengths and modalities. 3) Pairing with our proposed spatial-temporal mask modeling mechanism, it becomes a general-purpose video diffuser for harnessing a range of tasks, including unconditional generation, video prediction, interpolation, animation, and completion, etc. Extensive experiments on these tasks spanning various scenarios, including autonomous driving, natural weather, human action, and physics-based simulation, demonstrate the effectiveness of VDT. Additionally, we present comprehensive studies on how \model handles conditioning information with the mask modeling mechanism, which we believe will benefit future research and advance the field. Project page: https:VDT-2023.github.io

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms

    cs.CV 2026-07 conditional novelty 6.5 of 10

    E-VAD plus the real visible–event TJUTCM Pha dataset improve weakly supervised video anomaly detection by contrastively aligning event, video, and text features and adaptively fusing them.

  2. CDP: Towards Robust Autoregressive Visuomotor Policy Learning via Causal Diffusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Causal Diffusion Policy adds historical action conditioning and attention cache sharing to diffusion-based robot policies, improving success rates on most tested manipulation tasks under degraded observations.

  3. Programmatic Video Prediction Using Large Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ProgGen uses language-model-written programs for perception, dynamics, and rendering to predict future video frames from about ten training examples, beating large diffusion baselines on two synthetic benchmarks.

  4. Ingredients: Blending Custom Photos with Video Diffusion Transformers

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Ingredients adds a mask-supervised identity router to a video diffusion transformer, enabling multi-person videos from a few reference photos without per-identity fine-tuning.

  5. Towards Precise Scaling Laws for Video Diffusion Transformers

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Video diffusion transformers follow scaling laws, and tuning learning rate and batch size per model and data size makes those laws precise enough to predict loss and optimal model size at larger scales.

  6. Occlusion-Aware Temporally Consistent Amodal Completion for 3D Human-Object Interaction Reconstruction

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A temporal feature warping and attention fusion module for amodal completion improves occlusion handling and temporal stability in monocular HOI videos, and the completed frames support 3D Gaussian Splatting reconstruction.

  7. SlotPi: Physics-informed Object-centric Reasoning Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SlotPi combines a learned Hamiltonian energy module with spatiotemporal attention to improve object-centric video prediction and visual question answering on several datasets.

  8. A Diffusion-Driven Temporal Super-Resolution and Spatial Consistency Enhancement Framework for 4D MRI imaging

    eess.IV 2025-06 conditional novelty 5.0 of 10

    TSSC-Net uses a diffusion model conditioned on start and end frames to generate 6x more MRI time frames, then a tri-directional Mamba network to fix spatial inconsistencies.

  9. Humans Coexist, So Must Embodied Artificial Agents

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Coexistence, defined as sustained meaningful and reciprocal interaction among an agent, humans, and environment, is presented as a necessary design goal for embodied AI.

  10. LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers

    cs.LG 2024-12 conditional novelty 5.0 of 10

    LazyDiT learns small gates that decide when to reuse cached layer outputs, cutting diffusion transformer compute by up to half while matching or beating DDIM quality.

  11. LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

    cs.CV 2024-12 conditional novelty 5.0 of 10

    LinGen replaces self-attention in diffusion transformers with a linear-complexity MATE block, enabling 512p 68-second video generation on a single H100 with quality comparable to Gen-3, LumaLabs, and Kling.

  12. Video Diffusion Transformers are In-Context Learners

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Concatenating multiple video clips into one input and fine-tuning a LoRA adapter lets a pretrained video diffusion transformer produce consistent multi-scene videos from a single prompt.

Pith tools