Pith. sign in

REVIEW 7 cited by

Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.03162 v2 pith:WEAA53AM submitted 2024-02-05 cs.CV

classification cs.CV
keywords cameramotionmovementobjectcontroldirect-a-videoindependentlyvideo
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent text-to-video diffusion models have achieved impressive progress. In practice, users often desire the ability to control object motion and camera movement independently for customized video creation. However, current methods lack the focus on separately controlling object motion and camera movement in a decoupled manner, which limits the controllability and flexibility of text-to-video models. In this paper, we introduce Direct-a-Video, a system that allows users to independently specify motions for multiple objects as well as camera's pan and zoom movements, as if directing a video. We propose a simple yet effective strategy for the decoupled control of object motion and camera movement. Object motion is controlled through spatial cross-attention modulation using the model's inherent priors, requiring no additional optimization. For camera movement, we introduce new temporal cross-attention layers to interpret quantitative camera movement parameters. We further employ an augmentation-based approach to train these layers in a self-supervised manner on a small-scale dataset, eliminating the need for explicit motion annotation. Both components operate independently, allowing individual or combined control, and can generalize to open-domain scenarios. Extensive experiments demonstrate the superiority and effectiveness of our method. Project page and code are available at https://direct-a-video.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding

    cs.CV 2025-08 conditional novelty 7.0 of 10

    LSD-3D generates explicit, 3D-consistent driving scenes by combining a generated proxy mesh with geometry-grounded distillation from a 2D diffusion model.

  2. Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Single-image novel view synthesis is decomposed into panorama outpainting plus keyframe-conditioned video diffusion, producing loop-consistent scene tours.

  3. Yume: An Interactive World Generation Model

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A diffusion-based video model generates extendable, keyboard-controlled walkthroughs from a single input image, using quantized camera actions as text prompts.

  4. CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    CAT4D uses a multi-view video diffusion model to convert monocular video into multi-view video and reconstruct a dynamic 3D Gaussian scene, with competitive results on 4D reconstruction benchmarks.

  5. AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

    cs.CV 2024-11 conditional novelty 6.0 of 10

    AC3D improves camera control in video diffusion transformers by conditioning only early denoising steps and the first 8 of 32 blocks, and by adding 20K static-camera dynamic videos to training.

  6. I2VControl: Disentangled and Unified Video Motion Synthesis Control

    cs.CV 2024-11 conditional novelty 6.0 of 10

    I2VControl unifies camera, drag, and brush controls into a single point-trajectory-based adapter for image-to-video diffusion models, enabling conflict-free combined motion control.

  7. Video Diffusion Transformers are In-Context Learners

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Concatenating multiple video clips into one input and fine-tuning a LoRA adapter lets a pretrained video diffusion transformer produce consistent multi-scene videos from a single prompt.

Pith tools