Pith. sign in

REVIEW 12 cited by

Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.14780 v3 pith:V5BALSAE submitted 2024-02-22 cs.CV

classification cs.CV
keywords motioncustomizationvideodiffusionmodelstemporalappearancecustomize-a-video
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image customization has been extensively studied in text-to-image (T2I) diffusion models, leading to impressive outcomes and applications. With the emergence of text-to-video (T2V) diffusion models, its temporal counterpart, motion customization, has not yet been well investigated. To address the challenge of one-shot video motion customization, we propose Customize-A-Video that models the motion from a single reference video and adapts it to new subjects and scenes with both spatial and temporal varieties. It leverages low-rank adaptation (LoRA) on temporal attention layers to tailor the pre-trained T2V diffusion model for specific motion modeling. To disentangle the spatial and temporal information during training, we introduce a novel concept of appearance absorbers that detach the original appearance from the reference video prior to motion learning. The proposed modules are trained in a staged pipeline and inferred in a plug-and-play fashion, enabling easy extensions to various downstream tasks such as custom video generation and editing, video appearance customization and multiple motion combination. Our project page can be found at https://customize-a-video.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SketchAnimator: Animate Sketch via Motion Customization of Text-to-Video Diffusion Models

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A three-stage method (appearance LoRA, motion LoRA, SDS stroke optimization) animates a user sketch with the motion of a reference video in a one-shot setting.

  2. Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A temporal attention purification and skip-connection rerouting method that separates motion learning from appearance learning in text-to-video diffusion model customization.

  3. Generative Physical AI in Vision: A Survey

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A structured review that categorizes physics-aware generative models in vision into explicit-simulation and implicit-learning families and proposes six integration paradigms.

  4. Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A training-free method that matches sparse-point inter-frame feature correlations to transfer reference motion to generated videos with improved temporal consistency.

  5. Ingredients: Blending Custom Photos with Video Diffusion Transformers

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Ingredients adds a mask-supervised identity router to a video diffusion transformer, enabling multi-person videos from a few reference photos without per-identity fine-tuning.

  6. CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training

    cs.CV 2024-12 conditional novelty 6.0 of 10

    By training appearance and motion LoRAs on separate layers and distilling from their single-LoRA teachers for 30 steps, the combined model produces videos with a custom subject and a custom motion.

  7. MotionBridge: Dynamic Video Inbetweening with Flexible Controls

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MotionBridge generates interpolated video frames between two images while following user-supplied trajectory, mask, keyframe, guide-pixel, and text controls.

  8. SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SnapGen-V prunes, searches, and adversarially distills a video diffusion model down to 0.6B parameters that generates a five-second, 512x512 video on an iPhone 16 Pro Max in under five seconds.

  9. Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning

    cs.CV 2024-12 conditional novelty 6.0 of 10

    UES adds a self-supervised video condition to text-to-video diffusion models, enabling them to edit videos from delta prompts without paired supervision.

  10. VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A controllable diffusion transformer generates VFX videos from a reference image, text, and mask and timestamp conditions, with a new 675-video dataset and a temporal accuracy metric.

  11. Human Motion Video Generation: A Survey

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A comprehensive survey with a five-phase pipeline model for human motion video generation, covering over 200 papers and adding a new benchmark comparison of nine pose-guided methods.

  12. Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion

    cs.CV 2025-01 reject novelty 4.0 of 10

    GE-Adapter combines a temporal smoothness loss, bilateral-filtered DDIM inversion, and shared plus frame-specific prompt tokens to improve text-to-video editing, though the reported evidence is inconsistent.

Pith tools