REVIEW 12 cited by
Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Image customization has been extensively studied in text-to-image (T2I) diffusion models, leading to impressive outcomes and applications. With the emergence of text-to-video (T2V) diffusion models, its temporal counterpart, motion customization, has not yet been well investigated. To address the challenge of one-shot video motion customization, we propose Customize-A-Video that models the motion from a single reference video and adapts it to new subjects and scenes with both spatial and temporal varieties. It leverages low-rank adaptation (LoRA) on temporal attention layers to tailor the pre-trained T2V diffusion model for specific motion modeling. To disentangle the spatial and temporal information during training, we introduce a novel concept of appearance absorbers that detach the original appearance from the reference video prior to motion learning. The proposed modules are trained in a staged pipeline and inferred in a plug-and-play fashion, enabling easy extensions to various downstream tasks such as custom video generation and editing, video appearance customization and multiple motion combination. Our project page can be found at https://customize-a-video.github.io.
Forward citations
Cited by 12 Pith papers
-
SketchAnimator: Animate Sketch via Motion Customization of Text-to-Video Diffusion Models
A three-stage method (appearance LoRA, motion LoRA, SDS stroke optimization) animates a user sketch with the motion of a reference video in a one-shot setting.
-
Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models
A temporal attention purification and skip-connection rerouting method that separates motion learning from appearance learning in text-to-video diffusion model customization.
-
Generative Physical AI in Vision: A Survey
A structured review that categorizes physics-aware generative models in vision into explicit-simulation and implicit-learning families and proposes six integration paradigms.
-
Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss
A training-free method that matches sparse-point inter-frame feature correlations to transfer reference motion to generated videos with improved temporal consistency.
-
Ingredients: Blending Custom Photos with Video Diffusion Transformers
Ingredients adds a mask-supervised identity router to a video diffusion transformer, enabling multi-person videos from a few reference photos without per-identity fine-tuning.
-
CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training
By training appearance and motion LoRAs on separate layers and distilling from their single-LoRA teachers for 30 steps, the combined model produces videos with a custom subject and a custom motion.
-
MotionBridge: Dynamic Video Inbetweening with Flexible Controls
MotionBridge generates interpolated video frames between two images while following user-supplied trajectory, mask, keyframe, guide-pixel, and text controls.
-
SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device
SnapGen-V prunes, searches, and adversarially distills a video diffusion model down to 0.6B parameters that generates a five-second, 512x512 video on an iPhone 16 Pro Max in under five seconds.
-
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
UES adds a self-supervised video condition to text-to-video diffusion models, enabling them to edit videos from delta prompts without paired supervision.
-
VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer
A controllable diffusion transformer generates VFX videos from a reference image, text, and mask and timestamp conditions, with a new 675-video dataset and a temporal accuracy metric.
-
Human Motion Video Generation: A Survey
A comprehensive survey with a five-phase pipeline model for human motion video generation, covering over 200 papers and adding a new benchmark comparison of nine pose-guided methods.
-
Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion
GE-Adapter combines a temporal smoothness loss, bilateral-filtered DDIM inversion, and shared plus frame-specific prompt tokens to improve text-to-video editing, though the reported evidence is inconsistent.
Discussion (0). Continue with ORCID to comment.