REVIEW 4 major objections 5 minor 58 references
MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read MoTrans claims that a text-to-video diffusion model can learn a specific human-centric motion from one or a few reference videos and transfer it to new subjects and scenes without copying the reference appearance.
desk verdict Solid motion-customization recipe with a plausible central claim, but the proposed MoFid metric is confounded by appearance and the quantitative evidence is weaker than the user study. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a two-stage LoRA-based decoupling pipeline inside a pretrained text-to-video UNet. In the appearance stage, the MLLM-based recaptioner expands the user prompt to describe the foreground subject and background, and only spatial self-attention and feed-forward LoRAs are updated. In the motion stage, the appearance injector takes a random frame from the reference video, encodes it with an OpenCLIP image encoder, passes it through a linear layer, and broadcasts the result into the spatial transformer's hidden states before the temporal transformers, so the temporal LoRAs are steered toward motion; concurrently, a motion enhancer MLP produces a residual embedding from the mean-pooled video embedding and the verb's text embedding, which is added to the verb token and regularized with an L2 term. These pieces together are what the paper claims separates what the video looks like from what the video does.
What would settle it
Train MoTrans on two reference sets with the same motion but very different appearances (for example, a human lifting weights and a fluffy teddy bear lifting weights), generate videos of a new subject such as a tiger, and compare the outputs for appearance differences beyond sampling noise. If the tiger's texture or identity systematically changes with the reference set, the appearance injector is leaking appearance and the decoupling claim fails.
Extended reading notes
Core claim
The core claim is that appearance and motion in a small reference video set can be decoupled by giving the model two complementary appearance signals—an MLLM-expanded textual description and a visual embedding of a random frame—while reserving a separate residual embedding to represent the motion itself. Concretely, the method trains spatial LoRAs on the recaptioned prompt, freezes them, then trains temporal LoRAs while broadcasting a linearly projected image embedding into the UNet hidden states just before the temporal transformers, compelling those layers to model only dynamics. The motion enhancer locates the verb in the prompt, concatenates the mean-pooled video embedding with the verb's text embedding, and forms a residual embedding added to that token, so the specific motion is represented at the conditioning level. The paper argues this two-stage multimodal decoupling solves the overfitting failure of prior fine-tuning approaches and achieves superior performance in one-shot and few-shot motion transfer.
Load-bearing premise
The decoupling claim rests on the assumption that broadcasting a randomly chosen frame's image embedding into the UNet hidden states just before the temporal transformers steers the temporal LoRAs to learn motion only, rather than leaking the reference video's appearance into the generated video.
Editorial extensions
If this is right
- One-shot motion customization is feasible: a single reference video can transfer a complex motion like skateboarding or playing guitar to an arbitrary new subject.
- Few-shot customization from about 4–10 reference videos further improves text alignment, entity alignment, and temporal consistency, and the method also supports simultaneous subject and motion customization from an exemplar image set.
- The decoupling is measurable: removing the recaptioner or the appearance injector lowers CLIP text/entity alignment while raising motion fidelity, consistent with overfitting to appearance and motion.
- Training fits on a single A100 GPU with roughly 600 steps, and inference produces 24-frame clips at 576×320 in about 19 seconds, so the pipeline is lightweight for short-clip production.
- The proposed Motion Fidelity (MoFid) metric, built on VideoMAE embeddings, gives a quantitative handle on whether generated videos actually perform the reference motion, complementing CLIP-based metrics.
Reading between the lines
- A consequence the paper does not state is that the appearance injector's broadcast weight acts as a dial between motion fidelity and subject alignment; a direct test would sweep this scale and measure the MoFid/CLIP-E trade-off.
- Because the motion enhancer depends on locating the verb via part-of-speech tagging, the method is sensitive to prompt phrasing; an unstated extension is testing whether the residual embedding still helps when the prompt has no explicit verb or contains several motion verbs.
- The frozen spatial LoRAs open a path toward composable customization: an independently trained motion LoRA from one reference set could be paired with a spatial LoRA from another subject, something the paper demonstrates with exemplar images but does not test for arbitrary independently trained adapters.
- If the decoupling mechanism is robust, the specific image encoder used for the appearance injector may not be essential; swapping OpenCLIP for another frame encoder and re-measuring MoFid and CLIP-E would probe how much of the decoupling is due to the architecture versus the encoder's inductive bias.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MoTrans, a two-stage method for customizing motion transfer with text-driven video diffusion models. In an appearance learning stage, an MLLM-based recaptioner expands the prompt and spatial LoRAs are trained; in a motion learning stage, an appearance injector pre-injects frame embeddings into the UNet and temporal LoRAs are trained together with a motion-specific residual embedding. The method supports one-shot and few-shot motion customization and is evaluated on a self-collected 12-motion benchmark against ZeroScope-finetune, Tune-a-Video, MotionDirector, LAMP, and DreamVideo using CLIP-T, CLIP-E, TempCons, and a newly proposed MoFid metric, plus a user study. The central claim is that MoTrans learns the specific motion from reference videos and transfers it to new subjects while avoiding appearance overfitting.
Significance. If the central claim holds, MoTrans is a practically useful contribution: it addresses appearance-motion decoupling in a minimally invasive way (LoRAs only, no per-frame pose control) and supports both single- and multi-video customization, including simultaneous subject and motion customization. The paper ships a relatively complete experimental package with qualitative comparisons, ablations, and a user study. However, the quantitative support for the motion-transfer claim depends on a newly introduced metric, MoFid, that is not validated against human judgment or shown to isolate motion from appearance. The absence of error bars or significance tests further weakens the numerical comparisons. The method itself is clearly described conceptually, but one core equation is inconsistent with its textual description, making the exact mechanism ambiguous. These issues are fixable but currently prevent the evidence from fully supporting the stated claims.
major comments (4)
- [Section 3.2, Eq. (2)] The text states that the linearly projected image embedding is 'summed' with the hidden states before the temporal transformer, but Eq. (2) uses the broadcast operator ⊙, which conventionally denotes element-wise product. This is a load-bearing ambiguity: an additive injection and a multiplicative injection are different mechanisms and would have different effects on how appearance information interacts with temporal layers. Please correct the equation or the text, specify the exact implementation, and, ideally, provide an ablation comparing additive vs. multiplicative injection to confirm that the reported behavior is due to the intended operation.
- [Section 4.1, Eq. (8), and Section B.2] The proposed MoFid metric is not validated and appears confounded by appearance. The paper's own analysis in B.2 says that a high MoFid with low CLIP-T/CLIP-E 'typically indicates an overfitting to the reference's appearance,' which means MoFid responds to appearance similarity as well as motion. No control experiment is provided to show that MoFid isolates motion (e.g., same subject with different motions, different subjects with the same motion, or correlation with human motion-similarity judgments). Consequently, the key quantitative comparison in Table 1 is ambiguous: for one-shot, ZeroScope-finetune achieves a higher MoFid (0.6011) than MoTrans (0.5679), and the paper attributes this to appearance overfitting, but without a validated metric this could also indicate higher motion fidelity. Please validate MoFid or supplement it with a human-validated motion fidelity evaluation before using it as the primary evidence for the motion-transfer claim.
- [Tables 1, 2, and 3] All quantitative results are reported as point estimates without error bars, multiple seeds, or significance tests. Given that several differences are small (e.g., one-shot MoFid 0.5679 for MoTrans vs. 0.5627 for Tune-a-Video; ablation MoFid 0.5679 vs. 0.5643 without motion enhancer), the reported improvements may be within run-to-run noise. Please report mean and variance over at least three seeds or reference-video subsets, and apply an appropriate significance test for the main comparisons against the strongest baselines.
- [Section 3.2, Appearance injector] The appearance injector randomly selects one image embedding from the reference video(s) at each training step. When multiple reference videos are provided for few-shot learning, the text does not specify whether the random selection is uniform across frames within a single randomly chosen video, or across all videos. This lack of detail affects reproducibility and could also introduce unintended variance in the learned motion representation. Please specify the sampling procedure precisely.
minor comments (5)
- [Section 4.1, Implementation details] Equation (1) refers to a VQ-VAE for compressing frames, but ZeroScope, the stated base model, uses a latent diffusion VAE rather than a VQ-VAE. Please correct the terminology.
- [Section 4.1, Eq. (8)] The notation |¯v_m| is ambiguous; it is later used as the count of generated videos for motion m. Please define it explicitly, for example as N_m, to avoid confusion with set cardinality or absolute value.
- [Section 4.2, User study] The user study reports 1536 answers from 32 participants, but no statistical significance testing is reported, and the recruitment process is not described. A paired comparison test (e.g., Wilcoxon signed-rank) over the per-participant preferences would strengthen the claim that MoTrans is preferred over each baseline.
- [Appendix A] The benchmark is self-constructed and the paper does not mention whether the dataset or the evaluation code will be released. For reproducibility of the quantitative comparisons, please state the release plan or provide full details of the collected videos and prompt templates.
- [Section B.2 and Figure 9] Figure 9 is referenced in the supplementary analysis but appears to be a scatter plot without axis labels or a clear description of what each point represents. Please add axis labels and a caption explaining the plotted points (e.g., one point per method per motion or per generated video).
Circularity Check
No significant circularity: the claimed motion-transfer result is supported by standard diffusion training, external baselines, and a user study, with no fitted quantity renamed as a prediction.
full rationale
The paper's derivation chain is self-contained in the relevant sense. The two training stages use standard denoising losses (Eqs. 1 and 6), each optimizing LoRA weights and a residual verb embedding toward the usual noise-prediction objective; no evaluation metric or benchmark value is optimized during training, and no parameter is fitted to the quantity that is later reported as a prediction. The proposed MoFid metric (Eq. 8) compares generated videos to a randomly selected training video, which is a natural way to measure resemblance to the reference motion; while the paper itself notes in Section B.2 that high MoFid with low CLIP scores can indicate appearance overfitting, that is a limitation of the metric's interpretability, not a circular derivation. The central quantitative claim is additionally triangulated by CLIP-T, CLIP-E, temporal consistency, a user study, and comparisons against external baselines (ZeroScope, VideoCrafter, Tune-a-Video, LAMP, MotionDirector, DreamVideo), so the result does not reduce to the authors' own metric or benchmark. The self-citations present ([43], [57]) are contextual references used for general statements about applications and multimodal benefit; they are not load-bearing, do not invoke a uniqueness theorem, and do not substitute for the paper's own experiments. No equation in the paper is defined in terms of the conclusion it is used to support, and no fitted input is renamed as a prediction. Accordingly, there is no significant circularity.
Assumptions & free parameters
free parameters (5)
- regularization coefficient lambda =
1e-4
- LoRA rank =
32
- learning rate =
5e-4
- training steps =
600
- recaption prompt length cap =
15 words
assumptions (5)
- domain assumption The pretrained ZeroScope T2V model can be effectively adapted to new motions by fine-tuning only LoRAs in spatial and temporal transformers.
- domain assumption The verb in the text prompt is the primary carrier of motion information.
- domain assumption The appearance of the reference subject can be fully described by a recaptioned text prompt (up to 15 words) and a single random frame embedding.
- domain assumption VideoMAE embeddings are a reliable measure of motion similarity.
- ad hoc to paper The appearance injector operation in Equation (2) is well-defined and implemented as intended.
Cite this review
Pith. "Pith review of MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models." pith.science (2026). https://pith.science/paper/FZ6UT6QC
@misc{pith2026241201343,
author = {Pith},
title = {Pith review of: MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/FZ6UT6QC}},
note = {Machine review of arXiv:2412.01343}
}
read the original abstract
Existing pretrained text-to-video (T2V) models have demonstrated impressive abilities in generating realistic videos with basic motion or camera movement. However, these models exhibit significant limitations when generating intricate, human-centric motions. Current efforts primarily focus on fine-tuning models on a small set of videos containing a specific motion. They often fail to effectively decouple motion and the appearance in the limited reference videos, thereby weakening the modeling capability of motion patterns. To this end, we propose MoTrans, a customized motion transfer method enabling video generation of similar motion in new context. Specifically, we introduce a multimodal large language model (MLLM)-based recaptioner to expand the initial prompt to focus more on appearance and an appearance injection module to adapt appearance prior from video frames to the motion modeling process. These complementary multimodal representations from recaptioned prompt and video frames promote the modeling of appearance and facilitate the decoupling of appearance and motion. In addition, we devise a motion-specific embedding for further enhancing the modeling of the specific motion. Experimental results demonstrate that our method effectively learns specific motion pattern from singular or multiple reference videos, performing favorably against existing methods in customized video generation.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2021. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 1728–1738
work page 2021
-
[2]
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al . 2023. Improving im- age generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf 2, 3 (2023), 8
2023
-
[3]
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. 2024. Video generation models as world simulators. (2024). https://openai.com/research/video-generation-models-as-world-simulators
2024
-
[4]
Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh. 2017. Realtime multi- person 2d pose estimation using part affinity fields. In Proceedings of the IEEE conference on computer vision and pattern recognition . 7291–7299
2017
-
[5]
cerspense. 2023. https://huggingface.co/cerspense/zeroscope_v2_576w
work page 2023
-
[6]
Di Chang, Yichun Shi, Quankai Gao, Jessica Fu, Hongyi Xu, Guoxian Song, Qing Yan, Xiao Yang, and Mohammad Soleymani. 2023. MagicDance: Realistic Human Dance Video Generation with Motions & Facial Expressions Transfer. arXiv preprint arXiv:2311.12052 (2023)
arXiv 2023
-
[7]
Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. 2024. VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models. arXiv:2401.09047 [cs.CV]
arXiv 2024
-
[8]
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhong- dao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. 2023. PixArt-𝛼: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis. arXiv:2310.00426 [cs.CV]
arXiv 2023
Show all 58 references
-
[9]
Xi Chen, Lianghua Huang, Yu Liu, Yujun Shen, Deli Zhao, and Hengshuang Zhao. 2023. Anydoor: Zero-shot object-level image customization. arXiv preprint arXiv:2307.09481 (2023)
2023 arXiv
-
[10]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv prepri...
2020 arXiv
-
[11]
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. 2023. Structure and content-guided video synthesis with diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 7346–7356
2023
-
[12]
Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. 2022. An image is worth one word: Personalizing text-to-image generation using textual inversion. arXiv preprint arXiv:2208.01618 (2022)
2022 arXiv
-
[13]
Rinon Gal, Moab Arar, Yuval Atzmon, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. 2023. Encoder-based domain tuning for fast personalization of text- to-image models. ACM Transactions on Graphics (TOG) 42, 4 (2023), 1–13
2023
-
[14]
Rohit Girdhar, Mannat Singh, Andrew Brown, Quentin Duval, Samaneh Azadi, Sai Saketh Rambhatla, Akbar Shah, Xi Yin, Devi Parikh, and Ishan Misra. 2023. Emu video: Factorizing text-to-video generation by explicit image conditioning. arXiv preprint arXiv:2311.10709 (2023)
2023 arXiv
-
[15]
Rıza Alp Güler, Natalia Neverova, and Iasonas Kokkinos. 2018. Densepose: Dense human pose estimation in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition . 7297–7306
2018
-
[16]
Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. 2023. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725 (2023)
2023 arXiv
-
[17]
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei, Irfan Essa, Lu Jiang, and José Lezama. 2023. Photorealistic video generation with diffusion models. arXiv preprint arXiv:2312.06662 (2023)
2023 arXiv
-
[18]
Dan Hendrycks and Kevin Gimpel. 2016. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415 (2016)
2016 arXiv
-
[19]
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. 2022. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303 (2022)
2022 arXiv
-
[20]
Jonathan Ho and Tim Salimans. 2022. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598 (2022)
2022 arXiv
-
[21]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. 2022. Video diffusion models. arXiv:2204.03458 (2022)
2022 arXiv
-
[22]
Li Hu, Xin Gao, Peng Zhang, Ke Sun, Bang Zhang, and Liefeng Bo. 2023. Ani- mate anyone: Consistent and controllable image-to-video synthesis for character animation. arXiv preprint arXiv:2311.17117 (2023)
2023 arXiv
-
[23]
Diederik P Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013)
2013 arXiv
-
[24]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruc- tion Tuning
2023
-
[25]
Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C Kot. 2020. NTU RGB+D 120: A large-scale benchmark for 3D human activity understanding. IEEE Transactions on Pattern Analysis and Machine Intelligence 42, 10 (2020), 2684–2701
2020
-
[26]
Zhiheng Liu, Yifei Zhang, Yujun Shen, Kecheng Zheng, Kai Zhu, Ruili Feng, Yu Liu, Deli Zhao, Jingren Zhou, and Yang Cao. 2023. Cones 2: Customizable image synthesis with multiple subjects. arXiv preprint arXiv:2305.19327 (2023)
2023 arXiv
-
[27]
Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101 (2017)
2017 arXiv
-
[28]
Jiafeng Mao, Xueting Wang, and Kiyoharu Aizawa. 2023. Guided image synthe- sis via initial image editing in diffusion model. In Proceedings of the 31st ACM International Conference on Multimedia . 5321–5329
2023
-
[29]
Joanna Materzynska, Josef Sivic, Eli Shechtman, Antonio Torralba, Richard Zhang, and Bryan Russell. 2023. Customizing Motion in Text-to-Video Diffusion Models. arXiv preprint arXiv:2312.04966 (2023)
2023 arXiv
-
[30]
OpenAI. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
2023 arXiv
-
[31]
William Peebles and Saining Xie. 2023. Scalable diffusion models with transform- ers. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 4195–4205
2023
-
[32]
Bo Peng, Xinyuan Chen, Yaohui Wang, Chaochao Lu, and Yu Qiao. 2023. Con- ditionVideo: Training-Free Condition-Guided Text-to-Video Generation. arXiv preprint arXiv:2310.07697 (2023)
2023 arXiv
-
[33]
pikalab. 2023. https://pika.art/home
2023
-
[34]
Yixuan Ren, Yang Zhou, Jimei Yang, Jing Shi, Difan Liu, Feng Liu, Mingi Kwon, and Abhinav Shrivastava. 2024. Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models. arXiv preprint arXiv:2402.14780 (2024)
2024 arXiv
-
[35]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2021. High-Resolution Image Synthesis with Latent Diffusion Models. arXiv:2112.10752 [cs.CV]
2021 arXiv
-
[36]
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2023. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 22500–22510
2023
-
[37]
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. 2022. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792 (2022)
2022 arXiv
-
[38]
Jiaming Song, Chenlin Meng, and Stefano Ermon. 2021. Denoising diffusion implicit models. ICLR (2021)
2021
-
[39]
Khurram Soomro and Amir R Zamir. 2015. Action recognition in realistic sports videos. In Computer vision in sports . Springer, 181–208
2015
-
[40]
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012)
2012 arXiv
-
[41]
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Advances in neural information processing systems 35 (2022), 10078–10093
2022
-
[42]
Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. 2023. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571 (2023)
2023 arXiv
-
[43]
Qinghe Wang, Xu Jia, Xiaomin Li, Taiqing Li, Liqian Ma, Yunzhi Zhuge, and Huchuan Lu. 2024. StableIdentity: Inserting Anybody into Anywhere at First Sight. arXiv preprint arXiv:2401.15975 (2024)
2024 arXiv
-
[44]
Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, et al . 2023. Lavie: High- quality video generation with cascaded latent diffusion models. arXiv preprint arXiv:2309.15103 (2023)
2023 arXiv
-
[45]
Yujie Wei, Shiwei Zhang, Zhiwu Qing, Hangjie Yuan, Zhiheng Liu, Yu Liu, Yingya Zhang, Jingren Zhou, and Hongming Shan. 2023. Dreamvideo: Composing your dream videos with customized subject and motion.arXiv preprint arXiv:2312.04433 (2023)
2023 arXiv
-
[46]
Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. 2023. Tune-a- video: One-shot tuning of image diffusion models for text-to-video generation. In Proceedings of the IEEE/CVF International Con...
2023
-
[47]
Ruiqi Wu, Liangyu Chen, Tong Yang, Chunle Guo, Chongyi Li, and Xiangyu Zhang. 2023. Lamp: Learn a motion pattern for few-shot-based video generation. arXiv preprint arXiv:2310.10769 (2023)
2023 arXiv
-
[48]
Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Chenxu Zhang, Jiashi Feng, and Mike Zheng Shou. 2023. Magicanimate: Temporally MM ’24, October 28–November 1, 2024, Melbourne, VIC, Australia. Xiaomin Li et al. consistent human image animation using diffusio...
2023 arXiv
-
[49]
Binxin Yang, Shuyang Gu, Bo Zhang, Ting Zhang, Xuejin Chen, Xiaoyan Sun, Dong Chen, and Fang Wen. 2023. Paint by example: Exemplar-based image editing with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18381–18391
2023
-
[50]
Ling Yang, Zhaochen Yu, Chenlin Meng, Minkai Xu, Stefano Ermon, and Bin Cui
-
[51]
Shiyuan Yang, Xiaodong Chen, and Jing Liao. 2023. Uni-paint: A unified frame- work for multimodal image inpainting with pretrained diffusion model. In Pro- ceedings of the 31st ACM International Conference on Multimedia . 3190–3199
2023
-
[52]
Yan Zeng, Guoqiang Wei, Jiani Zheng, Jiaxin Zou, Yang Wei, Yuchen Zhang, and Hang Li. 2023. Make pixels dance: High-dynamic video generation. arXiv preprint arXiv:2311.10982 (2023)
2023 arXiv
-
[53]
Yuxin Zhang, Fan Tang, Nisha Huang, Haibin Huang, Chongyang Ma, Weiming Dong, and Changsheng Xu. 2023. MotionCrafter: One-Shot Motion Customiza- tion of Diffusion Models. arXiv preprint arXiv:2312.05288 (2023)
2023 arXiv
-
[54]
Jing Zhao, Heliang Zheng, Chaoyue Wang, Long Lan, Wanrong Huang, and Wenjing Yang. 2023. Null-text Guidance in Diffusion Models is Secretly a Cartoon- style Creator. arXiv preprint arXiv:2305.06710 (2023)
2023 arXiv
-
[55]
Rui Zhao, Yuchao Gu, Jay Zhangjie Wu, David Junhao Zhang, Jiawei Liu, Weijia Wu, Jussi Keppo, and Mike Zheng Shou. 2023. MotionDirector: Motion Cus- tomization of Text-to-Video Diffusion Models. arXiv preprint arXiv:2310.08465 (2023)
2023 arXiv
-
[56]
Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. 2022. Magicvideo: Efficient video generation with latent diffusion models. arXiv preprint arXiv:2211.11018 (2022)
2022 arXiv
-
[57]
A tiger is drinking water in the forest
Jiawen Zhu, Simiao Lai, Xin Chen, Dong Wang, and Huchuan Lu. 2023. Vi- sual prompt multi-modal tracking. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 9516–9526. MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Model...
2023
-
[2024]
arXiv preprint arXiv:2401.11708 (2024)
Mastering text-to-image diffusion: Recaptioning, planning, and generating with multimodal llms. arXiv preprint arXiv:2401.11708 (2024)
2024 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.