REVIEW 4 major objections 6 minor 69 references
Coarse proxy videos can steer complex motion in pretrained video generators without any training, by noising latents region-wise and relaxing them onto the model manifold.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 00:18 UTC pith:T5QYNQM5
load-bearing objection Clean training-free recipe for proxy-as-dynamics; finite-K SFR is the real mechanism and the asymptotic proof does not explain it. the 4 major comments →
ProxyUp: Training-Free Proxy-Conditioned Video Generation for Controllable Dynamics
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper claims that a training-free combination of region-wise latent noising and Stochastic Flow Relaxation lets a pretrained video generator preserve essential dynamics from a coarse proxy video while synthesizing novel, prompt-aligned content and plausible foreground–background interactions—outperforming video editing and motion-transfer baselines on dynamic fidelity and text alignment for both simulated and real proxies.
What carries the argument
Region-wise latent noising plus Stochastic Flow Relaxation (SFR): invert the masked proxy to an intermediate noise level, keep those latents in motion-critical regions while replacing the rest with matched noise, then iteratively denoise and re-noise the hybrid latent so it approaches the model’s in-distribution manifold before deterministic ODE sampling.
Load-bearing premise
That a finite number of SFR re-noising rounds is enough to move a hand-composed, out-of-distribution latent onto the pretrained model’s manifold so that sampling invents coherent interactions the base model already knows how to draw.
What would settle it
If, on the same backbone and equal or larger inference budget, a region-wise-noised SDEdit or editing baseline matches or exceeds ProxyUp on Motion Rationality and Mechanics for the bread-cutting and curtain-pulling proxies (and similar held-out dynamics), the claimed benefit of SFR would not hold.
If this is right
- Physics simulations and casual real recordings can be reused as motion controllers for open-ended text-driven video synthesis without collecting paired training data.
- Video editing and motion-transfer pipelines that stay anchored to source appearance are not the right tools when the source is only a dynamics prior.
- A modest inference-time relaxation loop can repair hybrid latents that would otherwise break foreground–background coupling.
- Task-specific proxy–prompt–mask evaluation sets become necessary to measure dynamics-preserving regeneration rather than pure editing or pure generation.
Where Pith is reading between the lines
- The same recipe could turn low-fidelity game or robotics rollouts into large synthetic video corpora with controlled physics and varied visual styles.
- When the base generator lacks the needed interaction priors, better proxies alone will not fix failures—pointing toward models trained on richer physical contact data.
- Automatic or learned masks and force-application cues would reduce reliance on SAM-style foreground masks and hand-chosen t_init and K.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces proxy-conditioned video generation: a coarse proxy video (simulation or real recording) supplies foreground dynamics while a text prompt specifies novel content and interactions. Because paired proxy–target data are scarce, the authors propose ProxyUp, a training-free pipeline on pretrained rectified-flow video models (primarily Wan2.2). ProxyUp (i) ODE-inverts the masked proxy foreground to an intermediate noise level, (ii) composes a hybrid latent by region-wise latent noising (Eq. 5), (iii) applies Stochastic Flow Relaxation (SFR; Sec. 4.4, Alg. 1) to pull the hand-composed latent toward the model manifold, and (iv) finishes with deterministic ODE sampling. On a 76-clip custom set (simulation + real), ProxyUp reports gains over editing, inpainting, and motion-transfer baselines on Imaging Quality, Motion Rationality, Mechanics, and Material (Table 1), with qualitative and ablation support (Figs. 4–5).
Significance. If the empirical claims hold, the work offers a practical, training-free route to controllable dynamics that text alone cannot specify, and a useful intermediate between pure T2V, video editing, and motion transfer. Strengths include a clear three-stage pipeline, an explicit algorithm, an ablation isolating RLN and SFR (Fig. 5), hyperparameter sweeps (Fig. 6), cross-backbone checks in the appendix, and a formal (if asymptotic) Markov argument for SFR (Prop. 1, App. A). The task framing and the idea of using low-fidelity proxies as dynamics carriers are timely for physics-aware video generation. Significance is tempered by a small custom evaluation set, author-written VBench-style QA criteria, and a theory–practice gap on finite-K dynamics retention.
major comments (4)
- [Sec. 4.4, Prop. 1, App. A, Alg. 1] Sec. 4.4 / Prop. 1 / App. A: Proposition 1 only shows that the SFR Markov kernel is ergodic and D_KL(q_K || π_ID) → 0 as K → ∞ under a well-trained velocity field. That limit is pure text-conditioned sampling at t_init and would erase the inverted proxy structure. No mask is re-applied inside the SFR loop (Alg. 1 lines 6–10), and no finite-K mixing-time or information-retention bound is given. Dynamics preservation at the operating point K=15, t_init=0.922 (s=0.8) is therefore an empirical incomplete-mixing effect, not a consequence of the stated theory. The central mechanism claim—that RLN+SFR jointly preserve proxy dynamics while restoring fg–bg coupling—needs either a finite-K analysis (e.g., how much inverted foreground signal remains after K steps) or a clear reframing that SFR is a practical regularizer whose dynamics retention is empirical.
- [Sec. 5.1, Table 1, App. C.2–C.3] Sec. 5.1 / Table 1 / App. C.2–C.3: The evaluation set has only 76 custom clips with author-written multi-question criteria for MR, Mech., and Mat. conditioned on the same proxy/prompt pairs used for generation. This is acceptable for a new task but is load-bearing for the claim of consistent outperformance in dynamic fidelity. Please (i) release the full metric prompts and scoring protocol as promised, (ii) report inter-annotator or multi-run variance, and (iii) add at least one external or human preference study on dynamics fidelity vs. text alignment so that Table 1 is not solely self-defined QA.
- [Sec. 5.2, Table 1, App. C.4] Sec. 5.2 / App. C.4: Several baselines run on different backbones and default schedules (DiTFlow on CogVideoX-5B; FlowDirector on Wan2.1; VACE 14B). Appendix cross-backbone checks (Figs. 9–10) and the SDEdit step-budget study (Fig. 8) help, but the main Table 1 still mixes generators. For the primary comparison, either re-run the strongest motion-transfer and editing baselines on the same Wan2.2 backbone used by ProxyUp, or report a backbone-matched subset as the headline table so gains on MR/Mech. can be attributed to the method rather than model capacity.
- [Sec. 5.4, Fig. 6] Sec. 5.4 / Fig. 6: Free parameters K, t_init (strength s), and CFG scales are chosen by qualitative inspection. Given that the skeptic concern is precisely the incomplete-mixing regime, please quantify the trade-off: e.g., proxy-motion metrics (optical-flow or keypoint correlation in the masked region) vs. K and s, not only visual examples. Without this, it is hard to know how fragile the reported MR/Mech. gains are to hyperparameter choice.
minor comments (6)
- [Sec. 3, Fig. 2] Fig. 2 caption and body: the preliminary analysis is helpful; please state the exact strength/t values and masks used for each baseline so the trade-off narrative is reproducible.
- [Sec. 4.3, Eq. (4)] Eq. (4): the background noise variance ((1−t_init)^2 + t_init^2)I is nonstandard relative to the linear path Z_t=(1−t)Z_0+tZ_1; a one-sentence justification (marginal variance of the interpolation) would help readers.
- [Sec. 2.1] Related work (Sec. 2.1) mentions physics-related conditions but cites little recent simulation-to-video or physics-prior work; a few additional pointers would better situate proxy videos among existing control signals.
- [Sec. 6, Fig. 11] Limitation section (Sec. 6) and Fig. 11 are candid; consider moving one failure case into the main paper so readers see the dependence on the base model’s physical prior without opening the appendix.
- [Sec. 4–5] Notation: strength s is defined in a footnote and reused as s=0.8; define it once in the main text near t_init for clarity.
- [Throughout] Minor typos / consistency: “V ACE” spacing in Table 1 and captions; “out-of-distribution (OOD)latent” missing space (Sec. 4.2); arXiv id and “Preprint” header are fine for review but should be cleaned for camera-ready.
Circularity Check
No circular derivation: ProxyUp is an empirical training-free pipeline; Prop. 1 is a standard asymptotic Markov/DPI sketch, not a tautology that forces the reported metrics.
full rationale
The paper’s load-bearing claims are (i) a constructive inference procedure (region-wise latent noising + finite-K SFR + ODE sampling) and (ii) empirical outperformance on a task-specific set (Table 1, Figs. 4–5). Neither reduces to its inputs by definition. Region-wise composition (Eq. 5) is a hand-built hybrid latent, not a quantity fitted from the evaluation targets. SFR (Eqs. 6–7, Alg. 1) is a re-noising loop whose asymptotic claim (Prop. 1 / App. A) is a standard ergodicity + Data Processing Inequality argument that lim K→∞ D_KL(q_K ∥ π_ID)=0 under a well-trained velocity field; that argument does not algebraically force finite-K=15 dynamics retention or the MR/Mech. scores. Hyperparameters (K=15, s=0.8) are chosen by qualitative inspection (Fig. 6), which is empirical tuning, not a fitted-input-called-prediction. Baselines are external methods run under their defaults; metrics (IQ/MR/Mech./Mat. adapted from VBench with proxy-conditioned QA criteria) score generated videos after the fact and are not satisfied by construction of Eq. 5. Self-citations (e.g., [69]) are not used as uniqueness theorems or load-bearing premises for the central mechanism. Gaps between asymptotic theory and finite-K practice are correctness/validity concerns, not circularity. Derivation chain is self-contained against external benchmarks; score 0.
Axiom & Free-Parameter Ledger
free parameters (3)
- SFR iterations K =
15
- initial noise level t_init / strength s =
s=0.8 (t_init=0.922)
- CFG scales (inversion / SFR / sampling) =
1 / 3 / 4&3
axioms (4)
- domain assumption The pretrained rectified-flow video model’s velocity field is sufficiently accurate that one-step Euler denoising approximates the in-distribution posterior mean (Tweedie / flow matching).
- standard math Re-noising with positive Gaussian variance yields an ergodic Markov kernel whose unique stationary distribution is the model’s marginal at t_init.
- domain assumption A binary foreground mask M correctly isolates motion-critical regions of the proxy so that preserving M⊙Z_inv retains the intended dynamics.
- domain assumption The base generator already encodes enough physical interaction knowledge to complete interactions absent from the proxy once the latent is on-manifold.
invented entities (3)
-
proxy-conditioned video generation (task)
no independent evidence
-
Stochastic Flow Relaxation (SFR)
no independent evidence
-
region-wise latent noising
no independent evidence
read the original abstract
Precise control over complex dynamics remains challenging for modern video generative models, as text prompts alone often cannot specify physically plausible, fine-grained motion and interactions. We introduce $\textit{proxy-conditioned video generation}$, where a coarse proxy video from physics-based simulation or real-world recording serves as a dynamics carrier to control foreground object motion. Given a proxy video and a text prompt, the goal is to synthesize a new video that preserves the proxy dynamics while generating novel content and plausible interactions aligned with the prompt. Since paired proxy-target videos are difficult to obtain, we propose $\textbf{ProxyUp}$, a training-free framework built on pretrained video generative models. ProxyUp first inverts the proxy video into an intermediate latent representation and applies $\textbf{region-wise latent noising}$, preserving motion-critical proxy latents while injecting noise into regions intended for text-driven regeneration. To mitigate the distribution mismatch and weak foreground-background coupling introduced by this heuristic latent composition, we further propose $\textbf{Stochastic Flow Relaxation (SFR)}$, which progressively relaxes the composed latent toward the model's learned distribution before ODE sampling. Experiments on both simulation and real-world proxies show that ProxyUp outperforms strong video editing and motion transfer baselines in dynamic fidelity and text alignment.
Figures
Reference graph
Works this paper leans on
-
[1]
Recammaster: Camera-controlled generative rendering from a single video
Jianhong Bai, Menghan Xia, Xiao Fu, Xintao Wang, Lianrui Mu, Jinwen Cao, Zuozhu Liu, Haoji Hu, Xiang Bai, Pengfei Wan, et al. Recammaster: Camera-controlled generative rendering from a single video. InICCV, 2025
2025
-
[2]
Syncammaster: Synchronizing multi-camera video generation from diverse viewpoints
Jianhong Bai, Menghan Xia, Xintao Wang, Ziyang Yuan, Xiao Fu, Zuozhu Liu, Haoji Hu, Pengfei Wan, and Di Zhang. Syncammaster: Synchronizing multi-camera video generation from diverse viewpoints. InICLR, 2025
2025
-
[3]
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Do- minik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127, 2023
Pith/arXiv arXiv 2023
-
[4]
Align your latents: High-resolution video synthesis with latent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. InCVPR, 2023
2023
-
[5]
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Leo Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al. Video generation models as world simulators. OpenAI Blog, 2024
2024
-
[6]
Uni3c: Unifying precisely 3d-enhanced camera and human motion controls for video generation
Chenjie Cao, Jingkai Zhou, Shikai Li, Jingyun Liang, Chaohui Yu, Fan Wang, Xiangyang Xue, and Yanwei Fu. Uni3c: Unifying precisely 3d-enhanced camera and human motion controls for video generation. InProceedings of the SIGGRAPH Asia 2025 Conference Papers, pages 1–12, 2025
2025
-
[7]
Sam 3: Segment anything with concepts.arXiv preprint arXiv:2511.16719, 2025
Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al. Sam 3: Segment anything with concepts.arXiv preprint arXiv:2511.16719, 2025
Pith/arXiv arXiv 2025
-
[8]
Pix2video: Video editing using image diffusion
Duygu Ceylan, Chun-Hao P Huang, and Niloy J Mitra. Pix2video: Video editing using image diffusion. InICCV, 2023
2023
-
[9]
Everybody dance now
Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A Efros. Everybody dance now. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5933–5942, 2019
2019
-
[10]
Haoxin Chen, Menghan Xia, Yingqing He, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Jinbo Xing, Yaofang Liu, Qifeng Chen, Xintao Wang, et al. Videocrafter1: Open diffusion models for high-quality video generation.arXiv preprint arXiv:2310.19512, 2023
Pith/arXiv arXiv 2023
-
[11]
Control-a-video: Controllable text-to-video diffusion models with motion prior and reward feedback learning
Weifeng Chen, Yatai Ji, Jie Wu, Hefeng Wu, Pan Xie, Jiashi Li, Xin Xia, Xuefeng Xiao, and Liang Lin. Control-a-video: Controllable text-to-video diffusion models with motion prior and reward feedback learning. InICLR, 2024
2024
-
[12]
Yiyang Chen, Xuanhua He, Xiujun Ma, and Yue Ma. Contextflow: Training-free video object editing via adaptive context enrichment.arXiv preprint arXiv:2509.17818, 2025
arXiv 2025
-
[13]
Yuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen, Jiawei Ren, Yanping Xie, Juan- Manuel Perez-Rua, Bodo Rosenhahn, Tao Xiang, and Sen He. Flatten: optical flow-guided attention for consistent text-to-video editing.arXiv preprint arXiv:2310.05922, 2023. 10
Pith/arXiv arXiv 2023
-
[14]
Scaling rectified flow transform- ers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transform- ers for high-resolution image synthesis. InICML, 2024
2024
-
[15]
Michal Geyer, Omer Bar-Tal, Shai Bagon, and Tali Dekel. Tokenflow: Consistent diffusion features for consistent video editing.arXiv preprint arXiv:2307.10373, 2023
Pith/arXiv arXiv 2023
-
[16]
Videoswap: Customized video subject swapping with interactive semantic point correspondence
Yuchao Gu, Yipin Zhou, Bichen Wu, Licheng Yu, Jia-Wei Liu, Rui Zhao, Jay Zhangjie Wu, David Junhao Zhang, Mike Zheng Shou, and Kevin Tang. Videoswap: Customized video subject swapping with interactive semantic point correspondence. InCVPR, 2023
2023
-
[17]
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. InICLR, 2024
2024
-
[18]
Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, and Ceyuan Yang. Cameractrl: Enabling camera control for text-to-video generation.arXiv preprint arXiv:2404.02101, 2024
Pith/arXiv arXiv 2024
-
[19]
Jonathan Ho, Ajay Jain, and P. Abbeel. Denoising diffusion probabilistic models. InNeurIPS, 2020
2020
-
[20]
Classifier-free diffusion guidance.arXiv preprint arXiv:2207.12598, 2022
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance.arXiv preprint arXiv:2207.12598, 2022
Pith/arXiv arXiv 2022
-
[21]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video diffusion models. InNeurIPS, 2022
2022
-
[22]
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. InICLR, 2023
2023
-
[23]
Zhihao Hu and Dong Xu. Videocontrolnet: A motion-guided video-to-video translation frame- work by using diffusion model with controlnet.arXiv preprint arXiv:2307.14073, 2023
Pith/arXiv arXiv 2023
-
[24]
VBench: Comprehensive benchmark suite for video generative models
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. VBench: Comprehensive benchmark suite for video generative models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024
2024
-
[25]
VBench++: Comprehensive and versatile benchmark suite for video generative models.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
Ziqi Huang, Fan Zhang, Xiaojie Xu, Yinan He, Jiashuo Yu, Ziyue Dong, Qianli Ma, Nattapol Chanpaisit, Chenyang Si, Yuming Jiang, Yaohui Wang, Xinyuan Chen, Ying-Cong Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. VBench++: Comprehensive and versatile benchmark suite for video generative models.IEEE Transactions on Pattern Analysis and Machine Intelli...
2025
-
[26]
Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang, Shuaiwen Leon Song, Samyam Rajbhandari, and Yuxiong He. Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models.arXiv preprint arXiv:2309.14509, 2023
Pith/arXiv arXiv 2023
-
[27]
Vace: All-in- one video creation and editing
Zeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang, Yulin Pan, and Yu Liu. Vace: All-in- one video creation and editing. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 17191–17202, 2025
2025
-
[28]
Rave: Randomized noise shuffling for fast and consistent video editing with diffusion models
Ozgur Kara, Bariscan Kurtkaya, Hidir Yesiltepe, James M Rehg, and Pinar Yanardag. Rave: Randomized noise shuffling for fast and consistent video editing with diffusion models. In CVPR, 2024
2024
-
[29]
Text2video-zero: Text-to-image diffusion models are zero-shot video generators
Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Text2video-zero: Text-to-image diffusion models are zero-shot video generators. InICCV, 2023
2023
-
[30]
Min-Jung Kim, Dongjin Kim, Seokju Yun, and Jaegul Choo. Tv-live: Training-free, text-guided video editing via layer informed vitality exploitation.arXiv preprint arXiv:2506.07205, 2025. 11
Pith/arXiv arXiv 2025
-
[31]
Target-aware video diffusion models
Taeksoo Kim and Hanbyul Joo. Target-aware video diffusion models. InICLR, 2026
2026
-
[32]
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024
Pith/arXiv arXiv 2024
-
[33]
Flowedit: Inversion-free text-based editing using pre-trained flow models
Vladimir Kulikov, Matan Kleiner, Inbar Huberman-Spiegelglas, and Tomer Michaeli. Flowedit: Inversion-free text-based editing using pre-trained flow models. InICCV, 2025
2025
-
[34]
Guangzhao Li, Yanming Yang, Chenxi Song, and Chi Zhang. Flowdirector: Training-free flow steering for precise text-to-video editing.arXiv preprint arXiv:2506.05046, 2025
arXiv 2025
-
[35]
Trackdiffusion: Tracklet-conditioned video generation via diffusion models
Pengxiang Li, Kai Chen, Zhili Liu, Ruiyuan Gao, Lanqing Hong, Dit-Yan Yeung, Huchuan Lu, and Xu Jia. Trackdiffusion: Tracklet-conditioned video generation via diffusion models. In WACV, 2025
2025
-
[36]
Towards an end-to- end framework for flow-guided video inpainting
Zhen Li, Cheng-Ze Lu, Jianhua Qin, Chun-Le Guo, and Ming-Ming Cheng. Towards an end-to- end framework for flow-guided video inpainting. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 17562–17571, 2022
2022
-
[37]
Lipman, Ricky T
Y . Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. InICLR, 2022
2022
-
[38]
Fuseformer: Fusing fine-grained information in transformers for video inpainting
Rui Liu, Hanming Deng, Yangyi Huang, Xiaoyu Shi, Lewei Lu, Wenxiu Sun, Xiaogang Wang, Jifeng Dai, and Hongsheng Li. Fuseformer: Fusing fine-grained information in transformers for video inpainting. InProceedings of the IEEE/CVF international conference on computer vision, pages 14040–14049, 2021
2021
-
[39]
Video-p2p: Video editing with cross-attention control
Shaoteng Liu, Yuechen Zhang, Wenbo Li, Zhe Lin, and Jiaya Jia. Video-p2p: Video editing with cross-attention control. InCVPR, 2024
2024
-
[40]
Flow straight and fast: Learning to generate and transfer data with rectified flow
Xingchao Liu, Chengyue Gong, et al. Flow straight and fast: Learning to generate and transfer data with rectified flow. InICLR, 2023
2023
-
[41]
Follow your pose: Pose-guided text-to-video generation using pose-free videos
Yue Ma, Yingqing He, Xiaodong Cun, Xintao Wang, Siran Chen, Xiu Li, and Qifeng Chen. Follow your pose: Pose-guided text-to-video generation using pose-free videos. InAAAI, 2024
2024
-
[42]
Sdedit: Guided image synthesis and editing with stochastic differential equations
Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. Sdedit: Guided image synthesis and editing with stochastic differential equations. In International Conference on Learning Representations, 2022
2022
-
[43]
Motionflow: Attention-driven motion transfer in video diffusion models
Tuna Han Salih Meral, Hidir Yesiltepe, Connor Dunlop, and Pinar Yanardag. Motionflow: Attention-driven motion transfer in video diffusion models. InProceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 8043–8051, 2026
2026
-
[44]
Dreamix: Video diffusion models are general video editors
Eyal Molad, Eliahu Horwitz, Dani Valevski, Alex Rav Acha, Yossi Matias, Yael Pritch, Yaniv Leviathan, and Yedid Hoshen. Dreamix: Video diffusion models are general video editors. arXiv preprint arXiv:2302.01329, 2023
Pith/arXiv arXiv 2023
-
[45]
Optical-flow guided prompt opti- mization for coherent video generation
Hyelin Nam, Jaemin Kim, Dohun Lee, and Jong Chul Ye. Optical-flow guided prompt opti- mization for coherent video generation. InCVPR, 2025
2025
-
[46]
I2vedit: First-frame-guided video editing via image-to-video diffusion models
Wenqi Ouyang, Yi Dong, Lei Yang, Jianlou Si, and Xingang Pan. I2vedit: First-frame-guided video editing via image-to-video diffusion models. InSIGGRAPH Asia, 2024
2024
-
[47]
Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rombach
Dustin Podell, Zion English, Kyle Lacey, A. Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. InICLR, 2023
2023
-
[48]
Video motion transfer with diffusion transformers
Alexander Pondaven, Aliaksandr Siarohin, Sergey Tulyakov, Philip Torr, and Fabio Pizzati. Video motion transfer with diffusion transformers. InCVPR, 2025
2025
-
[49]
Fatezero: Fusing attentions for zero-shot text-based video editing
Chenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen. Fatezero: Fusing attentions for zero-shot text-based video editing. InICCV, 2023. 12
2023
-
[50]
Blattmann, Dominik Lorenz, Patrick Esser, and B
Robin Rombach, A. Blattmann, Dominik Lorenz, Patrick Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. InCVPR, 2021
2021
-
[51]
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. InNeurIPS, 2022
2022
-
[52]
First order motion model for image animation.Advances in neural information processing systems, 32, 2019
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. First order motion model for image animation.Advances in neural information processing systems, 32, 2019
2019
-
[53]
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make-a-video: Text-to-video generation without text-video data. InICLR, 2023
2023
-
[54]
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. InICLR, 2021
2021
-
[55]
Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Pith/arXiv arXiv 2025
-
[56]
Modelscope text-to-video technical report.arXiv preprint arXiv:2308.06571, 2023
Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report.arXiv preprint arXiv:2308.06571, 2023
Pith/arXiv arXiv 2023
-
[57]
Magicvideo-v2: Multi-stage high-aesthetic video generation.arXiv preprint arXiv:2401.04468, 2024
Weimin Wang, Jiawei Liu, Zhijie Lin, Jiangqiao Yan, Shuo Chen, Chetwin Low, Tuyen Hoang, Jie Wu, Jun Hao Liew, Hanshu Yan, et al. Magicvideo-v2: Multi-stage high-aesthetic video generation.arXiv preprint arXiv:2401.04468, 2024
Pith/arXiv arXiv 2024
-
[58]
Wen Wang, kangyang Xie, Zide Liu, Hao Chen, Yue Cao, Xinlong Wang, and Chunhua Shen. Zero-shot video editing using off-the-shelf image diffusion models.arXiv preprint arXiv:2303.17599, 2023
Pith/arXiv arXiv 2023
-
[59]
Videocomposer: Compositional video synthesis with motion controllability
Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou. Videocomposer: Compositional video synthesis with motion controllability. InNeurIPS, 2023
2023
-
[60]
Videodirector: Precise video editing via text-to-video models
Yukun Wang, Longguang Wang, Zhiyuan Ma, Qibin Hu, Kai Xu, and Yulan Guo. Videodirector: Precise video editing via text-to-video models. InCVPR, 2025
2025
-
[61]
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. InICCV, 2023
2023
-
[62]
Omnivdiff: Omni controllable video diffusion for generation and understanding
Dianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu, Yuchi Huo, Rui Wang, Chi Zhang, and Xuelong Li. Omnivdiff: Omni controllable video diffusion for generation and understanding. InAAAI, 2026
2026
-
[63]
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer.arXiv preprint arXiv:2408.06072, 2024
Pith/arXiv arXiv 2024
-
[64]
Motiondirector: Motion customization of text-to-video diffusion models
Rui Zhao, Yuchao Gu, Jay Zhangjie Wu, David Junhao Zhang, Jia-Wei Liu, Weijia Wu, Jussi Keppo, and Mike Zheng Shou. Motiondirector: Motion customization of text-to-video diffusion models. InEuropean Conference on Computer Vision, pages 273–290. Springer, 2024
2024
-
[65]
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, et al. Pytorch fsdp: experiences on scaling fully sharded data parallel.arXiv preprint arXiv:2304.11277, 2023
Pith/arXiv arXiv 2023
-
[66]
Dian Zheng, Ziqi Huang, Hongbo Liu, Kai Zou, Yinan He, Fan Zhang, Yuanhan Zhang, Jingwen He, Wei-Shi Zheng, Yu Qiao, and Ziwei Liu. VBench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness.arXiv preprint arXiv:2503.21755, 2025. 13
Pith/arXiv arXiv 2025
-
[67]
Posecrafter: One-shot personalized video synthesis following flexible pose control
Yong Zhong, Min Zhao, Zebin You, Xiaofeng Yu, Changwang Zhang, and Chongxuan Li. Posecrafter: One-shot personalized video synthesis following flexible pose control. InECCV, 2024
2024
-
[68]
Propainter: Improving propagation and transformer for video inpainting
Shangchen Zhou, Chongyi Li, Kelvin CK Chan, and Chen Change Loy. Propainter: Improving propagation and transformer for video inpainting. InProceedings of the IEEE/CVF international conference on computer vision, pages 10477–10486, 2023
2023
-
[69]
Few-step flow for 3d generation via marginal-data transport distillation
Zanwei Zhou, Taoran Yi, Jiemin Fang, Chen Yang, Lingxi Xie, Xinggang Wang, Wei Shen, and Qi Tian. Few-step flow for 3d generation via marginal-data transport distillation. InProceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 13853–13861, 2026. A Proof of Proposition 1 In this section, we provide the theoretical proof for Propo...
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.