Pith. sign in

REVIEW 4 major objections 6 minor 69 references

Coarse proxy videos can steer complex motion in pretrained video generators without any training, by noising latents region-wise and relaxing them onto the model manifold.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 00:18 UTC pith:T5QYNQM5

load-bearing objection Clean training-free recipe for proxy-as-dynamics; finite-K SFR is the real mechanism and the asymptotic proof does not explain it. the 4 major comments →

arxiv 2607.03732 v1 pith:T5QYNQM5 submitted 2026-07-04 cs.CV

ProxyUp: Training-Free Proxy-Conditioned Video Generation for Controllable Dynamics

classification cs.CV
keywords proxy-conditioned video generationtraining-free video synthesisregion-wise latent noisingstochastic flow relaxationcontrollable dynamicsrectified flowmotion transfervideo editing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Modern text-to-video models still struggle to specify fine-grained, physically plausible motion and interactions from language alone. This paper introduces proxy-conditioned video generation: a coarse video from simulation or real recording acts only as a dynamics carrier for foreground motion, while a text prompt drives new content and scene interactions. Because paired proxy–target videos are hard to collect, the authors propose ProxyUp, a training-free pipeline on pretrained flow-based video models. It inverts the proxy, keeps motion-critical latents in a masked region, injects noise elsewhere for regeneration, then runs Stochastic Flow Relaxation to pull the hybrid latent toward the model’s learned distribution before ODE sampling. On both physics-simulation and real-world proxies, the method improves dynamic fidelity and text alignment over strong editing and motion-transfer baselines.

Core claim

The paper claims that a training-free combination of region-wise latent noising and Stochastic Flow Relaxation lets a pretrained video generator preserve essential dynamics from a coarse proxy video while synthesizing novel, prompt-aligned content and plausible foreground–background interactions—outperforming video editing and motion-transfer baselines on dynamic fidelity and text alignment for both simulated and real proxies.

What carries the argument

Region-wise latent noising plus Stochastic Flow Relaxation (SFR): invert the masked proxy to an intermediate noise level, keep those latents in motion-critical regions while replacing the rest with matched noise, then iteratively denoise and re-noise the hybrid latent so it approaches the model’s in-distribution manifold before deterministic ODE sampling.

Load-bearing premise

That a finite number of SFR re-noising rounds is enough to move a hand-composed, out-of-distribution latent onto the pretrained model’s manifold so that sampling invents coherent interactions the base model already knows how to draw.

What would settle it

If, on the same backbone and equal or larger inference budget, a region-wise-noised SDEdit or editing baseline matches or exceeds ProxyUp on Motion Rationality and Mechanics for the bread-cutting and curtain-pulling proxies (and similar held-out dynamics), the claimed benefit of SFR would not hold.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Physics simulations and casual real recordings can be reused as motion controllers for open-ended text-driven video synthesis without collecting paired training data.
  • Video editing and motion-transfer pipelines that stay anchored to source appearance are not the right tools when the source is only a dynamics prior.
  • A modest inference-time relaxation loop can repair hybrid latents that would otherwise break foreground–background coupling.
  • Task-specific proxy–prompt–mask evaluation sets become necessary to measure dynamics-preserving regeneration rather than pure editing or pure generation.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same recipe could turn low-fidelity game or robotics rollouts into large synthetic video corpora with controlled physics and varied visual styles.
  • When the base generator lacks the needed interaction priors, better proxies alone will not fix failures—pointing toward models trained on richer physical contact data.
  • Automatic or learned masks and force-application cues would reduce reliance on SAM-style foreground masks and hand-chosen t_init and K.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces proxy-conditioned video generation: a coarse proxy video (simulation or real recording) supplies foreground dynamics while a text prompt specifies novel content and interactions. Because paired proxy–target data are scarce, the authors propose ProxyUp, a training-free pipeline on pretrained rectified-flow video models (primarily Wan2.2). ProxyUp (i) ODE-inverts the masked proxy foreground to an intermediate noise level, (ii) composes a hybrid latent by region-wise latent noising (Eq. 5), (iii) applies Stochastic Flow Relaxation (SFR; Sec. 4.4, Alg. 1) to pull the hand-composed latent toward the model manifold, and (iv) finishes with deterministic ODE sampling. On a 76-clip custom set (simulation + real), ProxyUp reports gains over editing, inpainting, and motion-transfer baselines on Imaging Quality, Motion Rationality, Mechanics, and Material (Table 1), with qualitative and ablation support (Figs. 4–5).

Significance. If the empirical claims hold, the work offers a practical, training-free route to controllable dynamics that text alone cannot specify, and a useful intermediate between pure T2V, video editing, and motion transfer. Strengths include a clear three-stage pipeline, an explicit algorithm, an ablation isolating RLN and SFR (Fig. 5), hyperparameter sweeps (Fig. 6), cross-backbone checks in the appendix, and a formal (if asymptotic) Markov argument for SFR (Prop. 1, App. A). The task framing and the idea of using low-fidelity proxies as dynamics carriers are timely for physics-aware video generation. Significance is tempered by a small custom evaluation set, author-written VBench-style QA criteria, and a theory–practice gap on finite-K dynamics retention.

major comments (4)
  1. [Sec. 4.4, Prop. 1, App. A, Alg. 1] Sec. 4.4 / Prop. 1 / App. A: Proposition 1 only shows that the SFR Markov kernel is ergodic and D_KL(q_K || π_ID) → 0 as K → ∞ under a well-trained velocity field. That limit is pure text-conditioned sampling at t_init and would erase the inverted proxy structure. No mask is re-applied inside the SFR loop (Alg. 1 lines 6–10), and no finite-K mixing-time or information-retention bound is given. Dynamics preservation at the operating point K=15, t_init=0.922 (s=0.8) is therefore an empirical incomplete-mixing effect, not a consequence of the stated theory. The central mechanism claim—that RLN+SFR jointly preserve proxy dynamics while restoring fg–bg coupling—needs either a finite-K analysis (e.g., how much inverted foreground signal remains after K steps) or a clear reframing that SFR is a practical regularizer whose dynamics retention is empirical.
  2. [Sec. 5.1, Table 1, App. C.2–C.3] Sec. 5.1 / Table 1 / App. C.2–C.3: The evaluation set has only 76 custom clips with author-written multi-question criteria for MR, Mech., and Mat. conditioned on the same proxy/prompt pairs used for generation. This is acceptable for a new task but is load-bearing for the claim of consistent outperformance in dynamic fidelity. Please (i) release the full metric prompts and scoring protocol as promised, (ii) report inter-annotator or multi-run variance, and (iii) add at least one external or human preference study on dynamics fidelity vs. text alignment so that Table 1 is not solely self-defined QA.
  3. [Sec. 5.2, Table 1, App. C.4] Sec. 5.2 / App. C.4: Several baselines run on different backbones and default schedules (DiTFlow on CogVideoX-5B; FlowDirector on Wan2.1; VACE 14B). Appendix cross-backbone checks (Figs. 9–10) and the SDEdit step-budget study (Fig. 8) help, but the main Table 1 still mixes generators. For the primary comparison, either re-run the strongest motion-transfer and editing baselines on the same Wan2.2 backbone used by ProxyUp, or report a backbone-matched subset as the headline table so gains on MR/Mech. can be attributed to the method rather than model capacity.
  4. [Sec. 5.4, Fig. 6] Sec. 5.4 / Fig. 6: Free parameters K, t_init (strength s), and CFG scales are chosen by qualitative inspection. Given that the skeptic concern is precisely the incomplete-mixing regime, please quantify the trade-off: e.g., proxy-motion metrics (optical-flow or keypoint correlation in the masked region) vs. K and s, not only visual examples. Without this, it is hard to know how fragile the reported MR/Mech. gains are to hyperparameter choice.
minor comments (6)
  1. [Sec. 3, Fig. 2] Fig. 2 caption and body: the preliminary analysis is helpful; please state the exact strength/t values and masks used for each baseline so the trade-off narrative is reproducible.
  2. [Sec. 4.3, Eq. (4)] Eq. (4): the background noise variance ((1−t_init)^2 + t_init^2)I is nonstandard relative to the linear path Z_t=(1−t)Z_0+tZ_1; a one-sentence justification (marginal variance of the interpolation) would help readers.
  3. [Sec. 2.1] Related work (Sec. 2.1) mentions physics-related conditions but cites little recent simulation-to-video or physics-prior work; a few additional pointers would better situate proxy videos among existing control signals.
  4. [Sec. 6, Fig. 11] Limitation section (Sec. 6) and Fig. 11 are candid; consider moving one failure case into the main paper so readers see the dependence on the base model’s physical prior without opening the appendix.
  5. [Sec. 4–5] Notation: strength s is defined in a footnote and reused as s=0.8; define it once in the main text near t_init for clarity.
  6. [Throughout] Minor typos / consistency: “V ACE” spacing in Table 1 and captions; “out-of-distribution (OOD)latent” missing space (Sec. 4.2); arXiv id and “Preprint” header are fine for review but should be cleaned for camera-ready.

Circularity Check

0 steps flagged

No circular derivation: ProxyUp is an empirical training-free pipeline; Prop. 1 is a standard asymptotic Markov/DPI sketch, not a tautology that forces the reported metrics.

full rationale

The paper’s load-bearing claims are (i) a constructive inference procedure (region-wise latent noising + finite-K SFR + ODE sampling) and (ii) empirical outperformance on a task-specific set (Table 1, Figs. 4–5). Neither reduces to its inputs by definition. Region-wise composition (Eq. 5) is a hand-built hybrid latent, not a quantity fitted from the evaluation targets. SFR (Eqs. 6–7, Alg. 1) is a re-noising loop whose asymptotic claim (Prop. 1 / App. A) is a standard ergodicity + Data Processing Inequality argument that lim K→∞ D_KL(q_K ∥ π_ID)=0 under a well-trained velocity field; that argument does not algebraically force finite-K=15 dynamics retention or the MR/Mech. scores. Hyperparameters (K=15, s=0.8) are chosen by qualitative inspection (Fig. 6), which is empirical tuning, not a fitted-input-called-prediction. Baselines are external methods run under their defaults; metrics (IQ/MR/Mech./Mat. adapted from VBench with proxy-conditioned QA criteria) score generated videos after the fact and are not satisfied by construction of Eq. 5. Self-citations (e.g., [69]) are not used as uniqueness theorems or load-bearing premises for the central mechanism. Gaps between asymptotic theory and finite-K practice are correctness/validity concerns, not circularity. Derivation chain is self-contained against external benchmarks; score 0.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 3 invented entities

The method rests on a pretrained rectified-flow video prior, hand-chosen noise/relaxation hyperparameters, and a theoretical claim that finite SFR iterations pull hybrid latents onto the learned manifold. No new physical constants; the main invented machinery is the SFR loop and the proxy-conditioned task framing.

free parameters (3)
  • SFR iterations K = 15
    Default K=15 chosen empirically to balance quality, dynamics, and cost (§5.4, Fig. 6); central quality claims depend on this choice.
  • initial noise level t_init / strength s = s=0.8 (t_init=0.922)
    t_init=0.922 (s=0.8) selected by hand to keep foreground dynamics while allowing background regeneration (§5.4); too low or high breaks the method.
  • CFG scales (inversion / SFR / sampling) = 1 / 3 / 4&3
    CFG set to 1 (inversion), 3 (SFR), and Wan2.2 defaults (4/3) during sampling (App. C.1); affects text alignment and visual quality.
axioms (4)
  • domain assumption The pretrained rectified-flow video model’s velocity field is sufficiently accurate that one-step Euler denoising approximates the in-distribution posterior mean (Tweedie / flow matching).
    Invoked in §4.4 and Prop. 1 proof (App. A) to justify SFR’s denoise step.
  • standard math Re-noising with positive Gaussian variance yields an ergodic Markov kernel whose unique stationary distribution is the model’s marginal at t_init.
    Used in App. A to claim lim D_KL(q_K || π_ID)=0; standard Markov theory under positivity assumptions.
  • domain assumption A binary foreground mask M correctly isolates motion-critical regions of the proxy so that preserving M⊙Z_inv retains the intended dynamics.
    Problem formulation §4.2–4.3; masks from rendering or SAM3 (App. C.2).
  • domain assumption The base generator already encodes enough physical interaction knowledge to complete interactions absent from the proxy once the latent is on-manifold.
    Stated as a limitation in §6 and failure cases (App. E); load-bearing for ‘plausible interactions’ claims.
invented entities (3)
  • proxy-conditioned video generation (task) no independent evidence
    purpose: Frame proxy videos as pure dynamics carriers rather than appearance references for prompt-driven regeneration.
    Introduced in abstract and §1 as a setting distinct from editing/motion transfer; evaluated on a custom set.
  • Stochastic Flow Relaxation (SFR) no independent evidence
    purpose: Iteratively relax hybrid OOD latents toward the flow model’s learned distribution before ODE sampling.
    Core algorithmic contribution §4.4; Prop. 1 offers a KL-convergence sketch but no external validation outside this paper’s ablations.
  • region-wise latent noising no independent evidence
    purpose: Compose inverted foreground latents with background noise at t_init to decouple motion from proxy appearance.
    §4.3; related to SDEdit-style noising but region-masked; no independent external evidence beyond paper experiments.

pith-pipeline@v1.1.0-grok45 · 19773 in / 3380 out tokens · 28894 ms · 2026-07-12T00:18:51.928688+00:00 · methodology

0 comments
read the original abstract

Precise control over complex dynamics remains challenging for modern video generative models, as text prompts alone often cannot specify physically plausible, fine-grained motion and interactions. We introduce $\textit{proxy-conditioned video generation}$, where a coarse proxy video from physics-based simulation or real-world recording serves as a dynamics carrier to control foreground object motion. Given a proxy video and a text prompt, the goal is to synthesize a new video that preserves the proxy dynamics while generating novel content and plausible interactions aligned with the prompt. Since paired proxy-target videos are difficult to obtain, we propose $\textbf{ProxyUp}$, a training-free framework built on pretrained video generative models. ProxyUp first inverts the proxy video into an intermediate latent representation and applies $\textbf{region-wise latent noising}$, preserving motion-critical proxy latents while injecting noise into regions intended for text-driven regeneration. To mitigate the distribution mismatch and weak foreground-background coupling introduced by this heuristic latent composition, we further propose $\textbf{Stochastic Flow Relaxation (SFR)}$, which progressively relaxes the composed latent toward the model's learned distribution before ODE sampling. Experiments on both simulation and real-world proxies show that ProxyUp outperforms strong video editing and motion transfer baselines in dynamic fidelity and text alignment.

Figures

Figures reproduced from arXiv: 2607.03732 by Chen Yang, Fanpeng Meng, Jiazhong Cen, Jiemin Fang, Qi Tian, Sikuang Li, Wei Shen, Yumeng He, Zanwei Zhou, Zhikuan Bao.

Figure 1
Figure 1. Figure 1: Given a proxy video (middle) as a dynamics carrier, [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Preliminary analysis of proxy-video-guided generation using existing pipelines. We evaluate [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Overview of ProxyUp. Given a proxy video as guidance, we first apply Region-wise Latent [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Qualitative comparison results. Methods presented below the proxy video use the same [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Ablation study on the key components of ProxyUp. Without region-wise latent noising [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Hyperparameter analysis on the strength s and the number of SFR steps K [PITH_FULL_IMAGE:figures/full_fig_p017_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Given the same proxy video, ProxyUp can generate diverse videos conditioned on different [PITH_FULL_IMAGE:figures/full_fig_p017_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Qualitative results for increasing the inference budget of SDEdit. From top to bottom, [PITH_FULL_IMAGE:figures/full_fig_p018_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Cross-backbone qualitative results of ProxyUp on Wan2.1. The grid shows uniformly [PITH_FULL_IMAGE:figures/full_fig_p018_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Cross-backbone qualitative results of DiTFlow on Wan2.2. Rows alternate between proxy [PITH_FULL_IMAGE:figures/full_fig_p019_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Failure cases of ProxyUp. When the proxy video contains uncommon or out-of-distribution [PITH_FULL_IMAGE:figures/full_fig_p019_11.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

69 extracted references · 19 linked inside Pith

  1. [1]

    Recammaster: Camera-controlled generative rendering from a single video

    Jianhong Bai, Menghan Xia, Xiao Fu, Xintao Wang, Lianrui Mu, Jinwen Cao, Zuozhu Liu, Haoji Hu, Xiang Bai, Pengfei Wan, et al. Recammaster: Camera-controlled generative rendering from a single video. InICCV, 2025

  2. [2]

    Syncammaster: Synchronizing multi-camera video generation from diverse viewpoints

    Jianhong Bai, Menghan Xia, Xintao Wang, Ziyang Yuan, Xiao Fu, Zuozhu Liu, Haoji Hu, Pengfei Wan, and Di Zhang. Syncammaster: Synchronizing multi-camera video generation from diverse viewpoints. InICLR, 2025

  3. [3]

    Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127, 2023

    Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Do- minik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127, 2023

  4. [4]

    Align your latents: High-resolution video synthesis with latent diffusion models

    Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. InCVPR, 2023

  5. [5]

    Video generation models as world simulators

    Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Leo Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al. Video generation models as world simulators. OpenAI Blog, 2024

  6. [6]

    Uni3c: Unifying precisely 3d-enhanced camera and human motion controls for video generation

    Chenjie Cao, Jingkai Zhou, Shikai Li, Jingyun Liang, Chaohui Yu, Fan Wang, Xiangyang Xue, and Yanwei Fu. Uni3c: Unifying precisely 3d-enhanced camera and human motion controls for video generation. InProceedings of the SIGGRAPH Asia 2025 Conference Papers, pages 1–12, 2025

  7. [7]

    Sam 3: Segment anything with concepts.arXiv preprint arXiv:2511.16719, 2025

    Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al. Sam 3: Segment anything with concepts.arXiv preprint arXiv:2511.16719, 2025

  8. [8]

    Pix2video: Video editing using image diffusion

    Duygu Ceylan, Chun-Hao P Huang, and Niloy J Mitra. Pix2video: Video editing using image diffusion. InICCV, 2023

  9. [9]

    Everybody dance now

    Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A Efros. Everybody dance now. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5933–5942, 2019

  10. [10]

    Videocrafter1: Open diffusion models for high-quality video generation.arXiv preprint arXiv:2310.19512, 2023

    Haoxin Chen, Menghan Xia, Yingqing He, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Jinbo Xing, Yaofang Liu, Qifeng Chen, Xintao Wang, et al. Videocrafter1: Open diffusion models for high-quality video generation.arXiv preprint arXiv:2310.19512, 2023

  11. [11]

    Control-a-video: Controllable text-to-video diffusion models with motion prior and reward feedback learning

    Weifeng Chen, Yatai Ji, Jie Wu, Hefeng Wu, Pan Xie, Jiashi Li, Xin Xia, Xuefeng Xiao, and Liang Lin. Control-a-video: Controllable text-to-video diffusion models with motion prior and reward feedback learning. InICLR, 2024

  12. [12]

    Contextflow: Training-free video object editing via adaptive context enrichment.arXiv preprint arXiv:2509.17818, 2025

    Yiyang Chen, Xuanhua He, Xiujun Ma, and Yue Ma. Contextflow: Training-free video object editing via adaptive context enrichment.arXiv preprint arXiv:2509.17818, 2025

  13. [13]

    Flatten: optical flow-guided attention for consistent text-to-video editing.arXiv preprint arXiv:2310.05922, 2023

    Yuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen, Jiawei Ren, Yanping Xie, Juan- Manuel Perez-Rua, Bodo Rosenhahn, Tao Xiang, and Sen He. Flatten: optical flow-guided attention for consistent text-to-video editing.arXiv preprint arXiv:2310.05922, 2023. 10

  14. [14]

    Scaling rectified flow transform- ers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transform- ers for high-resolution image synthesis. InICML, 2024

  15. [15]

    Tokenflow: Consistent diffusion features for consistent video editing.arXiv preprint arXiv:2307.10373, 2023

    Michal Geyer, Omer Bar-Tal, Shai Bagon, and Tali Dekel. Tokenflow: Consistent diffusion features for consistent video editing.arXiv preprint arXiv:2307.10373, 2023

  16. [16]

    Videoswap: Customized video subject swapping with interactive semantic point correspondence

    Yuchao Gu, Yipin Zhou, Bichen Wu, Licheng Yu, Jia-Wei Liu, Rui Zhao, Jay Zhangjie Wu, David Junhao Zhang, Mike Zheng Shou, and Kevin Tang. Videoswap: Customized video subject swapping with interactive semantic point correspondence. InCVPR, 2023

  17. [17]

    Animatediff: Animate your personalized text-to-image diffusion models without specific tuning

    Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. InICLR, 2024

  18. [18]

    Cameractrl: Enabling camera control for text-to-video generation.arXiv preprint arXiv:2404.02101, 2024

    Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, and Ceyuan Yang. Cameractrl: Enabling camera control for text-to-video generation.arXiv preprint arXiv:2404.02101, 2024

  19. [19]

    Jonathan Ho, Ajay Jain, and P. Abbeel. Denoising diffusion probabilistic models. InNeurIPS, 2020

  20. [20]

    Classifier-free diffusion guidance.arXiv preprint arXiv:2207.12598, 2022

    Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance.arXiv preprint arXiv:2207.12598, 2022

  21. [21]

    Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video diffusion models. InNeurIPS, 2022

  22. [22]

    Cogvideo: Large-scale pretraining for text-to-video generation via transformers

    Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. InICLR, 2023

  23. [23]

    Videocontrolnet: A motion-guided video-to-video translation frame- work by using diffusion model with controlnet.arXiv preprint arXiv:2307.14073, 2023

    Zhihao Hu and Dong Xu. Videocontrolnet: A motion-guided video-to-video translation frame- work by using diffusion model with controlnet.arXiv preprint arXiv:2307.14073, 2023

  24. [24]

    VBench: Comprehensive benchmark suite for video generative models

    Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. VBench: Comprehensive benchmark suite for video generative models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024

  25. [25]

    VBench++: Comprehensive and versatile benchmark suite for video generative models.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

    Ziqi Huang, Fan Zhang, Xiaojie Xu, Yinan He, Jiashuo Yu, Ziyue Dong, Qianli Ma, Nattapol Chanpaisit, Chenyang Si, Yuming Jiang, Yaohui Wang, Xinyuan Chen, Ying-Cong Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. VBench++: Comprehensive and versatile benchmark suite for video generative models.IEEE Transactions on Pattern Analysis and Machine Intelli...

  26. [26]

    Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models.arXiv preprint arXiv:2309.14509, 2023

    Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang, Shuaiwen Leon Song, Samyam Rajbhandari, and Yuxiong He. Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models.arXiv preprint arXiv:2309.14509, 2023

  27. [27]

    Vace: All-in- one video creation and editing

    Zeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang, Yulin Pan, and Yu Liu. Vace: All-in- one video creation and editing. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 17191–17202, 2025

  28. [28]

    Rave: Randomized noise shuffling for fast and consistent video editing with diffusion models

    Ozgur Kara, Bariscan Kurtkaya, Hidir Yesiltepe, James M Rehg, and Pinar Yanardag. Rave: Randomized noise shuffling for fast and consistent video editing with diffusion models. In CVPR, 2024

  29. [29]

    Text2video-zero: Text-to-image diffusion models are zero-shot video generators

    Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Text2video-zero: Text-to-image diffusion models are zero-shot video generators. InICCV, 2023

  30. [30]

    Tv-live: Training-free, text-guided video editing via layer informed vitality exploitation.arXiv preprint arXiv:2506.07205, 2025

    Min-Jung Kim, Dongjin Kim, Seokju Yun, and Jaegul Choo. Tv-live: Training-free, text-guided video editing via layer informed vitality exploitation.arXiv preprint arXiv:2506.07205, 2025. 11

  31. [31]

    Target-aware video diffusion models

    Taeksoo Kim and Hanbyul Joo. Target-aware video diffusion models. InICLR, 2026

  32. [32]

    Hunyuanvideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024

    Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024

  33. [33]

    Flowedit: Inversion-free text-based editing using pre-trained flow models

    Vladimir Kulikov, Matan Kleiner, Inbar Huberman-Spiegelglas, and Tomer Michaeli. Flowedit: Inversion-free text-based editing using pre-trained flow models. InICCV, 2025

  34. [34]

    Flowdirector: Training-free flow steering for precise text-to-video editing.arXiv preprint arXiv:2506.05046, 2025

    Guangzhao Li, Yanming Yang, Chenxi Song, and Chi Zhang. Flowdirector: Training-free flow steering for precise text-to-video editing.arXiv preprint arXiv:2506.05046, 2025

  35. [35]

    Trackdiffusion: Tracklet-conditioned video generation via diffusion models

    Pengxiang Li, Kai Chen, Zhili Liu, Ruiyuan Gao, Lanqing Hong, Dit-Yan Yeung, Huchuan Lu, and Xu Jia. Trackdiffusion: Tracklet-conditioned video generation via diffusion models. In WACV, 2025

  36. [36]

    Towards an end-to- end framework for flow-guided video inpainting

    Zhen Li, Cheng-Ze Lu, Jianhua Qin, Chun-Le Guo, and Ming-Ming Cheng. Towards an end-to- end framework for flow-guided video inpainting. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 17562–17571, 2022

  37. [37]

    Lipman, Ricky T

    Y . Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. InICLR, 2022

  38. [38]

    Fuseformer: Fusing fine-grained information in transformers for video inpainting

    Rui Liu, Hanming Deng, Yangyi Huang, Xiaoyu Shi, Lewei Lu, Wenxiu Sun, Xiaogang Wang, Jifeng Dai, and Hongsheng Li. Fuseformer: Fusing fine-grained information in transformers for video inpainting. InProceedings of the IEEE/CVF international conference on computer vision, pages 14040–14049, 2021

  39. [39]

    Video-p2p: Video editing with cross-attention control

    Shaoteng Liu, Yuechen Zhang, Wenbo Li, Zhe Lin, and Jiaya Jia. Video-p2p: Video editing with cross-attention control. InCVPR, 2024

  40. [40]

    Flow straight and fast: Learning to generate and transfer data with rectified flow

    Xingchao Liu, Chengyue Gong, et al. Flow straight and fast: Learning to generate and transfer data with rectified flow. InICLR, 2023

  41. [41]

    Follow your pose: Pose-guided text-to-video generation using pose-free videos

    Yue Ma, Yingqing He, Xiaodong Cun, Xintao Wang, Siran Chen, Xiu Li, and Qifeng Chen. Follow your pose: Pose-guided text-to-video generation using pose-free videos. InAAAI, 2024

  42. [42]

    Sdedit: Guided image synthesis and editing with stochastic differential equations

    Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. Sdedit: Guided image synthesis and editing with stochastic differential equations. In International Conference on Learning Representations, 2022

  43. [43]

    Motionflow: Attention-driven motion transfer in video diffusion models

    Tuna Han Salih Meral, Hidir Yesiltepe, Connor Dunlop, and Pinar Yanardag. Motionflow: Attention-driven motion transfer in video diffusion models. InProceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 8043–8051, 2026

  44. [44]

    Dreamix: Video diffusion models are general video editors

    Eyal Molad, Eliahu Horwitz, Dani Valevski, Alex Rav Acha, Yossi Matias, Yael Pritch, Yaniv Leviathan, and Yedid Hoshen. Dreamix: Video diffusion models are general video editors. arXiv preprint arXiv:2302.01329, 2023

  45. [45]

    Optical-flow guided prompt opti- mization for coherent video generation

    Hyelin Nam, Jaemin Kim, Dohun Lee, and Jong Chul Ye. Optical-flow guided prompt opti- mization for coherent video generation. InCVPR, 2025

  46. [46]

    I2vedit: First-frame-guided video editing via image-to-video diffusion models

    Wenqi Ouyang, Yi Dong, Lei Yang, Jianlou Si, and Xingang Pan. I2vedit: First-frame-guided video editing via image-to-video diffusion models. InSIGGRAPH Asia, 2024

  47. [47]

    Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rombach

    Dustin Podell, Zion English, Kyle Lacey, A. Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. InICLR, 2023

  48. [48]

    Video motion transfer with diffusion transformers

    Alexander Pondaven, Aliaksandr Siarohin, Sergey Tulyakov, Philip Torr, and Fabio Pizzati. Video motion transfer with diffusion transformers. InCVPR, 2025

  49. [49]

    Fatezero: Fusing attentions for zero-shot text-based video editing

    Chenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen. Fatezero: Fusing attentions for zero-shot text-based video editing. InICCV, 2023. 12

  50. [50]

    Blattmann, Dominik Lorenz, Patrick Esser, and B

    Robin Rombach, A. Blattmann, Dominik Lorenz, Patrick Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. InCVPR, 2021

  51. [51]

    Photorealistic text-to-image diffusion models with deep language understanding

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. InNeurIPS, 2022

  52. [52]

    First order motion model for image animation.Advances in neural information processing systems, 32, 2019

    Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. First order motion model for image animation.Advances in neural information processing systems, 32, 2019

  53. [53]

    Make-a-video: Text-to-video generation without text-video data

    Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make-a-video: Text-to-video generation without text-video data. InICLR, 2023

  54. [54]

    Score-based generative modeling through stochastic differential equations

    Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. InICLR, 2021

  55. [55]

    Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025

    Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025

  56. [56]

    Modelscope text-to-video technical report.arXiv preprint arXiv:2308.06571, 2023

    Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report.arXiv preprint arXiv:2308.06571, 2023

  57. [57]

    Magicvideo-v2: Multi-stage high-aesthetic video generation.arXiv preprint arXiv:2401.04468, 2024

    Weimin Wang, Jiawei Liu, Zhijie Lin, Jiangqiao Yan, Shuo Chen, Chetwin Low, Tuyen Hoang, Jie Wu, Jun Hao Liew, Hanshu Yan, et al. Magicvideo-v2: Multi-stage high-aesthetic video generation.arXiv preprint arXiv:2401.04468, 2024

  58. [58]

    Zero-shot video editing using off-the-shelf image diffusion models.arXiv preprint arXiv:2303.17599, 2023

    Wen Wang, kangyang Xie, Zide Liu, Hao Chen, Yue Cao, Xinlong Wang, and Chunhua Shen. Zero-shot video editing using off-the-shelf image diffusion models.arXiv preprint arXiv:2303.17599, 2023

  59. [59]

    Videocomposer: Compositional video synthesis with motion controllability

    Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou. Videocomposer: Compositional video synthesis with motion controllability. InNeurIPS, 2023

  60. [60]

    Videodirector: Precise video editing via text-to-video models

    Yukun Wang, Longguang Wang, Zhiyuan Ma, Qibin Hu, Kai Xu, and Yulan Guo. Videodirector: Precise video editing via text-to-video models. InCVPR, 2025

  61. [61]

    Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

    Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. InICCV, 2023

  62. [62]

    Omnivdiff: Omni controllable video diffusion for generation and understanding

    Dianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu, Yuchi Huo, Rui Wang, Chi Zhang, and Xuelong Li. Omnivdiff: Omni controllable video diffusion for generation and understanding. InAAAI, 2026

  63. [63]

    Cogvideox: Text-to-video diffusion models with an expert transformer.arXiv preprint arXiv:2408.06072, 2024

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer.arXiv preprint arXiv:2408.06072, 2024

  64. [64]

    Motiondirector: Motion customization of text-to-video diffusion models

    Rui Zhao, Yuchao Gu, Jay Zhangjie Wu, David Junhao Zhang, Jia-Wei Liu, Weijia Wu, Jussi Keppo, and Mike Zheng Shou. Motiondirector: Motion customization of text-to-video diffusion models. InEuropean Conference on Computer Vision, pages 273–290. Springer, 2024

  65. [65]

    Pytorch fsdp: experiences on scaling fully sharded data parallel.arXiv preprint arXiv:2304.11277, 2023

    Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, et al. Pytorch fsdp: experiences on scaling fully sharded data parallel.arXiv preprint arXiv:2304.11277, 2023

  66. [66]

    VBench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness.arXiv preprint arXiv:2503.21755, 2025

    Dian Zheng, Ziqi Huang, Hongbo Liu, Kai Zou, Yinan He, Fan Zhang, Yuanhan Zhang, Jingwen He, Wei-Shi Zheng, Yu Qiao, and Ziwei Liu. VBench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness.arXiv preprint arXiv:2503.21755, 2025. 13

  67. [67]

    Posecrafter: One-shot personalized video synthesis following flexible pose control

    Yong Zhong, Min Zhao, Zebin You, Xiaofeng Yu, Changwang Zhang, and Chongxuan Li. Posecrafter: One-shot personalized video synthesis following flexible pose control. InECCV, 2024

  68. [68]

    Propainter: Improving propagation and transformer for video inpainting

    Shangchen Zhou, Chongyi Li, Kelvin CK Chan, and Chen Change Loy. Propainter: Improving propagation and transformer for video inpainting. InProceedings of the IEEE/CVF international conference on computer vision, pages 10477–10486, 2023

  69. [69]

    Few-step flow for 3d generation via marginal-data transport distillation

    Zanwei Zhou, Taoran Yi, Jiemin Fang, Chen Yang, Lingxi Xie, Xinggang Wang, Wei Shen, and Qi Tian. Few-step flow for 3d generation via marginal-data transport distillation. InProceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 13853–13861, 2026. A Proof of Proposition 1 In this section, we provide the theoretical proof for Propo...