REVIEW 4 major objections 3 minor 69 references
Extension of generalized KYP lemma: from LTI systems to LPV systems
T0 review · 4 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The generalized KYP lemma extends to LPV systems only after the frequency range is widened to restore a time-domain inequality.
desk verdict The abstract advertises a control theory result, but the body is an unrelated computer vision paper, so there is no mathematical content to referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the frequency-dependent integral quadratic constraint (IQC) function, which gives the gKYP lemma its time-domain reading by weighting the input spectrum over the frequency range of interest. For LTI systems the lemma works because this function is non-negative over the band; the paper's reformulation replaces the original frequency range $\Omega$ with an enlarged set $\widetilde{\Omega}$ so that the IQC becomes non-negative for the LPV system. The minimal enlargement is claimed to be governed by the gap between the system poles and $\Omega$, together with controllability Gramians, which are the matrices that measure how strongly the inputs reach the system's states and make the enlargement computable.
What would settle it
Take a scalar or low-order LPV system with one pole at distance $d$ from the boundary of a chosen frequency band $\Omega$ and compute the frequency-dependent IQC function in closed form; the claimed pole–gap and Gramian formula predicts the minimal widening needed to restore non-negativity. If the true minimal widening differs, or if some LPV system admits no finite widening that preserves the original band's behavior, the extension is refuted.
Extended reading notes
Core claim
The central claim is that the obstruction to extending the gKYP lemma from LTI to LPV systems is the non-negativity of the frequency-dependent IQC function: the intermodulation between the input signal and the time-varying scheduling parameter can drive this function negative, invalidating existing results that assume it stays non-negative. The paper proposes to repair the failure by enlarging the originally specified frequency range, and asserts that the minimal such enlargement depends on the interaction (gap) between the system poles and that range, together with a set of controllability Gramians. On this basis it states an extension of the gKYP lemma for LPV systems, enabling finite-frequency analysis in a direct and reliable manner.
Load-bearing premise
The repair depends on the premise that some computable enlargement of the frequency range always restores the IQC non-negativity while still capturing the original finite-frequency behavior, and the submitted full text does not contain the proof of this premise or of the claimed pole–gap and Gramian dependence.
Editorial extensions
If this is right
- Finite-frequency analysis of LPV systems can be carried out directly in the frequency domain, without first assuming the IQC non-negativity that fails in the LPV setting.
- The enlargement criterion gives a computable rule: how far the frequency range must grow is read off from pole locations and controllability Gramians.
- The counterexample shows that any proposed LPV extension of the gKYP lemma that silently keeps the LTI non-negativity assumption is unsound.
- The extended lemma makes frequency-bounded performance analysis of LPV plants a matter of verifying an enlarged-range inequality, in the same spirit as the LTI case.
Reading between the lines
- A testable consequence the authors do not spell out: as the pole–frequency gap grows, the required enlargement should shrink toward zero, so systems with poles far from the band should recover the LTI behavior.
- The dependence on controllability Gramians suggests the enlargement can be computed numerically by solving Lyapunov-type equations restricted to the controllable subspace, which would give a concrete algorithm beyond the existence statement.
- If the extension holds, LTI finite-frequency synthesis tools could be adapted to LPV plants by solving the enlarged-range inequality, with the pole–gap formula quantifying the added conservatism.
- A closed-form scalar LPV example with a single pole near the band boundary could test whether the minimal enlargement really matches the claimed pole–gap and Gramian formula.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The abstract of arXiv:2508.08579 announces an extension of the generalized KYP lemma from LTI to LPV systems. It claims (i) a counterexample showing that the frequency-dependent IQC non-negativity property can fail for LPV systems; (ii) a reformulation that enlarges the original frequency range to restore non-negativity; (iii) a characterization of the minimal required expansion in terms of the pole-frequency gap and a set of controllability Gramians; and (iv) numerical examples demonstrating the potential and efficiency of the approach. The full text supplied for review, however, is the paper 'RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space,' a computer-vision paper on controllable human video generation. None of the mathematical claims in the abstract appears in the body: there is no definition of the IQC function, no LPV model class, no theorem statement, no proof, and no numerical experiment related to finite-frequency analysis.
Significance. If the result announced in the abstract were established, it would be a useful advance: the existing gKYP-lemma toolbox does not automatically transfer to LPV systems, and an enlargement procedure governed by system-theoretic data would fill a genuine gap in the literature. The claimed dependence of the minimal enlargement on the pole-frequency gap and controllability Gramians is a natural and falsifiable statement. However, the submitted manuscript provides no evidence from which these claims can be checked. The paper contains no derivations, no machine-checked proofs, no reproducible code, and no numerical support for the announced theorem; its strengths and validity cannot be assessed from the submitted text.
major comments (4)
- [Abstract / Full text] The submitted full text is not the paper described by the abstract. The body, titled 'RealisMotion,' contains no occurrence of the generalized KYP lemma, IQC, LPV, controllability Gramians, Iwasaki, Hara, or finite-frequency analysis; instead it presents a diffusion-based video generation method. Consequently, every mathematical claim in the abstract is unsupported by the manuscript. This is not a local gap in a derivation; the entire technical content of the announced result is absent.
- [Abstract, paragraph 2] The abstract states that 'we first demonstrate through a counterexample that the IQC non-negativity property may fail for LPV systems.' No counterexample appears anywhere in the submitted text. No LPV system, no scheduling-parameter model, no frequency-limited input class, and no numerical verification is defined, so the claimed invalidation of existing IQC-based results cannot be examined.
- [Abstract, paragraphs 2-3] The core reformulation is also absent: the abstract claims that replacing the original frequency range with an enlarged one restores non-negativity and that the minimal required expansion depends on the pole-frequency gap and a set of controllability Gramians. The manuscript gives no formula for the enlarged range, no theorem statement, no proof of restored non-negativity, and no comparison with existing LPV gKYP results. There are also no numerical examples of 'potential and efficiency' as promised in the abstract.
- [Section 5, Conclusions (Limitation paragraph)] The only limitation acknowledged in the body concerns 'foreground–background lighting inconsistencies' of the video-generation method. This limitation is unrelated to the announced control-theoretic results and reinforces that the submitted text cannot support the claims in the abstract.
minor comments (3)
- [Title] The title 'Extension of generalized KYP lemma: from LTI systems to LPV systems' does not match the content of the full text, which is a paper on human video generation; the metadata must be corrected to correspond to the actual submission.
- [Abstract] The abstract contains typographical and grammatical errors, including 'Building upon this results' and the missing space in 'interaction(or gap)'; these should be corrected in any resubmission.
- [References] The reference list of the full text contains no entry for Iwasaki and Hara (2005) or for any control-theory literature cited in the abstract; the bibliography is inconsistent with the claimed subject matter.
Circularity Check
No circularity can be established because the submitted full text contains no derivation of the claimed gKYP-LPV extension; a missing-proof problem is not a circularity finding.
full rationale
The abstract of arXiv:2508.08579 claims an extension of the generalized KYP lemma to LPV systems, with a minimal frequency-range expansion depending on pole-frequency gaps and controllability Gramians. However, the submitted full text is a computer vision paper, 'RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space', and contains no occurrence of IQC, gKYP, LPV, controllability Gramians, Iwasaki-Hara, or any related mathematical derivation. There is therefore no derivation chain that can be walked, and no specific equation or section can be quoted to exhibit a reduction of a prediction or theorem to its own inputs. The rules require quoting the paper and exhibiting the specific reduction, and explicitly state that a missing-evidence or missing-proof problem is not a circularity argument. The abstract's stated dependence on pole-frequency gaps and controllability Gramians refers to system data rather than to the target inequality itself, so even the abstract alone does not show a self-definitional or fitted-input circularity. Accordingly, the honest finding is no significant circularity, with score 0.
Assumptions & free parameters
assumptions (2)
- standard math The generalized KYP lemma and its IQC interpretation are valid for LTI systems as established by Iwasaki and Hara (2005).
- domain assumption LPV systems admit a finite-frequency characterization analogous to LTI systems via scheduling-parameter-dependent analysis.
Cite this review
Pith. "Pith review of Extension of generalized KYP lemma: from LTI systems to LPV systems." pith.science (2026). https://pith.science/paper/KTD4PZ53
@misc{pith2026250808579,
author = {Pith},
title = {Pith review of: Extension of generalized KYP lemma: from LTI systems to LPV systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/KTD4PZ53}},
note = {Machine review of arXiv:2508.08579}
}
read the original abstract
The generalized Kalman-Yakubovich-Popov (gKYP) lemma, established by Iwasaki and Hara (2005 IEEE TAC), has served as a fundamental tool for finite-frequency analysis and synthesis of linear time-invariant (LTI) systems. Over the past two decades, efforts to extend the gKYP lemma from LTI systems to linear parameter varying (LPV) systems have been hindered by the intricate time-frequency inter-modulation effect between the input signal and the time-varying scheduling parameter. A key element in this framework is the frequency-dependent Integral Quadratic Constraint (IQC) function, which enables time-domain interpretation of the gKYP lemma, as demonstrated by Iwasaki et al in their companion 2005 System and Control Letter paper. The non-negativity property of this IQC function plays a crucial role in characterizing system behavior under frequency-limited inputs. In this paper, we first demonstrate through a counterexample that the IQC non-negativity property may fail for LPV systems, thereby invalidating existing results that rely on this assumption. To address this issue, we propose a reformulation strategy that replaces the original frequency range with an enlarged one, thereby restoring the non-negativity property for LPV systems. Moreover, we establish that the minimal required expansion depends on the interaction(or gap) between the system poles and the original frequency range, as well as a set of controllability Gramians. Building upon this results, an extension of gKYP lemma is presented, which allows us to conduct finite-frequency analysis of LPV systems in a direct and reliable manner. The potential and efficiency compared to existing results are demonstrated through numerical examples.
Reference graph
Works this paper leans on
-
[1]
Dhruv Agrawal, Martin Guay, Jakob Buhmann, Dominik Borer, and Robert W. Sumner. Pose and skeleton-aware neu- ral ik for pose and motion editing. In SIGGRAPH Asia 2023 Conference Papers, pages 1–10, 2023. 5
work page 2023
-
[2]
Lan- guage2pose: Natural language grounded pose forecasting
Chaitanya Ahuja and Louis-Philippe Morency. Lan- guage2pose: Natural language grounded pose forecasting. In 2019 International conference on 3D vision (3DV), pages 719–728. IEEE, 2019. 3
work page 2019
-
[3]
Rhythm is a dancer: Music-driven motion synthesis with global structure
Andreas Aristidou, Anastasios Yiannakidis, Kfir Aberman, Daniel Cohen-Or, Ariel Shamir, and Yiorgos Chrysanthou. Rhythm is a dancer: Music-driven motion synthesis with global structure. IEEE transactions on visualization and computer graphics, 29(8):3519–3534, 2022. 3
work page 2022
-
[4]
Seamless human motion composition with blended posi- tional encodings
German Barquero, Sergio Escalera, and Cristina Palmero. Seamless human motion composition with blended posi- tional encodings. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 457–469, 2024. 3
work page 2024
-
[5]
Stable video diffusion: Scaling latent video diffusion models to large datasets
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023. 3
arXiv 2023
-
[6]
Depth pro: Sharp monocular metric depth in less than a second
Aleksei Bochkovskii, Ama ˜AG ¸ l Delaunoy, Hugo Germain, Marcel Santos, Yichao Zhou, Stephan R Richter, and Vladlen Koltun. Depth pro: Sharp monocular metric depth in less than a second. arXiv preprint arXiv:2410.02073, 2024. 4
arXiv 2024
-
[7]
Keep it smpl: Automatic estimation of 3d human pose and shape from a single image
Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J Black. Keep it smpl: Automatic estimation of 3d human pose and shape from a single image. In Computer Vision–ECCV 2016: 14th Euro- pean Conference, Amsterdam, The Netherlands, October 11- 14, 2016, Proceedings, Part V 14, pages 561–578. Springer,
2016
-
[8]
Implicit neural representations for variable length human motion generation
Pablo Cervantes, Yusuke Sekikawa, Ikuro Sato, and Koichi Shinoda. Implicit neural representations for variable length human motion generation. In European Conference on Com- puter Vision, pages 356–372. Springer, 2022. 3
work page 2022
Show all 69 references
-
[9]
Lumosflow: Motion-guided long video generation
Jiahao Chen, Hangjie Yuan, Yichen Qian, Jingyun Liang, Ji- azheng Xing, Pengwei Liu, Weihua Chen, Fan Wang, and Bing Su. Lumosflow: Motion-guided long video generation. arXiv preprint arXiv:2506.02497, 2025. 3
2025 arXiv
-
[10]
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12873–12883, 2021. 3, 6
2021
-
[11]
3dtrajmaster: Mastering 3d trajectory for multi-entity motion in video generation
Xiao Fu, Xian Liu, Xintao Wang, Sida Peng, Menghan Xia, Xiaoyu Shi, Ziyang Yuan, Pengfei Wan, Di Zhang, and Dahua Lin. 3dtrajmaster: Mastering 3d trajectory for multi-entity motion in video generation. arXiv preprint arXiv:2412.07759, 2024. 3, 7
2024 arXiv
-
[12]
Humans in 4d: Re- constructing and tracking humans with transformers
Shubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa, and Jitendra Malik. Humans in 4d: Re- constructing and tracking humans with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14783–14794, 2023. 2
2023
-
[13]
Generating diverse and natural 3d human motions from text
Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating diverse and natural 3d human motions from text. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 5152–5161, 2022. 3
2022
-
[14]
Learning an infant body model from rgb- d data for accurate full body motion analysis
Nikolas Hesse, Sergi Pujades, Javier Romero, Michael J Black, Christoph Bodensteiner, Michael Arens, Ulrich G Hofmann, Uta Tacke, Mijna Hadders-Algra, Raphael Wein- berger, et al. Learning an infant body model from rgb- d data for accurate full body motion analysis. In Medi- c...
2018
-
[15]
Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium. Advances in neural information processing systems , 30, 2017. 6
2017
-
[16]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 3
2020
-
[17]
Animate anyone: Consistent and controllable image- to-video synthesis for character animation
Li Hu. Animate anyone: Consistent and controllable image- to-video synthesis for character animation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8153–8163, 2024. 2, 3, 7
2024
-
[18]
Animate anyone 2: High-fidelity character image animation with environment affordance
Li Hu, Guangyuan Wang, Zhen Shen, Xin Gao, Dechao Meng, Lian Zhuo, Peng Zhang, Bang Zhang, and Liefeng Bo. Animate anyone 2: High-fidelity character image animation with environment affordance. arXiv preprint arXiv:2502.06145, 2025. 2, 3, 7
2025 arXiv
-
[19]
Vace: All-in-one video creation and editing
Zeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang, Yulin Pan, and Yu Liu. Vace: All-in-one video creation and editing. arXiv preprint arXiv:2503.07598, 2025. 3
2025 arXiv
-
[20]
End-to-end recovery of human shape and pose
Angjoo Kanazawa, Michael J Black, David W Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7122–7131, 2018. 2
2018
-
[21]
Flame: Free- form language-based motion synthesis & editing
Jihoon Kim, Jiseob Kim, and Sungjoon Choi. Flame: Free- form language-based motion synthesis & editing. In Pro- ceedings of the AAAI Conference on Artificial Intelligence , volume 37, pages 8255–8263, 2023. 3
2023
-
[22]
Hunyuanvideo: A systematic framework for large video generative models
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603, 2024. 3
2024 arXiv
-
[23]
Unipose: A unified multimodal framework for human pose comprehension, generation and editing
Yiheng Li, Ruibing Hou, Hong Chang, Shiguang Shan, and Xilin Chen. Unipose: A unified multimodal framework for human pose comprehension, generation and editing. arXiv preprint arXiv:2411.16781, 2024. 5
2024 arXiv
-
[24]
Movideo: Motion-aware video generation with diffusion model
Jingyun Liang, Yuchen Fan, Kai Zhang, Radu Timofte, Luc Van Gool, and Rakesh Ranjan. Movideo: Motion-aware video generation with diffusion model. In European Con- ference on Computer Vision , pages 56–74. Springer, 2024. 3 9
2024
-
[25]
Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models
Gaojie Lin, Jianwen Jiang, Jiaqi Yang, Zerong Zheng, and Chao Liang. Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models. arXiv preprint arXiv:2502.01061, 2025. 3
2025 arXiv
-
[26]
Smpl: A skinned multi- person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. Smpl: A skinned multi- person linear model. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2, pages 851–866. ACM, 2023. 2
2023
-
[27]
Mimo: Controllable character video synthesis with spatial decomposed modeling
Yifang Men, Yuan Yao, Miaomiao Cui, and Liefeng Bo. Mimo: Controllable character video synthesis with spatial decomposed modeling. arXiv preprint arXiv:2409.16160 ,
-
[28]
Agora: Avatars in geography optimized for regression analysis
Priyanka Patel, Chun-Hao P Huang, Joachim Tesch, David T Hoffmann, Shashank Tripathi, and Michael J Black. Agora: Avatars in geography optimized for regression analysis. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 13468–13478, 2021. 6
2021
-
[29]
Expressive body capture: 3d hands, face, and body from a single image
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black. Expressive body capture: 3d hands, face, and body from a single image. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognitio...
2019
-
[30]
Recon- structing hands in 3d with transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Recon- structing hands in 3d with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9826–9836, 2024. 5
2024
-
[31]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF inter- national conference on computer vision , pages 4195–4205,
-
[32]
Controlnext: Powerful and effi- cient control for image and video generation
Bohao Peng, Jian Wang, Yuechen Zhang, Wenbo Li, Ming- Chang Yang, and Jiaya Jia. Controlnext: Powerful and effi- cient control for image and video generation. arXiv preprint arXiv:2408.06070, 2024. 7
2024 arXiv
-
[33]
Babel: Bodies, action and behavior with english la- bels
Abhinanda R Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez, and Michael J Black. Babel: Bodies, action and behavior with english la- bels. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 722–731, 2021. 3
2021
-
[34]
Em- bodied hands: Modeling and capturing hands and bodies to- gether
Javier Romero, Dimitrios Tzionas, and Michael J Black. Em- bodied hands: Modeling and capturing hands and bodies to- gether. arXiv preprint arXiv:2201.02610, 2022. 5
2022 arXiv
-
[35]
Human motion diffusion as a generative prior
Yonatan Shafir, Guy Tevet, Roy Kapon, and Amit H Bermano. Human motion diffusion as a generative prior. arXiv preprint arXiv:2303.01418, 2023. 3
2023 arXiv
-
[36]
Human4dit: 360-degree human video generation with 4d diffusion transformer
Ruizhi Shao, Youxin Pang, Zerong Zheng, Jingxiang Sun, and Yebin Liu. Human4dit: 360-degree human video generation with 4d diffusion transformer. arXiv preprint arXiv:2405.17405, 2024. 3
2024 arXiv
-
[37]
World-grounded human motion recovery via gravity-view coordinates
Zehong Shen, Huaijin Pi, Yan Xia, Zhi Cen, Sida Peng, Zechen Hu, Hujun Bao, Ruizhen Hu, and Xiaowei Zhou. World-grounded human motion recovery via gravity-view coordinates. In SIGGRAPH Asia 2024 Conference Papers , pages 1–11, 2024. 2, 4
2024
-
[38]
Wham: Reconstructing world-grounded humans with accu- rate 3d motion
Soyong Shin, Juyong Kim, Eni Halilaj, and Michael J Black. Wham: Reconstructing world-grounded humans with accu- rate 3d motion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2070– 2080, 2024. 2
2024
-
[39]
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International confer- ence on machine learning, pages 2256–2265. PMLR, 2015. 3
2015
-
[40]
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063,
-
[41]
Animate-x: Universal character image ani- mation with enhanced motion representation
Shuai Tan, Biao Gong, Xiang Wang, Shiwei Zhang, Dandan Zheng, Ruobing Zheng, Kecheng Zheng, Jingdong Chen, and Ming Yang. Animate-x: Universal character image ani- mation with enhanced motion representation. arXiv preprint arXiv:2410.10306, 2024. 3, 7
-
[42]
Motionclip: Exposing human motion generation to clip space
Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. Motionclip: Exposing human motion generation to clip space. In European Conference on Com- puter Vision, pages 358–374. Springer, 2022. 3
2022
-
[43]
Human motion dif- fusion model
Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H Bermano. Human motion dif- fusion model. arXiv preprint arXiv:2209.14916, 2022. 3
2022 arXiv
-
[44]
Musepose: a pose-driven image-to-video framework for virtual human generation, 2024
Zhengyan Tong, Chao Li, Zhaokang Chen, Bin Wu, and Wenjiang Zhou. Musepose: a pose-driven image-to-video framework for virtual human generation, 2024. 7
2024
-
[45]
Least-squares estimation of transformation parameters between two point patterns
Shinji Umeyama. Least-squares estimation of transformation parameters between two point patterns. IEEE Transactions on Pattern Analysis & Machine Intelligence , 13(04):376– 380, 1991. 5
1991
-
[46]
Fvd: A new metric for video generation
Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Rapha¨el Marinier, Marcin Michalski, and Sylvain Gelly. Fvd: A new metric for video generation. arXiv preprint arXiv:1812.01717, 2019. 6
2019 arXiv
-
[47]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Il- lia Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017. 6
2017 arXiv
-
[48]
Wan: Open and advanced large-scale video gen- erative models
Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, et al. Wan: Open and advanced large-scale video gen- erative models. arXiv preprint arXiv:2503.20314, 2025. 2, 3, 6, 7
2025 arXiv
-
[49]
Disco: Disentangled control for referring human dance generation in real world
Tan Wang, Linjie Li, Kevin Lin, Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu, and Li- juan Wang. Disco: Disentangled control for referring human dance generation in real world. arXiv preprint arXiv:2307.00040, 2(3):4, 2023. 2, 3
2023 arXiv
-
[50]
Unianimate: Taming unified video diffusion mod- els for consistent human image animation
Xiang Wang, Shiwei Zhang, Changxin Gao, Jiayu Wang, Xiaoqiang Zhou, Yingya Zhang, Luxin Yan, and Nong Sang. Unianimate: Taming unified video diffusion mod- els for consistent human image animation. arXiv preprint arXiv:2406.01188, 2024. 3 10
2024 arXiv
-
[51]
Tram: Global trajectory and motion of 3d humans from in- the-wild videos
Yufu Wang, Ziyun Wang, Lingjie Liu, and Kostas Daniilidis. Tram: Global trajectory and motion of 3d humans from in- the-wild videos. In European Conference on Computer Vi- sion, pages 467–487. Springer, 2024. 2
2024
-
[52]
Humanvid: Demystifying training data for camera-controllable human image animation
Zhenzhi Wang, Yixuan Li, Yanhong Zeng, Youqing Fang, Yuwei Guo, Wenran Liu, Jing Tan, Kai Chen, Tianfan Xue, Bo Dai, et al. Humanvid: Demystifying training data for camera-controllable human image animation. In The Thirty- eight Conference on Neural Information Processing Syst...
2024
-
[53]
Motionctrl: A unified and flexible motion controller for video generation
Zhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li, Tianshui Chen, Menghan Xia, Ping Luo, and Ying Shan. Motionctrl: A unified and flexible motion controller for video generation. In ACM SIGGRAPH 2024 Conference Pa- pers, pages 1–11, 2024. 3, 6, 7
2024
-
[54]
Draganything: Motion control for any- thing using entity representation
Weijia Wu, Zhuang Li, Yuchao Gu, Rui Zhao, Yefei He, David Junhao Zhang, Mike Zheng Shou, Yan Li, Tingting Gao, and Di Zhang. Draganything: Motion control for any- thing using entity representation. In European Conference on Computer Vision, pages 331–348. Springer, 2024. 3
2024
-
[55]
Magicanimate: Temporally consistent human im- age animation using diffusion model
Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Chenxu Zhang, Jiashi Feng, and Mike Zheng Shou. Magicanimate: Temporally consistent human im- age animation using diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...
2024
-
[56]
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiao- han Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 3
2024 arXiv
-
[57]
Effec- tive whole-body pose estimation with two-stages distillation
Zhendong Yang, Ailing Zeng, Chun Yuan, and Yu Li. Effec- tive whole-body pose estimation with two-stages distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4210–4220, 2023. 2
2023
-
[58]
Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory
Shengming Yin, Chenfei Wu, Jian Liang, Jie Shi, Houqiang Li, Gong Ming, and Nan Duan. Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory. arXiv preprint arXiv:2308.08089, 2023. 3
2023 arXiv
-
[59]
Whac: World-grounded hu- mans and cameras
Wanqi Yin, Zhongang Cai, Ruisi Wang, Fanzhou Wang, Chen Wei, Haiyi Mei, Weiye Xiao, Zhitao Yang, Qingping Sun, Atsushi Yamashita, et al. Whac: World-grounded hu- mans and cameras. In European Conference on Computer Vision, pages 20–37. Springer, 2024. 2
2024
-
[60]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023. 2, 3, 6
2023
-
[61]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In IEEE Conference on Computer Vision and Pattern Recognition , pages 586–595,
-
[62]
Rohm: Robust human motion reconstruction via diffusion
Siwei Zhang, Bharat Lal Bhatnagar, Yuanlu Xu, Alexan- der Winkler, Petr Kadlecek, Siyu Tang, and Federica Bogo. Rohm: Robust human motion reconstruction via diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14606–14617, 2024. 2
2024
-
[63]
Mim- icmotion: High-quality human motion video generation with confidence-aware pose guidance
Yuang Zhang, Jiaxi Gu, Li-Wen Wang, Han Wang, Junqi Cheng, Yuefeng Zhu, and Fangyuan Zou. Mim- icmotion: High-quality human motion video generation with confidence-aware pose guidance. arXiv preprint arXiv:2406.19680, 2024. 3, 7
2024 arXiv
-
[64]
Tora: Trajectory-oriented diffusion transformer for video genera- tion
Zhenghao Zhang, Junchao Liao, Menghao Li, Zuozhuo Dai, Bingxue Qiu, Siyu Zhu, Long Qin, and Weizhi Wang. Tora: Trajectory-oriented diffusion transformer for video genera- tion. arXiv preprint arXiv:2407.21705, 2024. 3, 7
2024 arXiv
-
[65]
Realisdance: Equip controllable character anima- tion with realistic hands
Jingkai Zhou, Benzhi Wang, Weihua Chen, Jingqi Bai, Dongyang Li, Aixi Zhang, Hao Xu, Mingyang Yang, and Fan Wang. Realisdance: Equip controllable character anima- tion with realistic hands. arXiv preprint arXiv:2409.06202,
-
[66]
Realisdance-dit: Sim- ple yet strong baseline towards controllable character anima- tion in the wild
Jingkai Zhou, Yifan Wu, Shikai Li, Min Wei, Chao Fan, Wei- hua Chen, Wei Jiang, and Fan Wang. Realisdance-dit: Sim- ple yet strong baseline towards controllable character anima- tion in the wild. arXiv preprint arXiv:2504.14977, 2025. 3, 6, 7
2025 arXiv
-
[67]
Champ: Controllable and consistent human image an- imation with 3d parametric guidance
Shenhao Zhu, Junming Leo Chen, Zuozhuo Dai, Zilong Dong, Yinghui Xu, Xun Cao, Yao Yao, Hao Zhu, and Siyu Zhu. Champ: Controllable and consistent human image an- imation with 3d parametric guidance. In European Confer- ence on Computer Vision, pages 145–162. Springer, 2024. 2, 3
2024
-
[68]
Denoising dif- fusion models for plug-and-play image restoration
Yuanzhi Zhu, Kai Zhang, Jingyun Liang, Jiezhang Cao, Bi- han Wen, Radu Timofte, and Luc Van Gool. Denoising dif- fusion models for plug-and-play image restoration. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1219–1229, 2023. 3
2023
-
[69]
Parco: Part- coordinating text-to-motion synthesis
Qiran Zou, Shangyuan Yuan, Shian Du, Yu Wang, Chang Liu, Yi Xu, Jie Chen, and Xiangyang Ji. Parco: Part- coordinating text-to-motion synthesis. In European Confer- ence on Computer Vision , pages 126–143. Springer, 2024. 3 11
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.