REVIEW 5 major objections 4 minor 79 references
A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty Motions
T0 review · 5 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A mask-conditioned correction module plus a physics-based transfer module restores physical plausibility to high-difficulty video-captured motions such as gymnastics and martial arts, while keeping the original movement pattern.
desk verdict Genuine extension with strong results, but the unconstrained residual forces need a direct answer before the physical-restoration claim fully lands. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the sequence of MCM followed by PTM. MCM is a mask-conditioned diffusion in-betweening model: it uses the video's segmentation masks as a condition, and regenerates the frames flagged as flawed by a comparison of projected 3D joints against 2D keypoints and masks, so the corrected reference is both consistent with the image evidence and friendly to imitation. PTM is a pre-trained humanoid imitation controller that adapts on the fly to each motion sequence; its test-time adaptation uses a relative reward (ignoring the absolute root position to avoid error accumulation), a relative early-termination rule keyed to mean joint distance, and residual forces, which are extra torques compensating for missing environmental effects. The pretrain-and-adapt pattern is what lets a single controller handle a long tail of high-difficulty motions without catastrophic forgetting.
What would settle it
Run PTM on an already physically clean reference motion that needs no external apparatus (for example, a slow push-up or squat from AMASS) and measure the residual forces learned during adaptation; if the residual forces remain large or track the reference's noise even though no external forces are needed, the physical plausibility gains are partly an artifact of unconstrained force fitting rather than genuine dynamics.
Extended reading notes
Core claim
The paper claims that the reason simulation-based motion imitation has failed on high-difficulty in-the-wild motions is twofold: the video capture output contains short flawed segments that break imitation, and the motions themselves are long-tail and force-intensive, beyond the reach of a single pre-trained controller. To address the first issue, MCM detects frames where the projected 3D body disagrees with the video's human segmentation mask or 2D keypoints, and regenerates those frames with a mask-conditioned diffusion in-betweening model, yielding a context-consistent reference motion. To address the second, PTM uses a pre-trained imitation controller and adapts it per motion sequence at test time via reinforcement learning, with a relative reward that drops the absolute root position, an early-termination rule based on mean relative joint distance, and residual forces that stand in for missing environmental apparatus such as trampolines and mats. The authors report that this two-stage pipeline, applied as a plug-and-play post-processor to the outputs of existing video mocap methods, keeps ground penetration below 0.5 on all tested datasets, reduces self-penetration by more than half, and heavily cuts foot sliding and floating, while preserving the original movement pattern.
Load-bearing premise
The whole approach rests on the assumption that the unconstrained residual forces applied during test-time adaptation compensate for genuinely missing environmental effects, rather than serving as free control signals that let the simulated humanoid track a noisy reference without actually being physically plausible.
Editorial extensions
If this is right
- Any existing video motion capture method can be post-processed by this plug-and-play pipeline to yield physically grounded 3D motion without retraining the capture method itself.
- Physical realism metrics (ground penetration, foot sliding, floating, self-penetration) drop by large margins on standard datasets and on the new in-the-wild set, making high-quality motion assets for games, VR, and animation obtainable from ordinary video.
- The pretrain-and-adapt scheme raises imitation success on the kungfu subset from about 76 percent (PHC+) to about 98 percent, indicating that test-time adaptation is a workable route to simulating long-tail, high-difficulty movements.
- The reported gains apply to single-person sequences; closely interactive multi-person motions are not restored by the current modules.
Reading between the lines
- Editorial inference: if residual forces are not tied to the video content, part of the reported physical plausibility could come from fitting the reference rather than modeling true dynamics; a testable check is whether the learned residual forces stay near zero on motions that need no external apparatus.
- Editorial inference: the method's ceiling is coupled to segmentation quality during fast, blurry frames, so improvements in segmentation robustness to motion blur should transfer directly to the correction module without changing the rest of the pipeline.
- Editorial inference: the pretrain-and-adapt pattern may transfer to other long-tail control problems, such as simulating diverse robot manipulation skills, where a broad prior covers common cases and per-instance adaptation handles rare ones.
- Editorial inference: the residual forces could be re-parameterized as contact forces with an estimated deformable surface such as a trampoline or mat, turning the current free-torque compensation into a physically grounded interaction model.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a two-stage plug-and-play pipeline for improving the physical plausibility of 3D human motions estimated by monocular video capture, with emphasis on high-difficulty in-the-wild motions such as gymnastics and martial arts. The first module (MCM) uses video segmentation masks and motion context to detect and correct short flawed motion segments via a diffusion-based in-betweening model. The second module (PTM) combines a pre-trained imitation controller with a test-time adaptation strategy that uses relative rewards, early termination, and residual forces. The authors collect a new 206-video benchmark and report substantial reductions in ground penetration, foot sliding, and floating on AIST++, a Kungfu subset, and EMDB, while preserving 2D/3D reconstruction fidelity.
Significance. If the claims are substantiated, the work addresses a real gap: existing physics-based imitation methods fail on high-difficulty motions estimated from monocular video. The proposed MCM-PTM pipeline and the new in-the-wild benchmark are potentially useful contributions to the human-motion-analysis community. The paper is thorough in its empirical coverage, including ablations of the adaptation settings and the correction module (Tables 4 and 5), and it correctly identifies the need to handle both flawed references and complex imitation tasks. However, the central claim of 'physical restoration' is undermined by the unconstrained residual forces and by the self-referential nature of the physical metrics, as detailed below.
major comments (5)
- [Section 3.3, Residual Force] The residual forces are introduced without any parameterization, bound, or restriction to contact phases; the text only says they compensate for 'the absence of these environmental conditions in our simulations.' Because the adaptation objective (Eq. 5) rewards tracking the noisy reference and Eq. 6 terminates on deviation, the optimizer can use these forces as hidden controls that drag the character along any trajectory, including airborne phases. If such forces are active during flight or are not physically realizable, the reported reductions in GP/FS/Float are partly fitting artifacts: any trajectory can be made contact-consistent with enough unmodeled actuation. The paper must constrain residual forces (e.g., to contact phases, with plausible magnitudes) or provide an analysis of the applied forces and an ablation without them showing that the physical metrics remain acceptable.
- [Section 3.3, Eq. (5)] The relative reward is written as a sum of positive exponentials, e^{w_p||...||} + e^{w_r||...||} + ..., so the reward increases with tracking error. This appears to be a sign error; the intended reward is presumably e^{-w||...||}. Since this reward drives the adaptation, the error must be corrected and the reported results must be re-verified with the correct sign.
- [Section 4.2 and Table 1] The physical authenticity metrics (SP, GP, Float, FS) are computed on the output of the physics simulator, which by construction enforces ground contact and friction. Comparing these against non-simulated baselines conflates the simulator's inductive bias with the method's contribution. The authors should report the metrics on the pre-simulation corrected motion, add a baseline that applies a simple ground-projection or any physics-based post-process to the reference motions, and discuss the extent to which the gains are due to simulation rather than the proposed modules.
- [Section 4.3] Implementation details are largely deferred to an appendix that is not present in the manuscript. The OKS mismatch threshold, early-termination distance d_term, reward weights (w_p, w_r, w_v, w_omega), PD gains, adaptation step budget, and the exact training schedule are not reported. These are essential for reproducing the 'plug-and-play' system; without them the method cannot be independently verified.
- [Tables 1–5] All results appear to be single-run with no error bars or significance tests. While the headline gains in physical metrics are large, several comparisons (e.g., OKS differences in Tables 1 and 5) are small, and the ablations would be more convincing with multiple seeds or runs.
minor comments (4)
- [Section 3.2, Eq. (4)] The OKS formula uses δ(v_i > 0) without defining v_i; state that v_i is the visibility flag from the 2D keypoint detector.
- [Table 1] The SP column has '--' entries for PhysPT rows; if self-penetration was not computed for that method, say so in the caption.
- [Section 4.4] The text says 'reduced the self-penetration by over 50%' but Table 1 shows AIST++ GVHMR goes from 0.072 to 0.046 (about 36%); adjust the claim or the table.
- [Section 4.3] The statement that pre-training takes around 2–3 days on a single A100 GPU should specify the number of environment steps, the number of training samples, and the software stack used for the simulation.
Circularity Check
No significant circularity: PTM is an optimization/refinement pipeline, and the central tracking claim is supported by independent GT-based benchmarks; physical-authenticity metrics check the simulator output by design rather than reducing a prediction to its input.
full rationale
The paper does not claim a derivation chain: PTM is an RL-based post-processing/optimization system, and MCM is a diffusion-based in-betweening module. No fitted constant is later renamed as a prediction. The central claim that high-difficulty motions can be tracked and physically refined is supported by external evaluations on EMDB, AIST++, and a new in-the-wild set with GT-based metrics (WA-MPJPE, W-MPJPE, RTE, MPJPE, PA-MPJPE), which are independent of the method's internal objectives. The physical-authenticity metrics (GP, FS, Float) are computed on the simulator output; contact and friction in the simulator naturally enforce low values, but this is an expected property of a physics-based post-processor rather than a derivation-level circularity. The Mask-Pose Similarity metric overlaps with the mask conditioning used by MCM, which weakens it as independent evidence, and the residual forces in Section 3.3 are unconstrained, which is a validity/robustness concern rather than a circularity. The only self-citation ([77], Section 1) is peripheral and not load-bearing. Accordingly, no specific circular step can be exhibited.
Assumptions & free parameters
free parameters (5)
- OKS/mask mismatch threshold for flaw detection =
not reported
- Early-termination distance threshold d_term =
not reported
- Relative reward weights (w_p, w_r, w_v, w_omega) =
not reported
- PD gains k_p, k_d =
not reported
- Adaptation step budget =
less than 500 for normal motions; 2000-4000 for high-difficulty motions
assumptions (4)
- domain assumption Segmentation masks from SAM provide reliable body localization even in blurred or fast-moving frames, and are stable where keypoint detectors fail.
- domain assumption Flawed motion appears as short segments surrounded by reliable motion context, so mask-guided diffusion in-betweening can replace them without drifting.
- domain assumption Residual forces can represent missing environmental supports such as trampolines or mats without violating the physical plausibility of the restored motion.
- standard math PPO converges to a policy that, after per-sequence adaptation, tracks high-difficulty reference motions while preserving the original pattern.
Cite this review
Pith. "Pith review of A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty Motions." pith.science (2026). https://pith.science/paper/TBXQNWB2
@misc{pith2026241217377,
author = {Pith},
title = {Pith review of: A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty Motions},
year = {2026},
howpublished = {\url{https://pith.science/paper/TBXQNWB2}},
note = {Machine review of arXiv:2412.17377}
}
read the original abstract
Extracting physically plausible 3D human motion from videos is a critical task. Although existing simulation-based motion imitation methods can enhance the physical quality of daily motions estimated from monocular video capture, extending this capability to high-difficulty motions remains an open challenge. This can be attributed to some flawed motion clips in video-based motion capture results and the inherent complexity in modeling high-difficulty motions. Therefore, sensing the advantage of segmentation in localizing human body, we introduce a mask-based motion correction module (MCM) that leverages motion context and video mask to repair flawed motions, producing imitation-friendly motions; and propose a physics-based motion transfer module (PTM), which employs a pretrain and adapt approach for motion imitation, improving physical plausibility with the ability to handle in-the-wild and challenging motions. Our approach is designed as a plug-and-play module to physically refine the video motion capture results, including high-difficulty in-the-wild motions. Finally, to validate our approach, we collected a challenging in-the-wild test set to establish a benchmark, and our method has demonstrated effectiveness on both the new benchmark and existing public datasets.https://physicalmotionrestoration.github.io
Figures
Reference graph
Works this paper leans on
-
[1]
Physics-based motion capture imitation with deep reinforcement learning
Nuttapong Chentanez, Matthias M ¨uller, Miles Macklin, Vik- tor Makoviychuk, and Stefan Jeschke. Physics-based motion capture imitation with deep reinforcement learning. In Pro- ceedings of the 11th ACM SIGGRAPH Conference on Mo- tion, Interaction and Games, pages 1–10, 2018. 3
work page 2018
-
[2]
Beyond static features for temporally consis- tent 3d human pose and shape from a video
Hongsuk Choi, Gyeongsik Moon, Ju Yong Chang, and Ky- oung Mu Lee. Beyond static features for temporally consis- tent 3d human pose and shape from a video. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1964–1973, 2021. 2
work page 1964
-
[3]
Flexible motion in-betweening with diffusion models
Setareh Cohan, Guy Tevet, Daniele Reda, Xue Bin Peng, and Michiel van de Panne. Flexible motion in-betweening with diffusion models. In ACM SIGGRAPH 2024 Conference Pa- pers, pages 1–9, 2024. 4
work page 2024
-
[4]
Physically consistent whole-body kinematics assessment based on an rgb-d sensor
Jessica Colombel, Vincent Bonnet, David Daney, Raphael Dumas, Antoine Seilles, and Franc ¸ois Charpillet. Physically consistent whole-body kinematics assessment based on an rgb-d sensor. application to simple rehabilitation exercises. Sensors, 20(10):2848, 2020. 1
work page 2020
-
[5]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in neural informa- tion processing systems, 34:8780–8794, 2021. 6
work page 2021
-
[6]
Tokenhmr: Advancing human mesh recov- ery with a tokenized pose representation
Sai Kumar Dwivedi, Yu Sun, Priyanka Patel, Yao Feng, and Michael J Black. Tokenhmr: Advancing human mesh recov- ery with a tokenized pose representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1323–1333, 2024. 2
work page 2024
-
[7]
Assistive gym: A physics simulation framework for assistive robotics
Zackory Erickson, Vamsee Gangaram, Ariel Kapusta, C Karen Liu, and Charles C Kemp. Assistive gym: A physics simulation framework for assistive robotics. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pages 10169–10176. IEEE, 2020. 1
work page 2020
-
[8]
Super- track: Motion tracking for physically simulated characters using supervised learning
Levi Fussell, Kevin Bergamin, and Daniel Holden. Super- track: Motion tracking for physically simulated characters using supervised learning. ACM Transactions on Graphics (TOG), 40(6):1–13, 2021. 3
work page 2021
Show all 79 references
-
[9]
Differentiable dynamics for articu- lated 3d human motion reconstruction
Erik G ¨artner, Mykhaylo Andriluka, Erwin Coumans, and Cristian Sminchisescu. Differentiable dynamics for articu- lated 3d human motion reconstruction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 13190–13200, 2022. 1, 3
2022
-
[10]
Trajectory optimization for physics-based re- construction of 3d human pose from monocular video
Erik G ¨artner, Mykhaylo Andriluka, Hongyi Xu, and Cristian Sminchisescu. Trajectory optimization for physics-based re- construction of 3d human pose from monocular video. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 13106–13115...
2022
-
[11]
3d human reconstruction in the wild with synthetic data using generative models
Yongtao Ge, Wenjia Wang, Yongfan Chen, Hao Chen, and Chunhua Shen. 3d human reconstruction in the wild with synthetic data using generative models. arXiv preprint arXiv:2403.11111, 2024. 2
2024
-
[12]
Posetriplet: Co-evolving 3d human pose estimation, imita- tion, and hallucination under self-supervision
Kehong Gong, Bingbing Li, Jianfeng Zhang, Tao Wang, Jing Huang, Michael Bi Mi, Jiashi Feng, and Xinchao Wang. Posetriplet: Co-evolving 3d human pose estimation, imita- tion, and hallucination under self-supervision. In Proceed- ings of the IEEE/CVF conference on computer visio...
2022
-
[13]
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In Proceedings...
2024
-
[14]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 6
2020
-
[15]
Neural mocon: Neural motion control for phys- ically plausible human motion capture
Buzhen Huang, Liang Pan, Yuan Yang, Jingyi Ju, and Yan- gang Wang. Neural mocon: Neural motion control for phys- ically plausible human motion capture. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6417–6426, 2022. 3
2022
-
[16]
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3. 6m: Large scale datasets and pre- dictive methods for 3d human sensing in natural environ- ments. IEEE transactions on pattern analysis and machine intelligence, 36(7):1325–1339, 2013. 5
2013
-
[17]
Learning 3d human dynamics from video
Angjoo Kanazawa, Jason Y Zhang, Panna Felsen, and Jiten- dra Malik. Learning 3d human dynamics from video. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 5614–5623, 2019. 2
2019
-
[18]
Gmd: Controllable human motion synthesis via guided diffusion models.arXiv preprint arXiv:2305.12577, 3, 2023
Korrawe Karunratanakul, Konpat Preechakul, Supasorn Suwajanakorn, and Siyu Tang. Gmd: Controllable human motion synthesis via guided diffusion models.arXiv preprint arXiv:2305.12577, 3, 2023. 4
2023 arXiv
-
[19]
Emdb: The electromagnetic database of global 3d human pose and shape in the wild
Manuel Kaufmann, Jie Song, Chen Guo, Kaiyue Shen, Tian- jian Jiang, Chengcheng Tang, Juan Jos ´e Z ´arate, and Otmar Hilliges. Emdb: The electromagnetic database of global 3d human pose and shape in the wild. In Proceedings of the IEEE/CVF International Conference on Computer ...
2023
-
[20]
Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C. Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick. Segment anything. arXiv:2304.02643, 2023. 2, 4
2023 arXiv
-
[21]
Vibe: Video inference for human body pose and shape estimation
Muhammed Kocabas, Nikos Athanasiou, and Michael J Black. Vibe: Video inference for human body pose and shape estimation. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 5253–5263, 2020. 2
2020
-
[22]
Pace: Human and motion estimation from in-the-wild videos
Muhammed Kocabas, Ye Yuan, Pavlo Molchanov, Yunrong Guo, Michael J Black, Otmar Hilliges, Jan Kautz, and Umar Iqbal. Pace: Human and motion estimation from in-the-wild videos. 3DV, 1(2):7, 2024. 2
2024
-
[23]
Optimal-state dynamics estimation for physics- based human motion capture from videos
Cuong Le, Viktor Johansson, Manon Kok, and Bastian Wandt. Optimal-state dynamics estimation for physics- based human motion capture from videos. arXiv preprint arXiv:2410.07795, 2024. 3
2024 arXiv
-
[24]
D &d: Learning human dynamics from dynamic camera
Jiefeng Li, Siyuan Bian, Chao Xu, Gang Liu, Gang Yu, and Cewu Lu. D &d: Learning human dynamics from dynamic camera. In European Conference on Computer Vision, pages 479–496. Springer, 2022. 1, 3
2022
-
[25]
Ross, and Angjoo Kanazawa
Ruilong Li, Shan Yang, David A. Ross, and Angjoo Kanazawa. Learn to dance with aist++: Music conditioned 3d dance generation, 2021. 5
2021
-
[26]
Motion-x: A large- scale 3d expressive whole-body human motion dataset
Jing Lin, Ailing Zeng, Shunlin Lu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, and Lei Zhang. Motion-x: A large- scale 3d expressive whole-body human motion dataset. Ad- vances in Neural Information Processing Systems, 2023. 5
2023
-
[27]
Dual-recommendation disentan- glement network for view fuzz in action recognition
Wenxuan Liu, Xian Zhong, Zhuo Zhou, Kui Jiang, Zheng Wang, and Chia-Wen Lin. Dual-recommendation disentan- glement network for view fuzz in action recognition. IEEE Trans. Image Process., 32:2719–2733, 2023. 1
2023
-
[28]
Smpl: A skinned multi- person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. Smpl: A skinned multi- person linear model. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2, pages 851–866. 2023. 2, 3
2023
-
[29]
Humantomato: Text-aligned whole-body motion generation
Shunlin Lu, Ling-Hao Chen, Ailing Zeng, Jing Lin, Ruimao Zhang, Lei Zhang, and Heung-Yeung Shum. Humantomato: Text-aligned whole-body motion generation. arXiv preprint arXiv:2310.12978, 2023. 3
2023 arXiv
-
[30]
3d human motion estimation via motion compression and re- finement
Zhengyi Luo, S Alireza Golestaneh, and Kris M Kitani. 3d human motion estimation via motion compression and re- finement. In Proceedings of the Asian Conference on Com- puter Vision, 2020. 2
2020
-
[31]
Dynamics-regulated kinematic policy for egocentric pose es- timation
Zhengyi Luo, Ryo Hachiuma, Ye Yuan, and Kris Kitani. Dynamics-regulated kinematic policy for egocentric pose es- timation. Advances in Neural Information Processing Sys- tems, 34:25019–25032, 2021. 3
2021
-
[32]
Embod- ied scene-aware human pose estimation
Zhengyi Luo, Shun Iwase, Ye Yuan, and Kris Kitani. Embod- ied scene-aware human pose estimation. Advances in Neural Information Processing Systems, 35:6815–6828, 2022. 1
2022
-
[33]
Perpetual humanoid control for real-time simulated avatars
Zhengyi Luo, Jinkun Cao, Kris Kitani, Weipeng Xu, et al. Perpetual humanoid control for real-time simulated avatars. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10895–10904, 2023. 2, 3, 4, 8
2023
-
[34]
Universal hu- manoid motion representations for physics-based control
Zhengyi Luo, Jinkun Cao, Josh Merel, Alexander Winkler, Jing Huang, Kris Kitani, and Weipeng Xu. Universal hu- manoid motion representations for physics-based control. arXiv preprint arXiv:2310.04582, 2023. 3, 6
2023 arXiv
-
[35]
Smplolympics: Sports environ- ments for physically simulated humanoids
Zhengyi Luo, Jiashun Wang, Kangni Liu, Haotian Zhang, Chen Tessler, Jingbo Wang, Ye Yuan, Jinkun Cao, Zihui Lin, Fengyi Wang, et al. Smplolympics: Sports environ- ments for physically simulated humanoids. arXiv preprint arXiv:2407.00187, 2024. 3
2024 arXiv
-
[36]
Grammar: Ground-aware motion model for 3d human motion reconstruction
Sihan Ma, Qiong Cao, Hongwei Yi, Jing Zhang, and Dacheng Tao. Grammar: Ground-aware motion model for 3d human motion reconstruction. In Proceedings of the 31st ACM International Conference on Multimedia, pages 2817– 2828, 2023. 2
2023
-
[37]
Amass: Archive of motion capture as surface shapes
Naureen Mahmood, Nima Ghorbani, Nikolaus F Troje, Ger- ard Pons-Moll, and Michael J Black. Amass: Archive of motion capture as surface shapes. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5442–5451, 2019. 5
2019
-
[38]
Vnect: Real-time 3d human pose estimation with a single rgb cam- era
Dushyant Mehta, Srinath Sridhar, Oleksandr Sotnychenko, Helge Rhodin, Mohammad Shafiei, Hans-Peter Seidel, Weipeng Xu, Dan Casas, and Christian Theobalt. Vnect: Real-time 3d human pose estimation with a single rgb cam- era. Acm transactions on graphics (tog) , 36(4):1–14, 2017. 1
2017
-
[39]
Catch & carry: reusable neural controllers for vision-guided whole-body tasks
Josh Merel, Saran Tunyasuvunakool, Arun Ahuja, Yuval Tassa, Leonard Hasenclever, Vu Pham, Tom Erez, Greg Wayne, and Nicolas Heess. Catch & carry: reusable neural controllers for vision-guided whole-body tasks. ACM Trans- actions on Graphics (TOG), 39(4):39–1, 2020. 3
2020
-
[40]
Deeploco: Dynamic locomotion skills using hierarchical deep reinforcement learning
Xue Bin Peng, Glen Berseth, KangKang Yin, and Michiel Van De Panne. Deeploco: Dynamic locomotion skills using hierarchical deep reinforcement learning. Acm transactions on graphics (tog), 36(4):1–13, 2017
2017
-
[41]
Deepmimic: Example-guided deep reinforce- ment learning of physics-based character skills
Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel Van de Panne. Deepmimic: Example-guided deep reinforce- ment learning of physics-based character skills. ACM Trans- actions On Graphics (TOG), 37(4):1–14, 2018. 4
2018
-
[42]
Mcp: Learning composable hierarchical control with multiplicative compositional policies.Advances in neural information processing systems, 32, 2019
Xue Bin Peng, Michael Chang, Grace Zhang, Pieter Abbeel, and Sergey Levine. Mcp: Learning composable hierarchical control with multiplicative compositional policies.Advances in neural information processing systems, 32, 2019
2019
-
[43]
Amp: Adversarial motion priors for styl- ized physics-based character control
Xue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine, and Angjoo Kanazawa. Amp: Adversarial motion priors for styl- ized physics-based character control. ACM Transactions on Graphics (ToG), 40(4):1–20, 2021. 4
2021
-
[44]
Ase: Large-scale reusable adversarial skill embeddings for physically simulated characters
Xue Bin Peng, Yunrong Guo, Lina Halper, Sergey Levine, and Sanja Fidler. Ase: Large-scale reusable adversarial skill embeddings for physically simulated characters. ACM Transactions On Graphics (TOG), 41(4):1–17, 2022. 3
2022
-
[45]
Humor: 3d human motion model for robust pose estimation
Davis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang, Srinath Sridhar, and Leonidas J Guibas. Humor: 3d human motion model for robust pose estimation. In Proceedings of the IEEE/CVF international conference on computer vision, pages 11488–11499, 2021. 2
2021
-
[46]
Trace and pace: Controllable pedestrian animation via guided trajec- tory diffusion
Davis Rempe, Zhengyi Luo, Xue Bin Peng, Ye Yuan, Kris Kitani, Karsten Kreis, Sanja Fidler, and Or Litany. Trace and pace: Controllable pedestrian animation via guided trajec- tory diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, ...
2023
-
[47]
Diffmimic: Efficient motion mimicking with differentiable physics
Jiawei Ren, Cunjun Yu, Siwei Chen, Xiao Ma, Liang Pan, and Ziwei Liu. Diffmimic: Efficient motion mimicking with differentiable physics. arXiv preprint arXiv:2304.03274 ,
-
[48]
Proximal policy optimization algo- rithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Rad- ford, and Oleg Klimov. Proximal policy optimization algo- rithms. arXiv preprint arXiv:1707.06347, 2017. 4
2017 arXiv
-
[49]
Global-to-local modeling for video-based 3d human pose and shape estimation
Xiaolong Shen, Zongxin Yang, Xiaohan Wang, Jianxin Ma, Chang Zhou, and Yi Yang. Global-to-local modeling for video-based 3d human pose and shape estimation. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8887–8896, 2023. 2
2023
-
[50]
World-grounded human motion recovery via gravity-view coordinates
Zehong Shen, Huaijin Pi, Yan Xia, Zhi Cen, Sida Peng, Zechen Hu, Hujun Bao, Ruizhen Hu, and Xiaowei Zhou. World-grounded human motion recovery via gravity-view coordinates. In SIGGRAPH Asia Conference Proceedings ,
-
[51]
Physcap: Physically plausible monocular 3d motion capture in real time
Soshi Shimada, Vladislav Golyanik, Weipeng Xu, and Chris- tian Theobalt. Physcap: Physically plausible monocular 3d motion capture in real time. ACM Transactions on Graphics (ToG), 39(6):1–16, 2020. 1, 3
2020
-
[52]
Neural monocular 3d human motion capture with physical awareness
Soshi Shimada, Vladislav Golyanik, Weipeng Xu, Patrick P´erez, and Christian Theobalt. Neural monocular 3d human motion capture with physical awareness. ACM Transactions on Graphics (ToG), 40(4):1–15, 2021. 3
2021
-
[53]
Wham: Reconstructing world-grounded humans with accu- rate 3d motion
Soyong Shin, Juyong Kim, Eni Halilaj, and Michael J Black. Wham: Reconstructing world-grounded humans with accu- rate 3d motion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2070– 2080, 2024. 2, 5
2024
-
[54]
Human mesh recovery from monocular images via a skeleton-disentangled representation
Yu Sun, Yun Ye, Wu Liu, Wenpeng Gao, Yili Fu, and Tao Mei. Human mesh recovery from monocular images via a skeleton-disentangled representation. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5349–5358, 2019. 2
2019
-
[55]
Trace: 5d temporal regression of avatars with dynamic cam- eras in 3d environments
Yu Sun, Qian Bao, Wu Liu, Tao Mei, and Michael J Black. Trace: 5d temporal regression of avatars with dynamic cam- eras in 3d environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 8856–8866, 2023. 2
2023
-
[56]
Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras
Zachary Teed and Jia Deng. Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras. Advances in neu- ral information processing systems, 34:16558–16569, 2021. 2
2021
-
[57]
Deep patch vi- sual odometry
Zachary Teed, Lahav Lipson, and Jia Deng. Deep patch vi- sual odometry. Advances in Neural Information Processing Systems, 36, 2024. 2
2024
-
[58]
Recovering 3d human mesh from monocular images: A sur- vey
Yating Tian, Hongwen Zhang, Yebin Liu, and Limin Wang. Recovering 3d human mesh from monocular images: A sur- vey. IEEE transactions on pattern analysis and machine in- telligence, 2023. 2
2023
-
[59]
3d hu- man pose estimation via intuitive physics
Shashank Tripathi, Lea M ¨uller, Chun-Hao P Huang, Omid Taheri, Michael J Black, and Dimitrios Tzionas. 3d hu- man pose estimation via intuitive physics. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4713–4725, 2023. 1, 3
2023
-
[60]
Aist dance video database: Multi-genre, multi-dancer, and multi-camera database for dance informa- tion processing
Shuhei Tsuchida, Satoru Fukayama, Masahiro Hamasaki, and Masataka Goto. Aist dance video database: Multi-genre, multi-dancer, and multi-camera database for dance informa- tion processing. In Proceedings of the 20th International Society for Music Information Retrieval Conferen...
2019
-
[61]
Mocapact: A multi-task dataset for simulated humanoid con- trol
Nolan Wagener, Andrey Kolobov, Felipe Vieira Frujeri, Ricky Loynd, Ching-An Cheng, and Matthew Hausknecht. Mocapact: A multi-task dataset for simulated humanoid con- trol. Advances in Neural Information Processing Systems , 35:35418–35431, 2022. 3
2022
-
[62]
Learn- ing human dynamics in autonomous driving scenarios
Jingbo Wang, Ye Yuan, Zhengyi Luo, Kevin Xie, Dahua Lin, Umar Iqbal, Sanja Fidler, and Sameh Khamis. Learn- ing human dynamics in autonomous driving scenarios. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 20796–20806, 2023. 1
2023
-
[63]
Pacer+: On-demand pedestrian animation controller in driving scenarios
Jingbo Wang, Zhengyi Luo, Ye Yuan, Yixuan Li, and Bo Dai. Pacer+: On-demand pedestrian animation controller in driving scenarios. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 718–728, 2024. 3
2024
-
[64]
Freeman: Towards benchmarking 3d human pose estimation under real-world conditions
Jiong Wang, Fengyu Yang, Bingliang Li, Wenbo Gou, Danqi Yan, Ailing Zeng, Yijun Gao, Junle Wang, Yanqing Jing, and Ruimao Zhang. Freeman: Towards benchmarking 3d human pose estimation under real-world conditions. In Proceedings of the IEEE/CVF Conference on Computer Vision and...
2024
-
[65]
Unicon: Universal neural controller for physics-based character motion
Tingwu Wang, Yunrong Guo, Maria Shugrina, and Sanja Fi- dler. Unicon: Universal neural controller for physics-based character motion. arXiv preprint arXiv:2011.15119, 2020. 3
2011 arXiv
-
[66]
Skillmimic: Learning reusable basketball skills from demonstrations
Yinhuai Wang, Qihan Zhao, Runyi Yu, Ailing Zeng, Jing Lin, Zhengyi Luo, Hok Wai Tsui, Jiwen Yu, Xiu Li, Qifeng Chen, et al. Skillmimic: Learning reusable basketball skills from demonstrations. arXiv preprint arXiv:2408.15270 ,
-
[67]
Tram: Global trajectory and motion of 3d humans from in- the-wild videos
Yufu Wang, Ziyun Wang, Lingjie Liu, and Kostas Daniilidis. Tram: Global trajectory and motion of 3d humans from in- the-wild videos. In European Conference on Computer Vi- sion, pages 467–487. Springer, 2025. 2, 5, 6, 7
2025
-
[68]
Quest- sim: Human motion tracking from sparse sensors with simu- lated avatars
Alexander Winkler, Jungdam Won, and Yuting Ye. Quest- sim: Human motion tracking from sparse sensors with simu- lated avatars. In SIGGRAPH Asia 2022 Conference Papers, pages 1–8, 2022. 3
2022
-
[69]
A scalable approach to control diverse behaviors for physi- cally simulated characters
Jungdam Won, Deepak Gopinath, and Jessica Hodgins. A scalable approach to control diverse behaviors for physi- cally simulated characters. ACM Transactions on Graphics (TOG), 39(4):33–1, 2020. 3
2020
-
[70]
Physics-based human motion es- timation and synthesis from videos
Kevin Xie, Tingwu Wang, Umar Iqbal, Yunrong Guo, Sanja Fidler, and Florian Shkurti. Physics-based human motion es- timation and synthesis from videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 11532–11541, 2021. 1, 3
2021
-
[71]
Decoupling human and camera motion from videos in the wild
Vickie Ye, Georgios Pavlakos, Jitendra Malik, and Angjoo Kanazawa. Decoupling human and camera motion from videos in the wild. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 21222–21232, 2023. 2
2023
-
[72]
Residual force control for ag- ile human behavior imitation and extended motion synthe- sis
Ye Yuan and Kris Kitani. Residual force control for ag- ile human behavior imitation and extended motion synthe- sis. Advances in Neural Information Processing Systems, 33: 21763–21774, 2020. 3, 5
2020
-
[73]
Simpoe: Simulated character control for 3d hu- man pose estimation
Ye Yuan, Shih-En Wei, Tomas Simon, Kris Kitani, and Ja- son Saragih. Simpoe: Simulated character control for 3d hu- man pose estimation. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 7159–7169, 2021. 1, 3
2021
-
[74]
Physdiff: Physics-guided human motion diffusion model
Ye Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat, and Jan Kautz. Physdiff: Physics-guided human motion diffusion model. In Proceedings of the IEEE/CVF international con- ference on computer vision, pages 16010–16021, 2023. 4
2023
-
[75]
Learning motion priors for 4d human body capture in 3d scenes
Siwei Zhang, Yan Zhang, Federica Bogo, Marc Pollefeys, and Siyu Tang. Learning motion priors for 4d human body capture in 3d scenes. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , pages 11343– 11353, 2021. 3
2021
-
[76]
Physpt: Physics-aware pretrained transformer for estimating human dynamics from monocular videos
Yufei Zhang, Jeffrey O Kephart, Zijun Cui, and Qiang Ji. Physpt: Physics-aware pretrained transformer for estimating human dynamics from monocular videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2305–2317, 2024. 3, 6, 7
2024
-
[77]
Bi-causal: Group activity recognition via bidi- rectional causality
Youliang Zhang, Wenxuan Liu, Danni Xu, Zhuo Zhou, and Zheng Wang. Bi-causal: Group activity recognition via bidi- rectional causality. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 1450–1459, 2024. 1
2024
-
[78]
Prox- ycap: Real-time monocular full-body capture in world space via human-centric proxy-to-motion learning
Yuxiang Zhang, Hongwen Zhang, Liangxiao Hu, Jiajun Zhang, Hongwei Yi, Shengping Zhang, and Yebin Liu. Prox- ycap: Real-time monocular full-body capture in world space via human-centric proxy-to-motion learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and P...
1954
-
[79]
On the continuity of rotation representations in neural networks
Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. On the continuity of rotation representations in neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5745–5753,
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.