Pith. sign in

REVIEW 2 major objections 5 minor 74 references

Improving Human Motion Plausibility with Body Momentum

T0 review · 2 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read Adding a momentum-matching loss to human motion models reduces foot sliding and jitter while preserving accuracy.

desk verdict Useful momentum loss for motion plausibility, but the spectrum term is Parseval-redundant with the time-domain angular-momentum term, undermining the frequency story. read the letter →

arxiv 2509.09496 v1 pith:EJ3JCEGU submitted 2025-09-11 cs.CV

classification cs.CV
keywords humanmotionmomentumplausibilitylossfunctionglobaltrajectoryrecoveryfootslidingangular
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the local joint rotations and the global root trajectory of human motion are physically coupled, and this coupling is captured by whole-body linear and angular momentum. It introduces a loss term, TMo, that pushes a model's generated momentum profiles to match ground-truth profiles. Adding TMo to existing motion recovery and generation models consistently reduces foot sliding, jitter, and balance errors without hurting accuracy. The loss is simple, requires no architectural changes, and works across multiple baselines and tasks.

What carries the argument

The central object is the whole-body linear momentum and angular momentum of the human body, computed by partitioning the SMPL model into 20 body parts with per-part masses, centroids, and inertia tensors. The loss TMo combines three terms: an angular-momentum matching term (with time derivative), a linear-momentum matching term (with time derivative), and a spectrum term that matches the discrete Fourier/cosine transform of angular momentum to ground truth. This machinery links local joint behavior to global movement without needing explicit force or torque estimation.

What would settle it

Measure the high-frequency content of linear and angular momentum in ground-truth motions that involve hard impacts or rapid direction changes (e.g., parkour landings, quick punches). If these motions exhibit large high-frequency momentum components comparable to the artifacts the loss is meant to suppress, then the premise underlying the spectrum loss is contradicted, and one would predict that L_S either does not help or actively harms plausibility on such motions.

Watch

Extended reading notes

Core claim

The central claim is that enforcing consistency between generated and ground-truth whole-body linear and angular momentum—computed in a world frame—improves the physical plausibility of reconstructed and generated human motion. The momentum terms aggregate the effect of all joint-level dynamics, so matching them provides a physically grounded bridge between local pose and global displacement. The proposed loss has three parts: matching linear momentum, matching angular momentum, and matching the frequency spectrum of angular momentum to suppress unnatural high-frequency content. Experiments on global trajectory recovery, full motion recovery, and text-to-motion generation show that the loss

Load-bearing premise

The frequency-domain justification for the spectrum loss assumes that external forces and torques acting on the body have small high-frequency content, so if a motion involves sharp impacts or very fast force changes (e.g., acrobatic landings), that specific loss term's physical grounding weakens.

Editorial extensions

If this is right

  • If the central claim holds, any kinematic motion model—reconstruction, prediction, or generation—can be made more physically plausible by adding the TMo loss during training, without redesigning the architecture.
  • The loss yields consistent improvements across diverse baselines (GLAMR, WHAM, PhysPT, TEMOS) and across datasets including in-the-wild and acrobatic motions, suggesting the coupling is general rather than task-specific.
  • The method performs better in low-data regimes, implying the momentum constraint acts as a useful inductive bias that reduces the amount of motion data needed.
  • The frequency-spectrum component provides a new, physically motivated detector of implausible motion: sequences with large high-frequency momentum components are likely unrealistic.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper leaves implicit is using the momentum-matching loss as a self-supervised test-time refinement objective, since momentum can be computed from the model's own outputs without ground-truth labels.
  • The momentum plausibility detector (based on high-frequency components) could be repurposed as a standalone evaluation metric or a filtering step for motion datasets, not just a training loss.
  • The same formulation could transfer to other articulated body models or even non-human characters, as long as a part-based mass and inertia model can be defined.
  • The loss might also serve as a regularizer in motion prediction tasks, where the future trajectory must remain dynamically consistent with the evolving pose.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper proposes TMo, an auxiliary training loss that encourages consistency between predicted/generated human motion and ground-truth whole-body linear momentum (LMo) and angular momentum (AMo), plus a spectrum-based term (L_S) on AMo. The loss is integrated into existing motion models (GLAMR, PhysPT, WHAM, TEMOS) and evaluated on global-trajectory prediction, global motion recovery, and text-to-motion generation. The authors report reduced foot sliding, lower jitter, and improved balance, with comparable accuracy, and provide ablations, weight-sensitivity analysis, a perceptual study, and a stated public code/data release.

Significance. The core idea is appealing and useful: whole-body linear and angular momentum are aggregate physical quantities that couple local joint motion to global root translation/rotation, and the proposed loss is simple and model-agnostic. The multi-task evaluation, component ablations, and perceptual study are strengths, and the paper ships code and data. If the empirical claim holds, this is a low-cost way to improve the plausibility of kinematic motion models. However, the spectrum-loss component is not actually frequency-selective, and the physical/frequency narrative needs substantial correction before the paper can be accepted as written.

major comments (2)
  1. [Sec. 3.3, Eq. (4c)] The spectrum loss L_S is mathematically redundant with the first term of L_AMo in Eq. (4a). When F is the DFT or an orthonormal DCT, Parseval's theorem gives ||F(AMo(hat R,hat theta)) - F(AMo(R,theta))||^2 = c ||AMo(hat R,hat theta) - AMo(R,theta)||^2 for a positive constant c. Thus L_S contains no frequency weighting or masking and is exactly a scaled version of the ||Delta AMo||^2 component already in L_AMo. The argument in Sec. 3.2 about high-frequency attenuation therefore does not justify L_S as implemented. Table 4 is consistent: the L_AMo-only and L_S-only rows give nearly identical jitter (15.52 vs 15.53) and FS (5.11 vs 5.19). This does not invalidate the momentum-alignment approach, but it removes L_S as a distinct contribution and forces a reinterpretation of the ablations. The authors should either remove L_S, implement a genuinely frequency-selective penalty (e.g., weighting
  2. [Sec. 3.2] The frequency-domain derivation relies on the external claim, citing [4], that external forces and torques have small high-frequency content. The supplementary (D.1) verifies that high-frequency momentum components are small on AMASS, but this is not verified for the evaluation datasets (EMDB, Kungfu, RICH) that contain dynamic or high-impact motions. Because L_S as written is not frequency-selective (see previous comment), the AMASS validation cannot rescue the physical grounding of L_S. Please either verify the assumption on the evaluation data or revise the motivation to describe the loss as aligning full momentum profiles rather than specifically suppressing high-frequency content.
minor comments (5)
  1. [Table 2] The abstract claims the loss 'preserves the accuracy' of the recovered motion, but WHAM+LTMo shows slightly worse RTE on both EMDB (4.3 vs 4.1) and RICH (4.4 vs 4.1). The degradation is small, but the claim should be qualified, and ideally the main metrics should be accompanied by error bars or significance tests.
  2. [Supplementary D.2] The text says the gaps are 'non-increasing', but the reported numbers are 3.06, 4.12, 2.74, 3.13, which are not monotonically non-increasing. Please correct the description.
  3. [Eq. (4c) / Sec. 3.3] The text says 'We use the discrete Fourier transform F and the discrete cosine transform' for the spectrum loss, but it is unclear whether both are used and how they are combined (e.g., averaged, summed, or used separately). Please specify the exact implementation.
  4. [Sec. 4.1] The PhysPT† baseline is described as using 'the same Transformer architecture as PhysPT, using only position based loss and our global trajectory predictor.' Please clarify the training protocol and how it differs from the original PhysPT, since it is a key comparison.
  5. [References / General] There are minor typos, e.g., reference [43] contains 'V ol.3' and the running header 'NGUYEN ET AL: BODY MOMENTUM IN HUMAN MOTION' is repeated. Also, the composite measure m_AB in Sec. 4.4 would benefit from a clearer explanation of the reference direction (baseline at full size).

Circularity Check

1 steps flagged · score 2.0 of 10

One loss component (LS) is a frequency-domain renaming of the time-domain AMo loss, but the central momentum-based plausibility claim remains an independent, externally evaluated supervised regression.

  1. other [Section 3.3, Eq. (4c); Section 3.2]
    "LS =∥F(AMo( ˆR, ˆθ))− F(AMo(R,θ))∥2,(4c) ... We use the discrete Fourier transform F and the discrete cosine transform [51] for the spectrum loss LS."

    Eq. (4c) defines LS as the squared frequency-domain distance between predicted and ground-truth angular momentum. For the DFT or DCT used in the paper, Parseval's theorem gives Σω |F(a)(ω)-F(b)(ω)|² = c Σt |a(t)-b(t)|² for a constant c>0. Hence LS = c∥AMo(R̂,θ̂)-AMo(R,θ)∥² = c∥ΔAMo∥², which is exactly the first term of LAMo in Eq. (4a), up to scaling. The loss therefore contains no frequency-selective weighting or high-frequency masking; the Sec. 3.2 argument that high-frequency momentum content should be small is not operationalized by LS. The ablation in Table 4 is consistent with this redundancy: LAMo-only and LS-only give nearly identical jitter (15.52 vs 15.53) and FS (5.11 vs 5.19).

full rationale

The central claim—that adding a momentum-consistency loss improves plausibility—is not circular. The loss is a supervised regression to ground-truth momentum profiles; the ground truth is external, no fitted parameter is renamed as a prediction, and no load-bearing self-citation or imported uniqueness theorem carries the argument. The experiments compare against GLAMR, WHAM, PhysPT, and TEMOS on external benchmarks and a perceptual study, so the empirical conclusion stands independently. The only notable reduction-by-construction is that the spectrum loss LS is equivalent, under an orthogonal transform, to the time-domain angular-momentum term already present in LAMo. This is a mathematical redundancy in one of three loss components and means the frequency-domain motivation is not actually enforced as a separate constraint, but it does not make the overall derivation self-referential. Accordingly, the circularity score is low.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central claim does not rest on any newly invented physical entity. The main burdens are the modeling assumptions of the SMPL part masses and the empirical frequency-domain premise. These are standard or validated approximations.

free parameters (5)
  • lambda_AMo = Not stated; tuned on validation
    Weight for angular momentum loss. Supplementary D.2 shows performance varies with weight, so the final choice is a free parameter.
  • lambda_LMo = Not stated; tuned on validation
    Weight for linear momentum loss. Same rationale as above.
  • lambda_S = Not stated; tuned on validation
    Weight for spectrum loss. Same rationale as above.
  • k0 = Depends on T and sampling frequency
    Frequency threshold in the HF plausibility detector (Sec. D.1). Chosen by hand.
  • K = 20
    Multiplier for implausibility threshold in Sec. D.1. Set to encompass 98.9% of training sequences; arbitrary.
assumptions (5)
  • domain assumption Uniform mass distribution over each SMPL body part
    Sec. 3.1: 'Assuming a uniform distribution of mass'. Needed to compute part masses and moments of inertia.
  • domain assumption Body part centroids are fixed in the part frame
    Sec. 3.1 and supplementary A. Validated on AMASS with mean error 4.9mm, but still an approximation.
  • domain assumption External torque and force spectra have negligible high-frequency content
    Sec. 3.2, citing [4]. Used to justify the spectrum loss L_S.
  • domain assumption SMPL body model represents the human body sufficiently for momentum computation
    Sec. 3.1. Standard assumption in the field.
  • standard math Newtonian physics applies to human motion
    The momentum definitions and conservation arguments are based on classical mechanics, which is standard for biomechanics.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Improving Human Motion Plausibility with Body Momentum." pith.science (2026). https://pith.science/paper/EJ3JCEGU

@misc{pith2026250909496,
  author       = {Pith},
  title        = {Pith review of: Improving Human Motion Plausibility with Body Momentum},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EJ3JCEGU}},
  note         = {Machine review of arXiv:2509.09496}
}
read the original abstract

Many studies decompose human motion into local motion in a frame attached to the root joint and global motion of the root joint in the world frame, treating them separately. However, these two components are not independent. Global movement arises from interactions with the environment, which are, in turn, driven by changes in the body configuration. Motion models often fail to precisely capture this physical coupling between local and global dynamics, while deriving global trajectories from joint torques and external forces is computationally expensive and complex. To address these challenges, we propose using whole-body linear and angular momentum as a constraint to link local motion with global movement. Since momentum reflects the aggregate effect of joint-level dynamics on the body's movement through space, it provides a physically grounded way to relate local joint behavior to global displacement. Building on this insight, we introduce a new loss term that enforces consistency between the generated momentum profiles and those observed in ground-truth data. Incorporating our loss reduces foot sliding and jitter, improves balance, and preserves the accuracy of the recovered motion. Code and data are available at the project page https://hlinhn.github.io/momentum_bmvc.

Figures

Figures reproduced from arXiv: 2509.09496 by the authors.

Figure 1
Figure 1. Root movements unexplained by local joint configurations lead to implausible [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Left: improvement in body stability, reflected by error of the change in the swing around gravity. Right: Root joint height during a jumping sequence. PhysPT [72] predicts unnatural changes of the root trajectory along the gravity direction, highlighted by red circles. Perceptual study. To assess the increase in plausibility due to our loss, we conduct a human study on Prolific [1]. We select 40 sequences from AMASS… view at source ↗
Figure 4
Figure 4. Our partition of the SMPL body model. While the body deforms with change in poses, the assumption that the centroid position is fixed with respect to the part’s frame of reference usually holds. We sample N = 1000 random sequences on AMASS training set and calculate the difference between the centroids position computed directly from the mesh, and computed indirectly from rotating them with respective body part’s ro… view at source ↗
Figures from the paper (3 more)
Figure 5
Figure 5. Figure 5: The effect of our loss on angular momentum’s high frequency components distribu [PITH_FULL_IMAGE:figures/full_fig_p018_5.png]
Figure 6
Figure 6. Figure 6: Comparisons of preference rating between ground truth, baseline and our method. [PITH_FULL_IMAGE:figures/full_fig_p020_6.png]
Figure 7
Figure 7. Figure 7: Top: baseline, bottom: ours. Notice the lack of leg flexion on the baseline TEMOS [PITH_FULL_IMAGE:figures/full_fig_p021_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

74 extracted references · 3 linked inside Pith

  1. [4]

    Boehm, Kieran M Nichols, and Kreg G

    Wendy L. Boehm, Kieran M Nichols, and Kreg G. Gruben. Frequency-dependent contri- butions of sagittal-plane foot force to upright human standing.Journal of biomechanics, 83:305–309, 2019

  2. [1]

    URLhttps://www.prolific.com/

  3. [2]

    PoseBERT: A Generic Transformer Module for Tem- poral 3D Human Modeling.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45:12798–12815, 2022

    Fabien Baradel, Romain Br’egier, Thibault Groueix, Philippe Weinzaepfel, Yannis Kalantidis, and Grégory Rogez. PoseBERT: A Generic Transformer Module for Tem- poral 3D Human Modeling.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45:12798–12815, 2022

  4. [3]

    Kender, and Zicheng Liu

    Emad Barsoum, John R. Kender, and Zicheng Liu. HP-GAN: Probabilistic 3D Human Motion Prediction via GAN.2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 1499–149909, 2017

  5. [5]

    Executing your Commands via Motion Diffusion in Latent Space.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18000–18010, 2022

    Xin Chen, Biao Jiang, Wen Liu, Zilong Huang, Bin Fu, Tao Chen, Jingyi Yu, and Gang Yu. Executing your Commands via Motion Diffusion in Latent Space.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18000–18010, 2022

  6. [6]

    Huber, Dagmar Sternad, and Martin A

    Enrico Chiovetto, Meghan E. Huber, Dagmar Sternad, and Martin A. Giese. Low- dimensional organization of angular momentum during walking on a narrow beam. Scientific Reports, 8, 2018

  7. [7]

    Efficient Human Motion Reconstruction from Monocular Videos with Physical Consistency Loss.SIGGRAPH Asia 2023 Conference Papers, 2023

    Lin Cong, Philipp Ruppel, Yizhou Wang, Xiang Pan, Norman Hendrich, and Jianwei Zhang. Efficient Human Motion Reconstruction from Monocular Videos with Physical Consistency Loss.SIGGRAPH Asia 2023 Conference Papers, 2023

  8. [8]

    Differ- entiable Dynamics for Articulated 3d Human Motion Reconstruction.2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13180–13190, 2022

    Erik Gartner, Mykhaylo Andriluka, Erwin Coumans, and Cristian Sminchisescu. Differ- entiable Dynamics for Articulated 3d Human Motion Reconstruction.2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13180–13190, 2022

Show all 74 references
  1. [9]

    Erik Gärtner, Mykhaylo Andriluka, Hongyi Xu, and Cristian Sminchisescu. Trajectory Optimization for Physics-Based Reconstruction of 3d Human Pose from Monocular Video.2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13096–13105, 2022

  2. [10]

    Geman and S

    D. Geman and S. Geman. Bayesian Image Analysis. In E. Bienenstock, F. F. Soulié, and G. Weisbuch, editors,Disordered Systems and Biological Organization, NATO ASI Series, vol. 20, pages 709–743. Springer, Berlin, Heidelberg, 1986. doi: 10.1007/ 978-3-642-82657-3_30

  3. [11]

    Humans in 4D: Reconstructing and Tracking Humans with Transformers

    Shubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa, and Jitendra Malik. Humans in 4D: Reconstructing and Tracking Humans with Transformers. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 14737– 14748, 2023. NGUYEN ET AL: BODY MOMEN...

  4. [12]

    Generating Diverse and Natural 3D Human Motions from Text.2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5142–5151, 2022

    Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating Diverse and Natural 3D Human Motions from Text.2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5142–5151, 2022

  5. [13]

    NeMF: Neural Motion Fields for Kinematic Animation

    Chengan He, Jun Saito, James Zachary, Holly Rushmeier, and Yi Zhou. NeMF: Neural Motion Fields for Kinematic Animation. InNeurIPS, 2022

  6. [14]

    Henning, Tristan Laidlow, and Stefan Leutenegger

    Dorian F. Henning, Tristan Laidlow, and Stefan Leutenegger. BodySLAM: Joint Camera Localisation, Mapping, and Human Motion Tracking.ArXiv, abs/2205.02301, 2022

  7. [15]

    MoGlow.ACM Transac- tions on Graphics (TOG), 39:1 – 14, 2019

    Gustav Eje Henter, Simon Alexanderson, and Jonas Beskow. MoGlow.ACM Transac- tions on Graphics (TOG), 39:1 – 14, 2019

  8. [16]

    Angular momentum in human walking.Journal of Experimental Biology, 211:467 – 481, 2008

    Hugh Herr and Marko Popovic. Angular momentum in human walking.Journal of Experimental Biology, 211:467 – 481, 2008

  9. [17]

    Buzhen Huang, Liang Pan, Yuan Yang, Jingyi Ju, and Yangang Wang. Neural Mo- Con: Neural Motion Control for Physically Plausible Human Motion Capture.2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6407–6416, 2022

  10. [18]

    Chun-Hao Paul Huang, Hongwei Yi, Markus Hoschle, Matvey Safroshkin, Tsvetelina Alexiadis, Senya Polikovsky, Daniel Scharstein, and Michael J. Black. Capturing and Inferring Dense Full-Body Human-Scene Contact.2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition ...

  11. [19]

    Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environ- ments.IEEE Transactions on Pattern Analysis and Machine Intelligence, 36:1325–1339, 2014

  12. [20]

    Computing the Moment of Inertia of a Solid Defined by a Triangle Mesh.Journal of Graphics Tools, 11:51 – 57, 2006

    Michael Kallay. Computing the Moment of Inertia of a Solid Defined by a Triangle Mesh.Journal of Graphics Tools, 11:51 – 57, 2006

  13. [21]

    Zhang, Panna Felsen, and Jitendra Malik

    Angjoo Kanazawa, Jason Y . Zhang, Panna Felsen, and Jitendra Malik. Learning 3D Human Dynamics From Video.2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5607–5616, 2018

  14. [22]

    Optimizing Diffusion Noise Can Serve As Universal Motion Priors.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1334–1345, 2023

    Korrawe Karunratanakul, Konpat Preechakul, Emre Aksan, Thabo Beeler, Supasorn Suwajanakorn, and Siyu Tang. Optimizing Diffusion Noise Can Serve As Universal Motion Priors.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1334–1345, 2023

  15. [23]

    EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 14586–14597, 2023

    Manuel Kaufmann, Jie Song, Chen Guo, Kaiyue Shen, Tianjian Jiang, Chengcheng Tang, Juan José Zárate, and Otmar Hilliges. EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 145...

  16. [24]

    Muhammed Kocabas, Nikos Athanasiou, and Michael J. Black. VIBE: Video Inference for Human Body Pose and Shape Estimation.2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5252–5262, 2019. 12NGUYEN ET AL: BODY MOMENTUM IN HUMAN MOTION

  17. [25]

    Black, Otmar Hilliges, Jan Kautz, and Umar Iqbal

    Muhammed Kocabas, Ye Yuan, Pavlo Molchanov, Yunrong Guo, Michael J. Black, Otmar Hilliges, Jan Kautz, and Umar Iqbal. PACE: Human and Camera Motion Estimation from in-the-wild Videos.2024 International Conference on 3D Vision (3DV), pages 397–408, 2023

  18. [26]

    Fast and flexible multi- legged locomotion using learned centroidal dynamics.ACM Transactions on Graphics (TOG), 39:46:1 – 46:17, 2020

    Tae-Joung Kwon, Yoonsang Lee, and Michiel van de Panne. Fast and flexible multi- legged locomotion using learned centroidal dynamics.ACM Transactions on Graphics (TOG), 39:46:1 – 46:17, 2020

  19. [27]

    Task-Generic Hierarchical Human Motion Prior using V AEs.2021 International Conference on 3D Vision (3DV), pages 771–781, 2021

    Jiaman Li, Ruben Villegas, Duygu Ceylan, Jimei Yang, Zhengfei Kuang, Hao Li, and Yajie Zhao. Task-Generic Hierarchical Human Motion Prior using V AEs.2021 International Conference on 3D Vision (3DV), pages 771–781, 2021

  20. [28]

    D&D: Learning Human Dynamics from Dynamic Camera

    Jiefeng Li, Siyuan Bian, Chao Xu, Gang Liu, Gang Yu, and Cewu Lu. D&D: Learning Human Dynamics from Dynamic Camera. InEuropean Conference on Computer Vision, 2022

  21. [29]

    CLIFF: Carrying Location Information in Full Frames into Human Pose and Shape Estimation

    Zhihao Li, Jianzhuang Liu, Zhensong Zhang, Songcen Xu, and Youliang Yan. CLIFF: Carrying Location Information in Full Frames into Human Pose and Shape Estimation. InEuropean Conference on Computer Vision, 2022

  22. [30]

    Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset

    Jing Lin, Ailing Zeng, Shunlin Lu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, and Lei Zhang. Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset. Advances in Neural Information Processing Systems, 2023

  23. [31]

    Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. SMPL: A Skinned Multi-Person Linear Model.Seminal Graphics Papers: Pushing the Boundaries, Volume 2, 2015

  24. [32]

    GraMMaR: Ground- aware Motion Model for 3D Human Motion Reconstruction.Proceedings of the 31st ACM International Conference on Multimedia, 2023

    Sihan Ma, Qiong Cao, Hongwei Yi, Jing Zhang, and Dacheng Tao. GraMMaR: Ground- aware Motion Model for 3D Human Motion Reconstruction.Proceedings of the 31st ACM International Conference on Multimedia, 2023

  25. [33]

    Zordan, and Christian R

    Adriano Macchietto, Victor B. Zordan, and Christian R. Shelton. Momentum control for balance.ACM SIGGRAPH 2009 papers, 2009

  26. [34]

    Troje, Gerard Pons-Moll, and Michael J

    Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. AMASS: Archive of Motion Capture As Surface Shapes.2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 5441–5450, 2019

  27. [35]

    On the coordina- tion of highly dynamic human movements: an extension of the uncontrolled manifold approach applied to precision jump in parkour.Scientific Reports, 8, 2018

    Galo Maldonado, François Bailly, Philippe Souéres, and Bruno Watier. On the coordina- tion of highly dynamic human movements: an extension of the uncontrolled manifold approach applied to precision jump in parkour.Scientific Reports, 8, 2018

  28. [36]

    Pose Trans- formers (POTR): Human Motion Prediction with Non-Autoregressive Transformers

    Ángel Martínez-González, Michael Villamizar, and Jean-Marc Odobez. Pose Trans- formers (POTR): Human Motion Prediction with Non-Autoregressive Transformers. 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), pages 2276–2284, 2021. NGUYEN ET AL: BODY M...

  29. [37]

    Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt

    Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal V . Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt. Monocular 3D Human Pose Estimation in the Wild Using Improved CNN Supervision.2017 International Conference on 3D Vision (3DV), pages 506–516, 2016

  30. [38]

    A review of 3D human body pose estimation and mesh recovery.Digit

    Zaka-Ud-Din Muhammad, Zhangjin Huang, and Rashid Khan. A review of 3D human body pose estimation and mesh recovery.Digit. Signal Process., 128:103628, 2022

  31. [39]

    Regulation of whole-body angular momentum during human walking.Scientific Reports, 13, 2023

    Takuo Negishi and Naomichi Ogihara. Regulation of whole-body angular momentum during human walking.Scientific Reports, 13, 2023

  32. [40]

    Black, and Gül Varol

    Mathis Petrovich, Michael J. Black, and Gül Varol. Action-Conditioned 3D Human Motion Synthesis with Transformer V AE. InInternational Conference on Computer Vision (ICCV), 2021

  33. [41]

    Black, and Gül Varol

    Mathis Petrovich, Michael J. Black, and Gül Varol. TEMOS: Generating diverse human motions from textual descriptions. InEuropean Conference on Computer Vision (ECCV), 2022

  34. [42]

    The KIT Motion-Language Dataset.Big Data, 4(4):236–252, dec 2016

    Matthias Plappert, Christian Mandery, and Tamim Asfour. The KIT Motion-Language Dataset.Big Data, 4(4):236–252, dec 2016. doi: 10.1089/big.2016.0028

  35. [43]

    Popovic, Andreas G

    Marko B. Popovic, Andreas G. Hofmann, and Hugh M. Herr. Angular momentum regulation during human walking: biomechanics and control.IEEE International Conference on Robotics and Automation, 2004. Proceedings. ICRA ’04. 2004, 3:2405– 2411 V ol.3, 2004

  36. [44]

    Reisman, John P

    Darcy S. Reisman, John P. Scholz, and Gregor Schöner. Coordination underlying the control of whole body momentum during sit-to-stand.Gait & posture, 15 1:45–55, 2002

  37. [45]

    Guibas, Aaron Hertzmann, Bryan C

    Davis Rempe, Leonidas J. Guibas, Aaron Hertzmann, Bryan C. Russell, Ruben Villegas, and Jimei Yang. Contact and Human Dynamics from Monocular Video. InSymposium on Computer Animation, 2020

  38. [46]

    Davis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang, Srinath Sridhar, and Leonidas J. Guibas. HuMoR: 3D Human Motion Model for Robust Pose Estima- tion.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 11468–11479, 2021

  39. [47]

    Human Motion Prediction via Spatio-Temporal Inpainting.2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 7133–7142, 2018

    Alejandro Hernandez Ruiz, Juergen Gall, and Francesc Moreno-Noguer. Human Motion Prediction via Spatio-Temporal Inpainting.2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 7133–7142, 2018

  40. [48]

    Soshi Shimada, Vladislav Golyanik, Weipeng Xu, and Christian Theobalt. PhysCap. ACM Transactions on Graphics (TOG), 39:1 – 16, 2020

  41. [49]

    Neural monocular 3D human motion capture with physical awareness.ACM Transac- tions on Graphics (TOG), 40:1 – 15, 2021

    Soshi Shimada, Vladislav Golyanik, Weipeng Xu, Patrick P’erez, and Christian Theobalt. Neural monocular 3D human motion capture with physical awareness.ACM Transac- tions on Graphics (TOG), 40:1 – 15, 2021. 14NGUYEN ET AL: BODY MOMENTUM IN HUMAN MOTION

  42. [50]

    Soyong Shin, Juyong Kim, Eni Halilaj, and Michael J. Black. WHAM: Reconstructing World-Grounded Humans with Accurate 3D Motion.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2070–2080, 2023

  43. [51]

    The Discrete Cosine Transform.SIAM Rev., 41:135–147, 1999

    Gilbert Strang. The Discrete Cosine Transform.SIAM Rev., 41:135–147, 1999

  44. [52]

    Yu Sun, Qian Bao, Wu Liu, Tao Mei, and Michael J. Black. TRACE: 5D Temporal Regression of Avatars with Dynamic Cameras in 3D Environments.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8856–8866, 2023

  45. [53]

    Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H. Bermano. Human Motion Diffusion Model.ArXiv, abs/2209.14916, 2022

  46. [54]

    Black, and Dimitrios Tzionas

    Shashank Tripathi, Lea Muller, Chun-Hao Paul Huang, Omid Taheri, Michael J. Black, and Dimitrios Tzionas. 3D Human Pose Estimation via Intuitive Physics.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4713–4725, 2023

  47. [55]

    Black, Daniel Holden, and Carsten Stoll

    Shashank Tripathi, Omid Taheri, Christoph Lassner, Michael J. Black, Daniel Holden, and Carsten Stoll. HUMOS: Human Motion Model Conditioned on Body Shape. In European Conference on Computer Vision (ECCV), 2024

  48. [56]

    van Dieën, Sjoerd M

    Jaap H. van Dieën, Sjoerd M. Bruijn, Koen K. Lemaire, and Dinant A. Kistemaker. Simultaneous stabilizing feedback control of linear and angular momentum in human walking.bioRxiv, 2025

  49. [57]

    Black, Bodo Rosenhahn, and Gerard Pons-Moll

    Timo von Marcard, Roberto Henschel, Michael J. Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering Accurate 3D Human Pose in the Wild Using IMUs and a Moving Camera. InEuropean Conference on Computer Vision, 2018

  50. [58]

    Zero-Moment Point - Thirty Five Years of its Life.Int

    Miomir Vukobratovic and Branislav Borovac. Zero-Moment Point - Thirty Five Years of its Life.Int. J. Humanoid Robotics, 1:157–173, 2004

  51. [59]

    TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos

    Yufu Wang, Ziyun Wang, Lingjie Liu, and Kostas Daniilidis. TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos. InEuropean Conference on Computer Vision, 2024

  52. [60]

    Ong, Antoine Falisse, Shardul Sapkota, Aidan Chandra, Joshua Autton Carter, Ezio Preatoni, Benjamin Fregly, Jennifer Hicks, Scott L

    Keenon Werling, Janelle Kaneda, Alan Tan, Rishi Agarwal, Six Skov, Tom Van Wouwe, Scott Uhlrich, Nicholas Bianco, Carmichael F. Ong, Antoine Falisse, Shardul Sapkota, Aidan Chandra, Joshua Autton Carter, Ezio Preatoni, Benjamin Fregly, Jennifer Hicks, Scott L. Delp, and C. Kar...

  53. [61]

    Winkler, C

    Alexander W. Winkler, C. Dario Bellicoso, Marco Hutter, and Jonas Buchli. Gait and trajectory optimization for legged systems through phase-based end-effector parameteri- zation.IEEE Robotics and Automation Letters, 3:1560–1567, 2018

  54. [62]

    Physics-based Human Motion Estimation and Synthesis from Videos.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 11512–11521, 2021

    Kevin Xie, Tingwu Wang, Umar Iqbal, Yunrong Guo, Sanja Fidler, and Florian Shkurti. Physics-based Human Motion Estimation and Synthesis from Videos.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 11512–11521, 2021. NGUYEN ET AL: BODY MOMENTUM IN HUMAN MOTION15

  55. [63]

    Tan, Yuhong Tan, Siheng Chen, Yu Wang, Xinchao Wang, and Yanfeng Wang

    Chenxin Xu, Robby T. Tan, Yuhong Tan, Siheng Chen, Yu Wang, Xinchao Wang, and Yanfeng Wang. EqMotion: Equivariant Multi-Agent Motion Prediction with Invariant Interaction Reasoning.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1410–1420, 2023

  56. [64]

    Decoupling Human and Camera Motion from Videos in the Wild.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 21222–21232, 2023

    Vickie Ye, Georgios Pavlakos, Jitendra Malik, and Angjoo Kanazawa. Decoupling Human and Camera Motion from Videos in the Wild.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 21222–21232, 2023

  57. [65]

    Residual Force Control for Agile Human Behavior Imitation and Extended Motion Synthesis.ArXiv, abs/2006.07364, 2020

    Ye Yuan and Kris Kitani. Residual Force Control for Agile Human Behavior Imitation and Extended Motion Synthesis.ArXiv, abs/2006.07364, 2020

  58. [66]

    Ye Yuan and Kris M. Kitani. DLow: Diversifying Latent Flows for Diverse Human Motion Prediction. InEuropean Conference on Computer Vision, 2020

  59. [67]

    GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic Cameras.2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11028–11039, 2021

    Ye Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani, and Jan Kautz. GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic Cameras.2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11028–11039, 2021

  60. [68]

    Ye Yuan, Shih-En Wei, Tomas Simon, Kris Kitani, and Jason M. Saragih. SimPoE: Sim- ulated Character Control for 3D Human Pose Estimation.2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7155–7165, 2021

  61. [69]

    PhysDiff: Physics- Guided Human Motion Diffusion Model.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 15964–15975, 2022

    Ye Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat, and Jan Kautz. PhysDiff: Physics- Guided Human Motion Diffusion Model.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 15964–15975, 2022

  62. [70]

    T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations

    Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Shaoli Huang, Yong Zhang, Hong- wei Zhao, Hongtao Lu, and Xi Shen. T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...

  63. [71]

    Learning Motion Priors for 4D Human Body Capture in 3D Scenes.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 11323–11333, 2021

    Siwei Zhang, Yan Zhang, Federica Bogo, Marc Pollefeys, and Siyu Tang. Learning Motion Priors for 4D Human Body Capture in 3D Scenes.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 11323–11333, 2021

  64. [72]

    Kephart, Zijun Cui, and Qiang Ji

    Yufei Zhang, Jeffrey O. Kephart, Zijun Cui, and Qiang Ji. PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular Videos.2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2305–2317, 2024

  65. [73]

    Kephart, and Qiang Ji

    Yufei Zhang, Jeffrey O. Kephart, and Qiang Ji. Incorporating Physics Principles for Precise Human Motion Prediction. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 6164–6174, January 2024. NGUYEN ET AL: BODY MOMENTUM IN HUMAN M...

  66. [74]

    Additionally, we train TEMOS with a dynamics stability term recently proposed by HUMOS [55] to compare against TMo

    against a version of TEMOS trained with our loss function, LTMo. Additionally, we train TEMOS with a dynamics stability term recently proposed by HUMOS [55] to compare against TMo. HUMOS stability termHUMOS [ 55] dynamics stability term extends the concept of pose stability pr...

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.