Pith. sign in

REVIEW 3 major objections 4 minor 58 references

WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human

T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read WATCH reconstructs world-grounded human and camera trajectories from monocular video by analytically decomposing camera heading and softly integrating camera motion, reporting the top world-space accuracy among human-motion-centric…

desk verdict Solid incremental step in global human motion reconstruction: clean heading decomposition and consistent gains over GVHMR, but the SOTA claim is only within a method family and the DPVO dependence is under-tested. read the letter →

arxiv 2509.04600 v1 pith:UV7GBRJM submitted 2025-09-04 cs.CV

classification cs.CV
keywords globalhumanmotionreconstructionworld-spacetrajectorycamera-humancouplingheadingangledecompositionmonocularvideoSLAMcameraSMPL-Xposeestimation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Global human motion reconstruction from ordinary monocular video fails when the camera itself moves, because the person's visible motion mixes the camera's motion with their own. This paper argues that the fix is to model both motions jointly: it introduces WATCH, a single network that reconstructs the camera and the human in a shared world frame. The method's central move is to reduce the hardest orientation problem—the camera's heading around gravity—to an analytical formula that needs only the camera's roll and pitch, obtained by recursive integration of the camera's body-frame angular velocity. It then feeds camera velocity into the model as soft context rather than as a hard positional constraint, avoiding the pose discontinuities typical of direct SLAM-to-body mapping. On in-the-wild benchmarks, WATCH reports the best world-space joint errors, smoothness, and foot-contact quality among human-motion-centric methods, indicating that explicit camera-human coupling is a viable path to stable global trajectories.

What carries the argument

The load-bearing object is the analytical heading decomposition. It uses a fixed-axis Euler decomposition in a gravity-aligned world frame, $R^{\mathrm{cam}} = R^{\mathrm{cam}}_{\mathrm{yaw}} R^{\mathrm{cam}}_{\mathrm{rp}}$, and isolates heading changes by conjugating the body-frame angular velocity with the roll-pitch component: $\Delta R^{\mathrm{cam}}_{\mathrm{yaw},t} = R^{\mathrm{cam}}_{\mathrm{rp},t} \Delta R^{\mathrm{cam}}_t (R^{\mathrm{cam}}_{\mathrm{rp},t+1})^T$. This matters because it converts heading estimation into a recursion: predict only roll-pitch, then multiply the heading increments and obtain both camera and human world orientations. The second mechanism is the camera trajectory integration module: camera local velocity is fused into the embedding space through MLPs, and an auxiliary decoder predicts camera local velocity and integrates it into a trajectory under teacher-forcing consistency losses, so translation cues shape the human trajectory without imposing direct depth constraints.

What would settle it

Run WATCH on a sequence with ground-truth camera and human trajectories in which the camera points nearly straight down for part of the clip; the paper's supplementary reports that this near-vertical pose makes the horizontal projection used in heading decomposition ill-conditioned. If world-space joint error spikes precisely in that segment, the analytical decomposition's stability assumption is violated. A second check: add controlled scale or drift errors to the camera angular velocity and measure W-MPJPE100 as a function of corruption; an early steep rise would indict the recursive heading integration as the bottleneck.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that world-space human orientation can be recovered without estimating heading directly. WATCH decomposes the camera rotation into a heading (yaw) part and a roll-pitch part, $R^{\mathrm{cam}}_t = R^{\mathrm{cam}}_{\mathrm{yaw},t} R^{\mathrm{cam}}_{\mathrm{rp},t}$. Because the body-frame angular velocity $\Delta R^{\mathrm{cam}}_t$ measures rotation changes in the camera's local coordinates, the heading increment can be isolated analytically: $\Delta R^{\mathrm{cam}}_{\mathrm{yaw},t} = R^{\mathrm{cam}}_{\mathrm{rp},t} \Delta R^{\mathrm{cam}}_t (R^{\mathrm{cam}}_{\mathrm{rp},t+1})^T$. So the network only needs to predict the more intuitive roll-pitch component, and the heading follows by integration; combining it with the camera-space human orientation gives the world-space orientation $R^{\mathrm{h,w}}_t = R^{\mathrm{cam}}_{\mathrm{yaw},t} R^{\mathrm{cam}}_{\mathrm{rp},t} R^{\mathrm{h,c}}_t$. For translation, the paper deliberately avoids hard-decoding camera trajectories. Instead, camera local velocity is encoded as a soft contextual feature, the decoder predicts both human and camera local velocities, and trajectories are produced by integration with teacher-forcing consistency losses. The combined effect, according to the paper's experiments, is lower world-aligned joint error and less jitter and foot sliding than previous human-motion-centric approaches, while retaining camera-space accuracy.

Load-bearing premise

The load-bearing premise is that the camera's turning-rate estimates from odometry or a gyroscope are accurate enough that repeatedly adding up the extracted heading does not drift, and that the network's estimate of how the camera tilts relative to gravity is correct; if either fails, the person's world-space orientation and trajectory degrade.

Editorial extensions

If this is right

  • If the central claim holds, any human-motion-centric pipeline can adopt the heading decomposition to obtain world-space orientation from only roll-pitch and angular velocity, making camera-orientation cues cheaper and more interpretable than geometric projection operators.
  • Soft camera-trajectory integration should improve global trajectory accuracy and smoothness without the pose discontinuities that direct SLAM-to-human mapping produces.
  • The benefit is largest when the camera moves; on the static-camera RICH setting, the paper reports camera-space results slightly below the strongest baseline, so the gain is tied to dynamic cameras.
  • The framework retains its advantage when camera motion comes from odometry as well as from an ideal gyroscope, so the method is intended to transfer to realistic SLAM-derived inputs.
  • Joint camera and human velocity decoding gives a route to end-to-end training that avoids cascading errors from separately estimated camera trajectories over long sequences.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Going beyond the paper: because heading is constructed analytically, one could calibrate the roll-pitch network with cheap IMU gravity measurements instead of full camera-pose ground truth, easing deployment on unseen handheld rigs.
  • Going beyond the paper: feeding the model deliberately corrupted camera angular velocities with controlled scale and drift errors and measuring world-space joint error as a function of corruption would isolate how much of the gain depends on odometry accuracy, a sensitivity analysis the paper does not report.
  • Going beyond the paper: the soft camera-trajectory module is architecture-agnostic enough that it could be grafted onto other human-centric baselines; if it transfers, camera-translation integration would become a plug-in competence rather than a bespoke design.
  • Going beyond the paper: the near-vertical camera pose case flagged in the supplementary suggests that augmenting training with downward- and upward-looking footage, or reparameterizing heading for that regime, would be a natural stress test.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces WATCH, a unified framework for world-grounded human pose and trajectory reconstruction from monocular video. The core technical contributions are an analytic heading-angle decomposition that extracts the camera's world-yaw component from body-frame angular velocity using only predicted roll and pitch, and a camera trajectory integration module that fuses camera velocity features and supervises camera and human trajectory consistency. Experiments on RICH, 3DPW, and EMDB-2 report improvements over human-motion-centric baselines such as GVHMR and WHAM, with particular gains in global trajectory metrics and smoothness. The central claim is state-of-the-art end-to-end trajectory reconstruction.

Significance. If the reported results hold, the work provides a simple, interpretable alternative to the view-gravity operator of GVHMR and a soft way to exploit camera translation cues without hard-decoding camera trajectories. The derivation in Section 3.2 is mathematically sound: Eq. (2) correctly isolates the heading increment by conjugation of the body-frame angular velocity, and the rotation-invariance property of Eq. (4) is valid. The teacher-forcing trajectory losses are a sensible way to decouple orientation and velocity supervision, and the commitment to release code supports reproducibility. The main open risk is that the central SOTA claim is broader than what Table 1 supports, and the recursive integrator's dependence on DPVO rotation accuracy is not stress-tested; these are fixable with additional analysis and experiments.

major comments (3)
  1. [Abstract and Section 4.3] The abstract's unqualified claim of "state-of-the-art performance in end-to-end trajectory reconstruction" is not supported by Table 1: on EMDB-2, TRAM and PromptHMR-vid report W-MPJPE100 of 222.4 and 216.5 mm and RTE of 1.4 and 1.3, whereas WATCH reports 269.3 mm and 1.7. Section 4.3 narrows the claim to "leading performance among human motion-centric methods," which is accurate, but the abstract and contributions section do not include that qualification. The claims should be aligned, or the comparisons with camera-trajectory-centric methods should be justified explicitly rather than deferred to qualitative physical-plausibility arguments.
  2. [Section 3.2, Eqs. (2)-(3), and Table 3] The recursive heading integration in Eq. (3) is an open-loop integrator: any per-frame bias or temporally correlated error in the DPVO-derived body-frame angular velocity Delta_R_cam,t accumulates in the heading and propagates into the human world orientation through Eq. (4). The evidence offered in Table 3 compares only two operating points (DPVO and GT gyro) and shows small aggregate gaps, which does not bound the failure regime. The paper should quantify sensitivity to DPVO rotation errors, for example by injecting controlled rotation noise or drift, reporting heading error as a function of sequence length, and measuring the effect of the network-estimated roll-pitch error on the extracted heading. Without this, the load-bearing robustness claim in Section 4.5 is not demonstrated.
  3. [Supplementary A.2, Eq. (9)] The near-vertical fallback in Eq. (9) assigns a fixed heading direction when the horizontal projection of the camera forward vector is below threshold epsilon; this can produce a discontinuous jump in the heading, and therefore in the human world orientation via Eq. (4). The paper does not report how often this degenerate regime occurs in RICH or EMDB-2, nor its effect on the reported metrics. Since the method is intended for in-the-wild moving cameras, this acknowledged ill-conditioned case should be quantified or explicitly discussed as a limitation.
minor comments (4)
  1. [Table 1] SLAHMR appears in both the "Camera-trajectory-centric" and "Human-motion-centric" groups with identical numbers; this is confusing and should be clarified or the method should be placed in only one category.
  2. [Figure 6 caption] The caption contains a typo: "trajctory" should be "trajectory".
  3. [Table 3] In the arXiv rendering of Table 3, the labels and method names are concatenated (e.g., "w/ DPVOWHAM"); the table should be reformatted so that rows and columns are clearly separated.
  4. [Section 3.3, Eqs. (5)-(6)] The "camera trajectory integration mechanism" as described is an additive feature embedding plus an auxiliary velocity integration loss; the distinction from simple feature fusion and the claim of being "inspired by world models" would benefit from a more concrete explanation of why this particular integration is preferable to hard-decoding approaches.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the heading decomposition is a rotation-matrix identity and the camera-trajectory cues are external supervised inputs, not fitted predictions.

full rationale

WATCH's central claims rest on the analytical heading decomposition of Eqs. (1)-(3), which is a rotation-matrix identity: given camera rotation R_cam and estimated roll-pitch R_rp, the heading increment Delta_R_yaw = R_rp,t * Delta_R_cam,t * (R_rp,t+1)^T is exactly the isolated heading step by construction, and R_yaw is integrated open-loop. This is not a fitted target or a prediction that assumes its own conclusion; it is a deterministic geometric transformation. Human world orientation in Eq. (4) is then a product of this integrated camera heading, roll-pitch, and camera-space human orientation, all supervised with external ground truth such as camera extrinsics, gyro data, and SMPL annotations. The camera trajectory integration in Eqs. (5)-(6) predicts velocities and integrates them, using camera motion as contextual features and as an auxiliary supervised task, not as the direct source of the human trajectory. The only notable self-citation is GVHMR (Shen et al. 2024), which shares authors with the present paper, but GVHMR is used as an empirical baseline on external benchmarks, not as the justification for WATCH's derivation; the decomposition is derived in-paper from rotation algebra. Therefore no load-bearing step reduces to its own input. Concerns about DPVO angular-velocity drift or degenerate near-vertical camera cases are robustness and sensitivity risks, not circularity.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central pipeline relies on standard rotation algebra, on the assumption that gravity defines the world vertical axis and is recoverable from network-estimated roll-pitch, and on the availability of camera angular velocity from DPVO/SLAM. Hyperparameters such as loss weights and contact thresholds are hand-chosen; no new physical entity is introduced.

free parameters (4)
  • Trajectory loss weights lambda_h and lambda_cam = 1 and 1
    Chosen by hand in Supplementary B.5; the balance between human and camera trajectory supervision depends on them.
  • Depth threshold for loss masking = 0.3 m
    Used in standard reconstruction losses following CLIFF and WHAM in Supplementary B.4; affects which 3D joints and vertices contribute to the loss.
  • Static contact velocity threshold = 0.15 m/s
    Contact labels are generated automatically with this threshold in Supplementary B.4; it affects foot-sliding supervision.
  • Degenerate heading epsilon = 1e-6
    Threshold in Supplementary A.2 for when the camera forward vector projected onto the horizontal plane is too small; it changes heading construction in near-vertical views.
assumptions (5)
  • standard math Every rotation R can be decomposed as R = Ryaw Rrp, with Ryaw a rotation about the world Y-axis and Rrp the roll-pitch component.
    Used in Eq. (1); fixed-axis Euler decomposition is a standard result.
  • domain assumption The world Y-axis is aligned with gravity and can be recovered from camera roll and pitch.
    Eqs. (4) and (8)-(10) in Section 3.2 and Supplementary A.2 rely on the vertical axis being known from R_rp; the network is trained to estimate roll-pitch from images.
  • domain assumption The camera body-frame angular velocity from DPVO/SLAM or GT gyro is accurate enough for recursive yaw integration.
    Eqs. (2)-(3) integrate heading from this angular velocity; Section 4.3 and Table 3 use DPVO or GT gyro as inputs, and the paper does not quantify sensitivity to SLAM error.
  • standard math Rotation composition of the recursive heading increments yields a valid orthonormal rotation.
    Eq. (3) assumes products of SO(3) elements stay in SO(3); this is standard.
  • domain assumption The camera-space human orientation Rh,c estimated by the network is a valid rotation from the body to the camera frame, so world orientation is R_cam Rh,c.
    Eq. (4) relies on this definition of Rh,c; if the network's camera-space orientation is not consistently framed, the world orientation will be wrong.

how reviews work

0 comments
Cite this review

Pith. "Pith review of WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human." pith.science (2026). https://pith.science/paper/UV7GBRJM

@misc{pith2026250904600,
  author       = {Pith},
  title        = {Pith review of: WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UV7GBRJM}},
  note         = {Machine review of arXiv:2509.04600}
}
read the original abstract

Global human motion reconstruction from in-the-wild monocular videos is increasingly demanded across VR, graphics, and robotics applications, yet requires accurate mapping of human poses from camera to world coordinates-a task challenged by depth ambiguity, motion ambiguity, and the entanglement between camera and human movements. While human-motion-centric approaches excel in preserving motion details and physical plausibility, they suffer from two critical limitations: insufficient exploitation of camera orientation information and ineffective integration of camera translation cues. We present WATCH (World-aware Allied Trajectory and pose reconstruction for Camera and Human), a unified framework addressing both challenges. Our approach introduces an analytical heading angle decomposition technique that offers superior efficiency and extensibility compared to existing geometric methods. Additionally, we design a camera trajectory integration mechanism inspired by world models, providing an effective pathway for leveraging camera translation information beyond naive hard-decoding approaches. Through experiments on in-the-wild benchmarks, WATCH achieves state-of-the-art performance in end-to-end trajectory reconstruction. Our work demonstrates the effectiveness of jointly modeling camera-human motion relationships and offers new insights for addressing the long-standing challenge of camera translation integration in global human motion reconstruction. The code will be available publicly.

Figures

Figures reproduced from arXiv: 2509.04600 by the authors.

Figure 2
Figure 2. Left: Model-predicted top-down view of camera poses and trajectories (top) alongside human poses and trajecto [PITH_FULL_IMAGE:figures/full_fig_p001_2.png] view at source ↗
Figure 3
Figure 3. Overview of the WATCH framework. Given a monocular video sequence, we extract multi-modal conditions in￾cluding ViT-based image features fimg, human bounding boxesfbbox, 2D keypoints fkp2d from ViTPose, and camera rotation angles fR∆ with local velocities fVcam. These conditions are processed through a RoPE Transformer backbone to predict camera-space human poses, camera roll-pitch angles, and local velocities. Thro… view at source ↗
Figure 4
Figure 4. The illustration of the relationship between the camera and the human body. The camera angular veloc￾ity ∆Rcam is expressed in the body-frame, while the head￾ing angular velocity ∆Rcam yaw represents the rotational com￾ponent around the gravity axis g. By knowing the roll and pitch angles (Rcam rp ), the camera heading orientation would be obtained using our heading decomposition approach. 3.2 Analytical Heading Dec… view at source ↗
Figures from the paper (5 more)
Figure 6
Figure 6. Figure 6: bird-eye view of trajctory comparison with [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 5
Figure 5. Figure 5: Qualitative comparison with PromptHMR￾vid (Wang et al. 2025b) and GVHMR (Shen et al. 2024) on global human motion estimation, frames in gray are failure cases. 4.5 Ablation Study [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 7
Figure 7. Figure 7: Two extreme occlusion cases where approximately [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: Comparison between WATCH and Prompt-HMR-vid on two fundamental motion scenarios. Top row: sitting scenario [PITH_FULL_IMAGE:figures/full_fig_p013_8.png]
Figure 9
Figure 9. Figure 9: WATCH’s performance on a long sequence with diverse motion patterns: (1) straight-line walking, (2) cross-stepping [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

58 extracted references · 32 canonical work pages

  1. [1]

    P.; Li, J.; Vetrivel, K.; Agarwal, R.; Wu, J.; Gopinath, D.; Clegg, A

    Ara \'u jo, J. P.; Li, J.; Vetrivel, K.; Agarwal, R.; Wu, J.; Gopinath, D.; Clegg, A. W.; and Liu, K. 2023. Circle: Capture in rich contextual environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 21211--21221

  2. [2]

    J.; Patel, P.; Tesch, J.; and Yang, J

    Black, M. J.; Patel, P.; Tesch, J.; and Yang, J. 2023. BEDLAM: A Synthetic Dataset of Bodies Exhibiting Detailed Lifelike Animated Motion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 8726--8737

  3. [3]

    Bogo, F.; Kanazawa, A.; Lassner, C.; Gehler, P.; Romero, J.; and Black, M. J. 2016. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In Computer Vision--ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part V 14, 561--578. Springer

  4. [4]

    Y.; and Lee, K

    Choi, H.; Moon, G.; Chang, J. Y.; and Lee, K. M. 2021. Beyond static features for temporally consistent 3d human pose and shape from a video. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 1964--1973

  5. [5]

    Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929

  6. [6]

    TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation

    Dwivedi, S. K.; Sun, Y.; Patel, P.; Feng, Y.; and Black, M. J. 2024. TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation. arXiv:2404.16752

  7. [7]

    Fu, Z.; Zhao, Q.; Wu, Q.; Wetzstein, G.; and Finn, C. 2024. Humanplus: Humanoid shadowing and imitation from humans. arXiv preprint arXiv:2406.10454

  8. [8]

    Goel, S.; Pavlakos, G.; Rajasegaran, J.; Kanazawa, A.; and Malik, J. 2023. Humans in 4D: Reconstructing and Tracking Humans with Transformers. arXiv preprint arXiv:2305.20091

Show all 58 references
  1. [9]

    F.; Choi, C.; Schaefer, S.; and Leutenegger, S

    Henning, D. F.; Choi, C.; Schaefer, S.; and Leutenegger, S. 2023. BodySLAM++: Fast and Tightly-Coupled Visual-Inertial Camera and Human Motion Tracking. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 3781--3788. IEEE

  2. [10]

    F.; Laidlow, T.; and Leutenegger, S

    Henning, D. F.; Laidlow, T.; and Leutenegger, S. 2022. BodySLAM: joint camera localisation, mapping, and human motion tracking. In European Conference on Computer Vision, 656--673. Springer

  3. [11]

    P.; Yi, H.; H \"o schle, M.; Safroshkin, M.; Alexiadis, T.; Polikovsky, S.; Scharstein, D.; and Black, M

    Huang, C.-H. P.; Yi, H.; H \"o schle, M.; Safroshkin, M.; Alexiadis, T.; Polikovsky, S.; Scharstein, D.; and Black, M. J. 2022. Capturing and Inferring Dense Full-Body Human-Scene Contact. In Proceedings IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 13274--13285

  4. [12]

    Ionescu, C.; Papava, D.; Olaru, V.; and Sminchisescu, C. 2013. Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE transactions on pattern analysis and machine intelligence, 36(7): 1325--1339

  5. [13]

    J.; Jacobs, D

    Kanazawa, A.; Black, M. J.; Jacobs, D. W.; and Malik, J. 2018. End-to-end recovery of human shape and pose. In Proceedings of the IEEE conference on computer vision and pattern recognition, 7122--7131

  6. [14]

    Y.; Felsen, P.; and Malik, J

    Kanazawa, A.; Zhang, J. Y.; Felsen, P.; and Malik, J. 2019. Learning 3d human dynamics from video. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 5614--5623

  7. [15]

    J.; and Hilliges, O

    Kaufmann, M.; Song, J.; Guo, C.; Shen, K.; Jiang, T.; Tang, C.; Z \'a rate, J. J.; and Hilliges, O. 2023. EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 14632--14643

  8. [16]

    P.; Hilliges, O.; and Black, M

    Kocabas, M.; Huang, C.-H. P.; Hilliges, O.; and Black, M. J. 2021. PARE: Part attention regressor for 3D human body estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 11127--11137

  9. [17]

    J.; Hilliges, O.; Kautz, J.; and Iqbal, U

    Kocabas, M.; Yuan, Y.; Molchanov, P.; Guo, Y.; Black, M. J.; Hilliges, O.; Kautz, J.; and Iqbal, U. 2023. PACE: Human and Camera Motion Estimation from in-the-wild Videos. arXiv preprint arXiv:2310.13768

  10. [18]

    Kocabas, M.; et al. 2020. Vibe: Video inference for human body pose and shape estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 5253--5263

  11. [19]

    J.; and Daniilidis, K

    Kolotouros, N.; Pavlakos, G.; Black, M. J.; and Daniilidis, K. 2019. Learning to reconstruct 3D human pose and shape via model-fitting in the loop. In Proceedings of the IEEE/CVF international conference on computer vision, 2252--2261

  12. [20]

    Li, J.; Bian, S.; Xu, C.; Liu, G.; Yu, G.; and Lu, C. 2022 a . D &D: Learning Human Dynamics from Dynamic Camera. In European Conference on Computer Vision, 479--496. Springer

  13. [21]

    Li, J.; Cao, J.; Zhang, H.; Rempe, D.; Kautz, J.; Iqbal, U.; and Yuan, Y. 2025. GENMO: A GENeralist Model for Human MOtion. arXiv preprint arXiv:2505.01425

  14. [22]

    Li, J.; Xu, C.; Chen, Z.; Bian, S.; Yang, L.; and Lu, C. 2021. Hybrik: A hybrid analytical-neural inverse kinematics solution for 3d human pose and shape estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 3383--3393

  15. [23]

    Li, Z.; Liu, J.; Zhang, Z.; Xu, S.; and Yan, Y. 2022 b . Cliff: Carrying location information in full frames into human pose and shape estimation. In Computer Vision--ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23--27, 2022, Proceedings, Part V, 590--606. Springer

  16. [24]

    Liu, K.; Fu, Y.; Yuan, W.; Lin, J.; Li, P.; Gu, X.; Qiu, L.; Wang, H.; Dong, Z.; and Han, X. 2025 a . Motions as Queries: One-Stage Multi-Person Holistic Human Motion Capture. In Proceedings of the Computer Vision and Pattern Recognition Conference, 17529--17539

  17. [25]

    Liu, Z.; Lin, J.; Wu, W.; and Zhou, B. 2025 b . Joint optimization for 4d human-scene reconstruction in the wild. arXiv preprint arXiv:2501.02158

  18. [26]

    Loper, M.; Mahmood, N.; Romero, J.; Pons-Moll, G.; and Black, M. J. 2015. SMPL: A skinned multi-person linear model. ACM transactions on graphics (TOG), 34(6): 1--16

  19. [27]

    A.; and Kitani, K

    Luo, Z.; Golestaneh, S. A.; and Kitani, K. M. 2020. 3d human motion estimation via motion compression and refinement. In Proceedings of the Asian Conference on Computer Vision

  20. [28]

    F.; Pons-Moll, G.; and Black, M

    Mahmood, N.; Ghorbani, N.; Troje, N. F.; Pons-Moll, G.; and Black, M. J. 2019. AMASS: Archive of motion capture as surface shapes. In Proceedings of the IEEE/CVF international conference on computer vision, 5442--5451

  21. [29]

    Mur-Artal, R.; and Tard \'o s, J. D. 2017. Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras. IEEE transactions on robotics, 33(5): 1255--1262

  22. [30]

    Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. 2023. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193

  23. [31]

    Patel, P.; and Black, M. J. 2025. CameraHMR: Aligning People with Perspective. In International Conference on 3D Vision (3DV)

  24. [32]

    Patel, P.; et al. 2025. CameraHMR: Aligning People with Perspective. arXiv preprint arXiv:2411.08128

  25. [33]

    A.; Tzionas, D.; and Black, M

    Pavlakos, G.; Choutas, V.; Ghorbani, N.; Bolkart, T.; Osman, A. A.; Tzionas, D.; and Black, M. J. 2019. Expressive body capture: 3d hands, face, and body from a single image. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 10975--10985

  26. [34]

    Shen, X.; Yang, Z.; Wang, X.; Ma, J.; Zhou, C.; and Yang, Y. 2023. Global-to-Local Modeling for Video-based 3D Human Pose and Shape Estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 8887--8896

  27. [35]

    Shen, Z.; Pi, H.; Xia, Y.; Cen, Z.; Peng, S.; Hu, Z.; Bao, H.; Hu, R.; and Zhou, X. 2024. World-Grounded Human Motion Recovery via Gravity-View Coordinates. In SIGGRAPH Asia 2024 Conference Papers, 1–11. ACM

  28. [36]

    Shimada, S.; Golyanik, V.; Xu, W.; and Theobalt, C. 2020. Physcap: Physically plausible monocular 3d motion capture in real time. ACM Transactions on Graphics (ToG), 39(6): 1--16

  29. [37]

    Shin, S.; Kim, J.; Halilaj, E.; and Black, M. J. 2023. WHAM: Reconstructing World-grounded Humans with Accurate 3D Motion. arXiv preprint arXiv:2312.07531

  30. [38]

    Shrestha, A.; Liu, P.; Ros, G.; Yuan, K.; and Fern, A. 2024. Generating physically realistic and directable human motions from multi-modal inputs. In European Conference on Computer Vision, 1--17. Springer

  31. [39]

    Sun, Y.; Bao, Q.; Liu, W.; Mei, T.; and Black, M. J. 2023. TRACE: 5D temporal regression of avatars with dynamic cameras in 3D environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 8856--8866

  32. [40]

    Teed, Z.; and Deng, J. 2021. Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras. Advances in neural information processing systems, 34: 16558--16569

  33. [41]

    Teed, Z.; et al. 2023. Deep Patch Visual Odometry. arXiv:2208.04726

  34. [42]

    J.; Rosenhahn, B.; and Pons-Moll, G

    Von Marcard, T.; Henschel, R.; Black, M. J.; Rosenhahn, B.; and Pons-Moll, G. 2018. Recovering accurate 3d human pose in the wild using imus and a moving camera. In Proceedings of the European conference on computer vision (ECCV), 601--617

  35. [43]

    Wan, Z.; Li, Z.; Tian, M.; Liu, J.; Yi, S.; and Li, H. 2021. Encoder-decoder with Multi-level Attention for 3D Human Shape and Pose Estimation. In The IEEE International Conference on Computer Vision (ICCV)

  36. [44]

    Wang, J.; Chen, M.; Karaev, N.; Vedaldi, A.; Rupprecht, C.; and Novotny, D. 2025 a . VGGT: Visual Geometry Grounded Transformer. arXiv:2503.11651

  37. [45]

    Wang, S.; Leroy, V.; Cabon, Y.; Chidlovskii, B.; and Revaud, J. 2024 a . DUSt3R: Geometric 3D Vision Made Easy. arXiv:2312.14132

  38. [46]

    Wang, Y.; and Daniilidis, K. 2023. Refit: Recurrent fitting network for 3d human recovery. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 14644--14654

  39. [47]

    J.; and Kocabas, M

    Wang, Y.; Sun, Y.; Patel, P.; Daniilidis, K.; Black, M. J.; and Kocabas, M. 2025 b . PromptHMR: Promptable Human Mesh Recovery. arXiv:2504.06397

  40. [48]

    Wang, Y.; Wang, Z.; Liu, L.; and Daniilidis, K. 2024 b . TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos. arXiv:2403.17346

  41. [49]

    Wei, W.-L.; Lin, J.-C.; Liu, T.-L.; and Liao, H.-Y. M. 2022. Capturing humans in motion: Temporal-attentive 3D human pose and shape estimation from monocular video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13211--13220

  42. [50]

    Xie, K.; Wang, T.; Iqbal, U.; Guo, Y.; Fidler, S.; and Shkurti, F. 2021. Physics-based human motion estimation and synthesis from videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision (CVPR)

  43. [51]

    Ye, V.; Pavlakos, G.; Malik, J.; and Kanazawa, A. 2023 a . Decoupling human and camera motion from videos in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 21222--21232

  44. [52]

    Ye, V.; Pavlakos, G.; Malik, J.; and Kanazawa, A. 2023 b . Decoupling Human and Camera Motion from Videos in the Wild. arXiv:2302.12827

  45. [53]

    Yin, W.; Cai, Z.; Wang, R.; Wang, F.; Wei, C.; Mei, H.; Xiao, W.; Yang, Z.; Sun, Q.; Yamashita, A.; Liu, Z.; and Yang, L. 2024. WHAC: World-grounded Humans and Cameras. arXiv:2403.12959

  46. [54]

    Yuan, Y.; Iqbal, U.; Molchanov, P.; Kitani, K.; and Kautz, J. 2022. GLAMR: Global occlusion-aware human mesh recovery with dynamic cameras. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 11038--11049

  47. [55]

    Zhang, H.; Tian, Y.; Zhou, X.; Ouyang, W.; Liu, Y.; Wang, L.; and Sun, Z. 2021. Pymaf: 3d human pose and shape regression with pyramidal mesh alignment feedback loop. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 11446--11456

  48. [56]

    O.; Cui, Z.; and Ji, Q

    Zhang, Y.; Kephart, J. O.; Cui, Z.; and Ji, Q. 2024. Physpt: Physics-aware pretrained transformer for estimating human dynamics from monocular videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

  49. [57]

    F.; Wang, H.; and Zhang, L

    Zhang, Y.; Wu, G.; Chen, L.-H.; Zhao, Z.; Lin, J.; Jiang, X.; Wu, J.; Li, Z.; Yang, H. F.; Wang, H.; and Zhang, L. 2025. HumanMM: Global Human Motion Recovery from Multi-shot Videos. arXiv:2503.07597

  50. [58]

    Y.; Raj, B.; Xu, M.; Yang, J.; and Huang, C.-H

    Zhao, Y.; Wang, T. Y.; Raj, B.; Xu, M.; Yang, J.; and Huang, C.-H. P. 2024. Synergistic Global-space Camera and Human Reconstruction from Videos. arXiv:2405.14855

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.