REVIEW 3 major objections 4 minor 58 references
WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human
T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read WATCH reconstructs world-grounded human and camera trajectories from monocular video by analytically decomposing camera heading and softly integrating camera motion, reporting the top world-space accuracy among human-motion-centric…
desk verdict Solid incremental step in global human motion reconstruction: clean heading decomposition and consistent gains over GVHMR, but the SOTA claim is only within a method family and the DPVO dependence is under-tested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the analytical heading decomposition. It uses a fixed-axis Euler decomposition in a gravity-aligned world frame, $R^{\mathrm{cam}} = R^{\mathrm{cam}}_{\mathrm{yaw}} R^{\mathrm{cam}}_{\mathrm{rp}}$, and isolates heading changes by conjugating the body-frame angular velocity with the roll-pitch component: $\Delta R^{\mathrm{cam}}_{\mathrm{yaw},t} = R^{\mathrm{cam}}_{\mathrm{rp},t} \Delta R^{\mathrm{cam}}_t (R^{\mathrm{cam}}_{\mathrm{rp},t+1})^T$. This matters because it converts heading estimation into a recursion: predict only roll-pitch, then multiply the heading increments and obtain both camera and human world orientations. The second mechanism is the camera trajectory integration module: camera local velocity is fused into the embedding space through MLPs, and an auxiliary decoder predicts camera local velocity and integrates it into a trajectory under teacher-forcing consistency losses, so translation cues shape the human trajectory without imposing direct depth constraints.
What would settle it
Run WATCH on a sequence with ground-truth camera and human trajectories in which the camera points nearly straight down for part of the clip; the paper's supplementary reports that this near-vertical pose makes the horizontal projection used in heading decomposition ill-conditioned. If world-space joint error spikes precisely in that segment, the analytical decomposition's stability assumption is violated. A second check: add controlled scale or drift errors to the camera angular velocity and measure W-MPJPE100 as a function of corruption; an early steep rise would indict the recursive heading integration as the bottleneck.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that world-space human orientation can be recovered without estimating heading directly. WATCH decomposes the camera rotation into a heading (yaw) part and a roll-pitch part, $R^{\mathrm{cam}}_t = R^{\mathrm{cam}}_{\mathrm{yaw},t} R^{\mathrm{cam}}_{\mathrm{rp},t}$. Because the body-frame angular velocity $\Delta R^{\mathrm{cam}}_t$ measures rotation changes in the camera's local coordinates, the heading increment can be isolated analytically: $\Delta R^{\mathrm{cam}}_{\mathrm{yaw},t} = R^{\mathrm{cam}}_{\mathrm{rp},t} \Delta R^{\mathrm{cam}}_t (R^{\mathrm{cam}}_{\mathrm{rp},t+1})^T$. So the network only needs to predict the more intuitive roll-pitch component, and the heading follows by integration; combining it with the camera-space human orientation gives the world-space orientation $R^{\mathrm{h,w}}_t = R^{\mathrm{cam}}_{\mathrm{yaw},t} R^{\mathrm{cam}}_{\mathrm{rp},t} R^{\mathrm{h,c}}_t$. For translation, the paper deliberately avoids hard-decoding camera trajectories. Instead, camera local velocity is encoded as a soft contextual feature, the decoder predicts both human and camera local velocities, and trajectories are produced by integration with teacher-forcing consistency losses. The combined effect, according to the paper's experiments, is lower world-aligned joint error and less jitter and foot sliding than previous human-motion-centric approaches, while retaining camera-space accuracy.
Load-bearing premise
The load-bearing premise is that the camera's turning-rate estimates from odometry or a gyroscope are accurate enough that repeatedly adding up the extracted heading does not drift, and that the network's estimate of how the camera tilts relative to gravity is correct; if either fails, the person's world-space orientation and trajectory degrade.
Editorial extensions
If this is right
- If the central claim holds, any human-motion-centric pipeline can adopt the heading decomposition to obtain world-space orientation from only roll-pitch and angular velocity, making camera-orientation cues cheaper and more interpretable than geometric projection operators.
- Soft camera-trajectory integration should improve global trajectory accuracy and smoothness without the pose discontinuities that direct SLAM-to-human mapping produces.
- The benefit is largest when the camera moves; on the static-camera RICH setting, the paper reports camera-space results slightly below the strongest baseline, so the gain is tied to dynamic cameras.
- The framework retains its advantage when camera motion comes from odometry as well as from an ideal gyroscope, so the method is intended to transfer to realistic SLAM-derived inputs.
- Joint camera and human velocity decoding gives a route to end-to-end training that avoids cascading errors from separately estimated camera trajectories over long sequences.
Reading between the lines
- Going beyond the paper: because heading is constructed analytically, one could calibrate the roll-pitch network with cheap IMU gravity measurements instead of full camera-pose ground truth, easing deployment on unseen handheld rigs.
- Going beyond the paper: feeding the model deliberately corrupted camera angular velocities with controlled scale and drift errors and measuring world-space joint error as a function of corruption would isolate how much of the gain depends on odometry accuracy, a sensitivity analysis the paper does not report.
- Going beyond the paper: the soft camera-trajectory module is architecture-agnostic enough that it could be grafted onto other human-centric baselines; if it transfers, camera-translation integration would become a plug-in competence rather than a bespoke design.
- Going beyond the paper: the near-vertical camera pose case flagged in the supplementary suggests that augmenting training with downward- and upward-looking footage, or reparameterizing heading for that regime, would be a natural stress test.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces WATCH, a unified framework for world-grounded human pose and trajectory reconstruction from monocular video. The core technical contributions are an analytic heading-angle decomposition that extracts the camera's world-yaw component from body-frame angular velocity using only predicted roll and pitch, and a camera trajectory integration module that fuses camera velocity features and supervises camera and human trajectory consistency. Experiments on RICH, 3DPW, and EMDB-2 report improvements over human-motion-centric baselines such as GVHMR and WHAM, with particular gains in global trajectory metrics and smoothness. The central claim is state-of-the-art end-to-end trajectory reconstruction.
Significance. If the reported results hold, the work provides a simple, interpretable alternative to the view-gravity operator of GVHMR and a soft way to exploit camera translation cues without hard-decoding camera trajectories. The derivation in Section 3.2 is mathematically sound: Eq. (2) correctly isolates the heading increment by conjugation of the body-frame angular velocity, and the rotation-invariance property of Eq. (4) is valid. The teacher-forcing trajectory losses are a sensible way to decouple orientation and velocity supervision, and the commitment to release code supports reproducibility. The main open risk is that the central SOTA claim is broader than what Table 1 supports, and the recursive integrator's dependence on DPVO rotation accuracy is not stress-tested; these are fixable with additional analysis and experiments.
major comments (3)
- [Abstract and Section 4.3] The abstract's unqualified claim of "state-of-the-art performance in end-to-end trajectory reconstruction" is not supported by Table 1: on EMDB-2, TRAM and PromptHMR-vid report W-MPJPE100 of 222.4 and 216.5 mm and RTE of 1.4 and 1.3, whereas WATCH reports 269.3 mm and 1.7. Section 4.3 narrows the claim to "leading performance among human motion-centric methods," which is accurate, but the abstract and contributions section do not include that qualification. The claims should be aligned, or the comparisons with camera-trajectory-centric methods should be justified explicitly rather than deferred to qualitative physical-plausibility arguments.
- [Section 3.2, Eqs. (2)-(3), and Table 3] The recursive heading integration in Eq. (3) is an open-loop integrator: any per-frame bias or temporally correlated error in the DPVO-derived body-frame angular velocity Delta_R_cam,t accumulates in the heading and propagates into the human world orientation through Eq. (4). The evidence offered in Table 3 compares only two operating points (DPVO and GT gyro) and shows small aggregate gaps, which does not bound the failure regime. The paper should quantify sensitivity to DPVO rotation errors, for example by injecting controlled rotation noise or drift, reporting heading error as a function of sequence length, and measuring the effect of the network-estimated roll-pitch error on the extracted heading. Without this, the load-bearing robustness claim in Section 4.5 is not demonstrated.
- [Supplementary A.2, Eq. (9)] The near-vertical fallback in Eq. (9) assigns a fixed heading direction when the horizontal projection of the camera forward vector is below threshold epsilon; this can produce a discontinuous jump in the heading, and therefore in the human world orientation via Eq. (4). The paper does not report how often this degenerate regime occurs in RICH or EMDB-2, nor its effect on the reported metrics. Since the method is intended for in-the-wild moving cameras, this acknowledged ill-conditioned case should be quantified or explicitly discussed as a limitation.
minor comments (4)
- [Table 1] SLAHMR appears in both the "Camera-trajectory-centric" and "Human-motion-centric" groups with identical numbers; this is confusing and should be clarified or the method should be placed in only one category.
- [Figure 6 caption] The caption contains a typo: "trajctory" should be "trajectory".
- [Table 3] In the arXiv rendering of Table 3, the labels and method names are concatenated (e.g., "w/ DPVOWHAM"); the table should be reformatted so that rows and columns are clearly separated.
- [Section 3.3, Eqs. (5)-(6)] The "camera trajectory integration mechanism" as described is an additive feature embedding plus an auxiliary velocity integration loss; the distinction from simple feature fusion and the claim of being "inspired by world models" would benefit from a more concrete explanation of why this particular integration is preferable to hard-decoding approaches.
Circularity Check
No significant circularity: the heading decomposition is a rotation-matrix identity and the camera-trajectory cues are external supervised inputs, not fitted predictions.
full rationale
WATCH's central claims rest on the analytical heading decomposition of Eqs. (1)-(3), which is a rotation-matrix identity: given camera rotation R_cam and estimated roll-pitch R_rp, the heading increment Delta_R_yaw = R_rp,t * Delta_R_cam,t * (R_rp,t+1)^T is exactly the isolated heading step by construction, and R_yaw is integrated open-loop. This is not a fitted target or a prediction that assumes its own conclusion; it is a deterministic geometric transformation. Human world orientation in Eq. (4) is then a product of this integrated camera heading, roll-pitch, and camera-space human orientation, all supervised with external ground truth such as camera extrinsics, gyro data, and SMPL annotations. The camera trajectory integration in Eqs. (5)-(6) predicts velocities and integrates them, using camera motion as contextual features and as an auxiliary supervised task, not as the direct source of the human trajectory. The only notable self-citation is GVHMR (Shen et al. 2024), which shares authors with the present paper, but GVHMR is used as an empirical baseline on external benchmarks, not as the justification for WATCH's derivation; the decomposition is derived in-paper from rotation algebra. Therefore no load-bearing step reduces to its own input. Concerns about DPVO angular-velocity drift or degenerate near-vertical camera cases are robustness and sensitivity risks, not circularity.
Assumptions & free parameters
free parameters (4)
- Trajectory loss weights lambda_h and lambda_cam =
1 and 1
- Depth threshold for loss masking =
0.3 m
- Static contact velocity threshold =
0.15 m/s
- Degenerate heading epsilon =
1e-6
assumptions (5)
- standard math Every rotation R can be decomposed as R = Ryaw Rrp, with Ryaw a rotation about the world Y-axis and Rrp the roll-pitch component.
- domain assumption The world Y-axis is aligned with gravity and can be recovered from camera roll and pitch.
- domain assumption The camera body-frame angular velocity from DPVO/SLAM or GT gyro is accurate enough for recursive yaw integration.
- standard math Rotation composition of the recursive heading increments yields a valid orthonormal rotation.
- domain assumption The camera-space human orientation Rh,c estimated by the network is a valid rotation from the body to the camera frame, so world orientation is R_cam Rh,c.
Cite this review
Pith. "Pith review of WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human." pith.science (2026). https://pith.science/paper/UV7GBRJM
@misc{pith2026250904600,
author = {Pith},
title = {Pith review of: WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human},
year = {2026},
howpublished = {\url{https://pith.science/paper/UV7GBRJM}},
note = {Machine review of arXiv:2509.04600}
}
read the original abstract
Global human motion reconstruction from in-the-wild monocular videos is increasingly demanded across VR, graphics, and robotics applications, yet requires accurate mapping of human poses from camera to world coordinates-a task challenged by depth ambiguity, motion ambiguity, and the entanglement between camera and human movements. While human-motion-centric approaches excel in preserving motion details and physical plausibility, they suffer from two critical limitations: insufficient exploitation of camera orientation information and ineffective integration of camera translation cues. We present WATCH (World-aware Allied Trajectory and pose reconstruction for Camera and Human), a unified framework addressing both challenges. Our approach introduces an analytical heading angle decomposition technique that offers superior efficiency and extensibility compared to existing geometric methods. Additionally, we design a camera trajectory integration mechanism inspired by world models, providing an effective pathway for leveraging camera translation information beyond naive hard-decoding approaches. Through experiments on in-the-wild benchmarks, WATCH achieves state-of-the-art performance in end-to-end trajectory reconstruction. Our work demonstrates the effectiveness of jointly modeling camera-human motion relationships and offers new insights for addressing the long-standing challenge of camera translation integration in global human motion reconstruction. The code will be available publicly.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
P.; Li, J.; Vetrivel, K.; Agarwal, R.; Wu, J.; Gopinath, D.; Clegg, A
Ara \'u jo, J. P.; Li, J.; Vetrivel, K.; Agarwal, R.; Wu, J.; Gopinath, D.; Clegg, A. W.; and Liu, K. 2023. Circle: Capture in rich contextual environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 21211--21221
work page 2023
-
[2]
J.; Patel, P.; Tesch, J.; and Yang, J
Black, M. J.; Patel, P.; Tesch, J.; and Yang, J. 2023. BEDLAM: A Synthetic Dataset of Bodies Exhibiting Detailed Lifelike Animated Motion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 8726--8737
work page 2023
-
[3]
Bogo, F.; Kanazawa, A.; Lassner, C.; Gehler, P.; Romero, J.; and Black, M. J. 2016. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In Computer Vision--ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part V 14, 561--578. Springer
2016
-
[4]
Choi, H.; Moon, G.; Chang, J. Y.; and Lee, K. M. 2021. Beyond static features for temporally consistent 3d human pose and shape from a video. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 1964--1973
work page 2021
-
[5]
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929
arXiv 2020
-
[6]
TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation
Dwivedi, S. K.; Sun, Y.; Patel, P.; Feng, Y.; and Black, M. J. 2024. TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation. arXiv:2404.16752
work page Pith review arXiv 2024
-
[7]
Fu, Z.; Zhao, Q.; Wu, Q.; Wetzstein, G.; and Finn, C. 2024. Humanplus: Humanoid shadowing and imitation from humans. arXiv preprint arXiv:2406.10454
arXiv 2024
-
[8]
Goel, S.; Pavlakos, G.; Rajasegaran, J.; Kanazawa, A.; and Malik, J. 2023. Humans in 4D: Reconstructing and Tracking Humans with Transformers. arXiv preprint arXiv:2305.20091
arXiv 2023
Show all 58 references
-
[9]
F.; Choi, C.; Schaefer, S.; and Leutenegger, S
Henning, D. F.; Choi, C.; Schaefer, S.; and Leutenegger, S. 2023. BodySLAM++: Fast and Tightly-Coupled Visual-Inertial Camera and Human Motion Tracking. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 3781--3788. IEEE
2023
-
[10]
F.; Laidlow, T.; and Leutenegger, S
Henning, D. F.; Laidlow, T.; and Leutenegger, S. 2022. BodySLAM: joint camera localisation, mapping, and human motion tracking. In European Conference on Computer Vision, 656--673. Springer
2022
-
[11]
P.; Yi, H.; H \"o schle, M.; Safroshkin, M.; Alexiadis, T.; Polikovsky, S.; Scharstein, D.; and Black, M
Huang, C.-H. P.; Yi, H.; H \"o schle, M.; Safroshkin, M.; Alexiadis, T.; Polikovsky, S.; Scharstein, D.; and Black, M. J. 2022. Capturing and Inferring Dense Full-Body Human-Scene Contact. In Proceedings IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 13274--13285
2022
-
[12]
Ionescu, C.; Papava, D.; Olaru, V.; and Sminchisescu, C. 2013. Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE transactions on pattern analysis and machine intelligence, 36(7): 1325--1339
2013
-
[13]
J.; Jacobs, D
Kanazawa, A.; Black, M. J.; Jacobs, D. W.; and Malik, J. 2018. End-to-end recovery of human shape and pose. In Proceedings of the IEEE conference on computer vision and pattern recognition, 7122--7131
2018
-
[14]
Y.; Felsen, P.; and Malik, J
Kanazawa, A.; Zhang, J. Y.; Felsen, P.; and Malik, J. 2019. Learning 3d human dynamics from video. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 5614--5623
2019
-
[15]
J.; and Hilliges, O
Kaufmann, M.; Song, J.; Guo, C.; Shen, K.; Jiang, T.; Tang, C.; Z \'a rate, J. J.; and Hilliges, O. 2023. EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 14632--14643
2023
-
[16]
P.; Hilliges, O.; and Black, M
Kocabas, M.; Huang, C.-H. P.; Hilliges, O.; and Black, M. J. 2021. PARE: Part attention regressor for 3D human body estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 11127--11137
2021
-
[17]
J.; Hilliges, O.; Kautz, J.; and Iqbal, U
Kocabas, M.; Yuan, Y.; Molchanov, P.; Guo, Y.; Black, M. J.; Hilliges, O.; Kautz, J.; and Iqbal, U. 2023. PACE: Human and Camera Motion Estimation from in-the-wild Videos. arXiv preprint arXiv:2310.13768
2023 arXiv
-
[18]
Kocabas, M.; et al. 2020. Vibe: Video inference for human body pose and shape estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 5253--5263
2020
-
[19]
J.; and Daniilidis, K
Kolotouros, N.; Pavlakos, G.; Black, M. J.; and Daniilidis, K. 2019. Learning to reconstruct 3D human pose and shape via model-fitting in the loop. In Proceedings of the IEEE/CVF international conference on computer vision, 2252--2261
2019
-
[20]
Li, J.; Bian, S.; Xu, C.; Liu, G.; Yu, G.; and Lu, C. 2022 a . D &D: Learning Human Dynamics from Dynamic Camera. In European Conference on Computer Vision, 479--496. Springer
2022
-
[21]
Li, J.; Cao, J.; Zhang, H.; Rempe, D.; Kautz, J.; Iqbal, U.; and Yuan, Y. 2025. GENMO: A GENeralist Model for Human MOtion. arXiv preprint arXiv:2505.01425
2025 arXiv
-
[22]
Li, J.; Xu, C.; Chen, Z.; Bian, S.; Yang, L.; and Lu, C. 2021. Hybrik: A hybrid analytical-neural inverse kinematics solution for 3d human pose and shape estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 3383--3393
2021
-
[23]
Li, Z.; Liu, J.; Zhang, Z.; Xu, S.; and Yan, Y. 2022 b . Cliff: Carrying location information in full frames into human pose and shape estimation. In Computer Vision--ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23--27, 2022, Proceedings, Part V, 590--606. Springer
2022
-
[24]
Liu, K.; Fu, Y.; Yuan, W.; Lin, J.; Li, P.; Gu, X.; Qiu, L.; Wang, H.; Dong, Z.; and Han, X. 2025 a . Motions as Queries: One-Stage Multi-Person Holistic Human Motion Capture. In Proceedings of the Computer Vision and Pattern Recognition Conference, 17529--17539
2025
-
[25]
Liu, Z.; Lin, J.; Wu, W.; and Zhou, B. 2025 b . Joint optimization for 4d human-scene reconstruction in the wild. arXiv preprint arXiv:2501.02158
2025 arXiv
-
[26]
Loper, M.; Mahmood, N.; Romero, J.; Pons-Moll, G.; and Black, M. J. 2015. SMPL: A skinned multi-person linear model. ACM transactions on graphics (TOG), 34(6): 1--16
2015
-
[27]
A.; and Kitani, K
Luo, Z.; Golestaneh, S. A.; and Kitani, K. M. 2020. 3d human motion estimation via motion compression and refinement. In Proceedings of the Asian Conference on Computer Vision
2020
-
[28]
F.; Pons-Moll, G.; and Black, M
Mahmood, N.; Ghorbani, N.; Troje, N. F.; Pons-Moll, G.; and Black, M. J. 2019. AMASS: Archive of motion capture as surface shapes. In Proceedings of the IEEE/CVF international conference on computer vision, 5442--5451
2019
-
[29]
Mur-Artal, R.; and Tard \'o s, J. D. 2017. Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras. IEEE transactions on robotics, 33(5): 1255--1262
2017
-
[30]
Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. 2023. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193
2023 arXiv
-
[31]
Patel, P.; and Black, M. J. 2025. CameraHMR: Aligning People with Perspective. In International Conference on 3D Vision (3DV)
2025
-
[32]
Patel, P.; et al. 2025. CameraHMR: Aligning People with Perspective. arXiv preprint arXiv:2411.08128
2025 arXiv
-
[33]
A.; Tzionas, D.; and Black, M
Pavlakos, G.; Choutas, V.; Ghorbani, N.; Bolkart, T.; Osman, A. A.; Tzionas, D.; and Black, M. J. 2019. Expressive body capture: 3d hands, face, and body from a single image. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 10975--10985
2019
-
[34]
Shen, X.; Yang, Z.; Wang, X.; Ma, J.; Zhou, C.; and Yang, Y. 2023. Global-to-Local Modeling for Video-based 3D Human Pose and Shape Estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 8887--8896
2023
-
[35]
Shen, Z.; Pi, H.; Xia, Y.; Cen, Z.; Peng, S.; Hu, Z.; Bao, H.; Hu, R.; and Zhou, X. 2024. World-Grounded Human Motion Recovery via Gravity-View Coordinates. In SIGGRAPH Asia 2024 Conference Papers, 1–11. ACM
2024
-
[36]
Shimada, S.; Golyanik, V.; Xu, W.; and Theobalt, C. 2020. Physcap: Physically plausible monocular 3d motion capture in real time. ACM Transactions on Graphics (ToG), 39(6): 1--16
2020
-
[37]
Shin, S.; Kim, J.; Halilaj, E.; and Black, M. J. 2023. WHAM: Reconstructing World-grounded Humans with Accurate 3D Motion. arXiv preprint arXiv:2312.07531
2023 arXiv
-
[38]
Shrestha, A.; Liu, P.; Ros, G.; Yuan, K.; and Fern, A. 2024. Generating physically realistic and directable human motions from multi-modal inputs. In European Conference on Computer Vision, 1--17. Springer
2024
-
[39]
Sun, Y.; Bao, Q.; Liu, W.; Mei, T.; and Black, M. J. 2023. TRACE: 5D temporal regression of avatars with dynamic cameras in 3D environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 8856--8866
2023
-
[40]
Teed, Z.; and Deng, J. 2021. Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras. Advances in neural information processing systems, 34: 16558--16569
2021
-
[41]
Teed, Z.; et al. 2023. Deep Patch Visual Odometry. arXiv:2208.04726
2023 arXiv
-
[42]
J.; Rosenhahn, B.; and Pons-Moll, G
Von Marcard, T.; Henschel, R.; Black, M. J.; Rosenhahn, B.; and Pons-Moll, G. 2018. Recovering accurate 3d human pose in the wild using imus and a moving camera. In Proceedings of the European conference on computer vision (ECCV), 601--617
2018
-
[43]
Wan, Z.; Li, Z.; Tian, M.; Liu, J.; Yi, S.; and Li, H. 2021. Encoder-decoder with Multi-level Attention for 3D Human Shape and Pose Estimation. In The IEEE International Conference on Computer Vision (ICCV)
2021
-
[44]
Wang, J.; Chen, M.; Karaev, N.; Vedaldi, A.; Rupprecht, C.; and Novotny, D. 2025 a . VGGT: Visual Geometry Grounded Transformer. arXiv:2503.11651
2025 arXiv
-
[45]
Wang, S.; Leroy, V.; Cabon, Y.; Chidlovskii, B.; and Revaud, J. 2024 a . DUSt3R: Geometric 3D Vision Made Easy. arXiv:2312.14132
2024 arXiv
-
[46]
Wang, Y.; and Daniilidis, K. 2023. Refit: Recurrent fitting network for 3d human recovery. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 14644--14654
2023
-
[47]
J.; and Kocabas, M
Wang, Y.; Sun, Y.; Patel, P.; Daniilidis, K.; Black, M. J.; and Kocabas, M. 2025 b . PromptHMR: Promptable Human Mesh Recovery. arXiv:2504.06397
2025 arXiv
-
[48]
Wang, Y.; Wang, Z.; Liu, L.; and Daniilidis, K. 2024 b . TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos. arXiv:2403.17346
2024 arXiv
-
[49]
Wei, W.-L.; Lin, J.-C.; Liu, T.-L.; and Liao, H.-Y. M. 2022. Capturing humans in motion: Temporal-attentive 3D human pose and shape estimation from monocular video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13211--13220
2022
-
[50]
Xie, K.; Wang, T.; Iqbal, U.; Guo, Y.; Fidler, S.; and Shkurti, F. 2021. Physics-based human motion estimation and synthesis from videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision (CVPR)
2021
-
[51]
Ye, V.; Pavlakos, G.; Malik, J.; and Kanazawa, A. 2023 a . Decoupling human and camera motion from videos in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 21222--21232
2023
-
[52]
Ye, V.; Pavlakos, G.; Malik, J.; and Kanazawa, A. 2023 b . Decoupling Human and Camera Motion from Videos in the Wild. arXiv:2302.12827
2023 arXiv
-
[53]
Yin, W.; Cai, Z.; Wang, R.; Wang, F.; Wei, C.; Mei, H.; Xiao, W.; Yang, Z.; Sun, Q.; Yamashita, A.; Liu, Z.; and Yang, L. 2024. WHAC: World-grounded Humans and Cameras. arXiv:2403.12959
2024 arXiv
-
[54]
Yuan, Y.; Iqbal, U.; Molchanov, P.; Kitani, K.; and Kautz, J. 2022. GLAMR: Global occlusion-aware human mesh recovery with dynamic cameras. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 11038--11049
2022
-
[55]
Zhang, H.; Tian, Y.; Zhou, X.; Ouyang, W.; Liu, Y.; Wang, L.; and Sun, Z. 2021. Pymaf: 3d human pose and shape regression with pyramidal mesh alignment feedback loop. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 11446--11456
2021
-
[56]
O.; Cui, Z.; and Ji, Q
Zhang, Y.; Kephart, J. O.; Cui, Z.; and Ji, Q. 2024. Physpt: Physics-aware pretrained transformer for estimating human dynamics from monocular videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2024
-
[57]
F.; Wang, H.; and Zhang, L
Zhang, Y.; Wu, G.; Chen, L.-H.; Zhao, Z.; Lin, J.; Jiang, X.; Wu, J.; Li, Z.; Yang, H. F.; Wang, H.; and Zhang, L. 2025. HumanMM: Global Human Motion Recovery from Multi-shot Videos. arXiv:2503.07597
2025 arXiv
-
[58]
Y.; Raj, B.; Xu, M.; Yang, J.; and Huang, C.-H
Zhao, Y.; Wang, T. Y.; Raj, B.; Xu, M.; Yang, J.; and Huang, C.-H. P. 2024. Synergistic Global-space Camera and Human Reconstruction from Videos. arXiv:2405.14855
2024 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.