REVIEW 2 major objections 6 minor 2 cited by
Reconstructing People, Places, and Cameras
T0 review · 2 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read HSfM claims that jointly optimizing human meshes, scene point clouds, and cameras from sparse uncalibrated multi-view images makes all three more accurate, with the statistical scale of human bodies supplying metric units to the whole…
desk verdict Solid integration paper with real gains, but the metric-scale claim needs a direct scale-error test before I'd trust the headline numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a joint objective $L_{\text{Humans}} + \lambda L_{\text{Places}}$ defined over human body parameters, depth maps, and camera parameters, coupled by a projection model $x_{2D} = K(Rx_{3D} + \alpha t)$ in which the scalar $\alpha$ rescales camera translations while preserving their directions. The human term is bundle adjustment on 2D keypoint detections: predicted 3D joints are projected through each estimated camera and compared with detected keypoints, normalized by bounding-box height and weighted by confidence, plus a shape regularizer that keeps body shapes near the SMPL-X mean, SMPL-X being a parametric human body model. The scene term is a global alignment loss that compares each per-view pointmap, transformed into the world frame, against cross-view pointmap predictions from the scene reconstruction model. The initialization procedure computes $\alpha$ analytically: it reads camera rotations from the human body orientation, reads camera translations from human positions using a similar-triangle bone-length ratio, and then least-squares fits the scene-derived camera centers to the human-derived ones. This chain is what transmits the metric statistics of the human body model into the scene and cameras.
What would settle it
Take a multi-view sequence of two people wearing identical uniforms, run HSfM with an automated re-ID module instead of ground-truth identities, and compare W-MPJPE and RRA@15 to the same run with correct identities; the paper's own supplement reports automated re-ID at 51% on EgoHumans, so the claim predicts a sharp degradation in both metrics.
Extended reading notes
Core claim
The central discovery is a synergy: human reconstructions, scene point clouds, and cameras reinforce one another inside a single optimization. The paper's argument is that 2D human keypoints act as reliable correspondences for bundle adjustment, 3D human mesh predictions act as a statistically anchored 3D structure that fixes the otherwise arbitrary scale of the scene, and the scene structure in turn prevents cameras from drifting or overfitting to keypoint noise. HSfM implements this by first estimating a global scale $\alpha$ and human world positions $\gamma$ that align the human-centric and scene-centric camera estimates, then minimizing $L_{\text{Humans}} + \lambda L_{\text{Places}}$ over human body parameters, depth maps, camera intrinsics and extrinsics. The human bundle-adjustment loss $L_{\text{Humans}}$ re-projects predicted 3D joints through the estimated cameras and compares them to 2D keypoint detections, while $L_{\text{Places}}$ is the global alignment loss that merges per-view scene pointmaps into one world space. The paper's ablations show that removing either loss degrades both human and camera metrics, and that increasing the number of people in the optimization steadily improves camera pose, which is the evidence for the claimed feedback loop.
Load-bearing premise
The method assumes a human operator (or oracle) has already matched every person across all camera views, and that the body-size statistics learned by the mesh regressor are reliable enough to set the world's metric scale.
Editorial extensions
If this is right
- A sparse, uncalibrated camera setup—even two phones—can produce a metric-scale world reconstruction whenever the images contain people, removing the need for calibration targets or known camera positions.
- Multi-person scenes are not just harder cases; they make the reconstruction better, since each additional person adds more keypoint correspondences for bundle adjustment and tightens camera pose and scale estimates.
- Human-only bundle adjustment is insufficient: without the scene alignment loss, camera metrics degrade sharply, so any production system should keep humans and scene in one optimization rather than treating human pose as a post-process.
- The scale source is modular: swapping the human mesh regressor for one trained on non-standard body sizes (as the paper notes) changes the scale anchor without changing the optimization geometry.
- The same joint objective extends to video or dynamic scenes if the scene model is replaced by a motion-aware version, since the optimization itself does not depend on static camera constraints.
Reading between the lines
- The method's current ceiling is person re-identification: the paper's own supplement reports only 51% accuracy for an automated re-ID module on EgoHumans, so the joint optimization's gains would largely evaporate in an end-to-end deployment until re-ID is solved.
- The human-as-scale-anchor idea is transferable: the same $\alpha$ mechanism could calibrate monocular metric depth or single-view reconstruction when a person appears in the frame, which the paper does not explore.
- A testable extension is to make the scale anchor robust to body-size bias by fusing several humans' predicted bone lengths (or known object sizes) in the least-squares fit, which would reduce the risk that one atypical body skews the whole scene scale.
- Because the optimization already runs on sparse views and produces metric output, it is a plausible front-end for egocentric or social-interaction capture where the goal is not just pose but spatial relationships between people and their environment, which the paper demonstrates only qualitatively.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes HSfM, an optimization-based method that jointly reconstructs SMPL-X human meshes, per-view scene pointmaps, and camera poses/intrinsics from sparse uncalibrated multi-view images, assuming known person correspondences across views. Initialization combines HMR2 human meshes and DUSt3R scene pointmaps; a global scale alpha is fit by aligning human-derived camera translations with DUSt3R camera positions, and a joint loss then combines human keypoint bundle adjustment, body-shape regularization, and a DUSt3R-style global alignment loss. Experiments on EgoHumans and EgoExo4D report consistent improvements over UnCaliPose, DUSt3R, and MASt3R in human and camera metrics, and ablations show that removing either the human loss or the scene loss degrades part of the results.
Significance. If the metric-scale contribution is properly validated, this is a useful and timely contribution: the paper is among the first to fuse dense scene pointmaps, multi-person human meshes, and cameras into one joint optimization, and it releases code. The ablations in Table 3 give concrete evidence that the human keypoint loss and the scene alignment loss help different parts of the reconstruction, and the improvements in scale-invariant camera metrics (RRA, s-CCA) are consistent across two benchmarks. The central novelty, however, is the recovery of approximate metric scale from a human statistical model, and that claim is currently not directly measured: the headline world metrics are computed after SE(3) camera alignment, and no experiment isolates global scale error. The paper also clearly states its re-identification assumption and documents in the supplementary material that automated re-ID is far from solved, which is an important scope limitation but not a hidden flaw.
major comments (2)
- [§4.1, Eq. (9), §5 Table 2] The metric-scale claim is not validated. The only scale anchor in the optimization is the shape regularizer L_beta = ||beta||^2 (Eq. 9), which pulls every person toward the SMPL-X mean body, and the initial scale alpha is fit by least squares between human-derived camera positions and DUSt3R camera positions (Sec. 4.1). A systematic bias in HMR2's body-size or bone-length predictions would therefore propagate directly into alpha and all world metrics, yet no experiment measures scale error alone (e.g., predicted vs. ground-truth scene-scale ratio) nor the sensitivity to body-size prior. The current absolute metrics, W-MPJPE and TE, are computed after SE(3) camera alignment (Suppl. S.4.1), and the scale-invariant metrics (s-TE, s-CCA, GA-MPJPE) are by construction insensitive to the claimed contribution. I request a direct per-scene scale-error evaluation and a sensitivity analysis (e.g., perturbing the body-scale prior by ±5–10%) so that the paper's 'approximate metric scale' claim is bounded. Note also that the scene loss in Eq. (10) leaves sigma unregularized, making L_Places scale-invariant; the human terms are the only scale anchor, which further motivates this experiment.
- [§5 Evaluation Metrics, Suppl. S.4.1] The headline W-MPJPE and TE are reported after SE(3) alignment of the predicted cameras to the ground-truth cameras, so the absolute world position is not evaluated. The abstract's phrasing ('human localization accuracy within the world coordinate frame') is stronger than what the metric supports, since the global translation and rotation are removed by the alignment. Please report an unaligned world-coordinate error, or at least report the distribution of the residual alignment transform (e.g., the global translation component) so readers can see how much of the error is absolute localization rather than relative reconstruction.
minor comments (6)
- [Abstract / Introduction / Table 2] The numbers in the abstract and introduction do not match Table 2: EgoExo4D W-MPJPE is 3.59m (baseline) and 0.50m (HSfM) in Table 2, but the abstract says 2.9m to 0.56m, and the introduction cites 3.51m for EgoHumans, which matches the table. Please harmonize these figures.
- [References] References [55] and [56] are the same paper (Xu and Kitani, BMVC 2022); one entry should be removed or the citations disambiguated.
- [Eq. (5)] The notation around Eq. (5) is under-specified: the variables \tilde{\gamma}_c and \hat{T}_c are not fully defined in terms of their coordinate frames, and the equation's relationship to the camera center is not immediately clear. Please clarify the derivation.
- [Suppl. S.3, Table S.1] The footnote for the 2-camera rows states that s-TE is omitted because 'the predictions become identical to the ground truth camera translations after scale alignment'; this is confusing and should be explained in more detail, since scale alignment does not normally make predictions identical to ground truth unless the method is degenerate.
- [Tables 2, 3, S.1] No variance or per-sequence breakdown is reported for any metric. Given that the paper's own conclusion notes sensitivity to hyperparameters, reporting error bars or per-sequence results would strengthen confidence in the reported improvements.
- [Sec. 1 / Suppl. S.2] The known-re-ID assumption is clearly stated in the main text, but the practical severity (automated re-ID reaches only 51.22% accuracy on EgoHumans, per Suppl. S.2) is only discussed in the supplementary material. Please mention this limitation in the main text so readers understand the manual-annotation requirement of the method.
Circularity Check
No significant circularity: HSfM's metric scale is an external input (HMR2/SMPL-X), and its alpha fit and joint optimization are standard data fitting against independent benchmarks.
full rationale
The paper's derivation chain is not circular. The metric scale enters through pretrained HMR2/SMPL-X human mesh predictions, which are external, independently trained models; the scene scale alpha is then fitted by least squares to align DUSt3R camera positions with human-derived camera positions (Sec. 4.1, Eq. 5), which is a calibration/fitting step rather than a prediction of the same quantity being evaluated. The joint optimization (Eqs. 6-10) minimizes reprojection error against 2D keypoints and a global pointmap alignment loss, and no ground-truth metric from the evaluation benchmarks is used as an optimization target. The shape regularizer L_beta = ||beta||^2 (Eq. 9) is a prior toward the average SMPL-X body, which is an assumption about body scale and not a circularity: it does not use the target errors or ground-truth data. Evaluation metrics (W-MPJPE, TE, RRA, etc.) are measured against external ground-truth annotations after rigid or Sim(3) alignment. Self-citations to Ye et al. [61] and HMR2 [16] provide pretrained, externally validated components rather than an unverified uniqueness theorem or an ansatz that already contains the claimed result. The stated assumption of known re-identification is a limitation (Suppl. S.2) and does not reduce the central claim to its inputs. Overall, the paper is an optimization pipeline with external data-driven initialization; no load-bearing step reduces by construction to its own inputs.
Assumptions & free parameters
free parameters (6)
- alpha (world scale) =
per scene, initialized via least squares in Sec. 4.1, refined by optimization
- gamma_h (global human translation) =
per person
- beta_h (body shape) =
per person
- camera poses and focal length (R_c, t_c, f) =
per camera
- lambda (scene loss weight) =
not reported
- optimization steps and learning rate =
min 500 steps, LR 0.015
assumptions (6)
- domain assumption Known person re-identification across all views
- domain assumption HMR2 body meshes carry metric scale from training statistics
- domain assumption DUSt3R pointmaps and cameras are a reliable initialization
- domain assumption Consistent human orientation across views
- domain assumption Simplified pinhole camera model with centered principal point and no distortion
- domain assumption The staged optimization avoids poor local minima
Cite this review
Pith. "Pith review of Reconstructing People, Places, and Cameras." pith.science (2026). https://pith.science/paper/7T4RO47C
@misc{pith2026241217806,
author = {Pith},
title = {Pith review of: Reconstructing People, Places, and Cameras},
year = {2026},
howpublished = {\url{https://pith.science/paper/7T4RO47C}},
note = {Machine review of arXiv:2412.17806}
}
read the original abstract
We present "Humans and Structure from Motion" (HSfM), a method for jointly reconstructing multiple human meshes, scene point clouds, and camera parameters in a metric world coordinate system from a sparse set of uncalibrated multi-view images featuring people. Our approach combines data-driven scene reconstruction with the traditional Structure-from-Motion (SfM) framework to achieve more accurate scene reconstruction and camera estimation, while simultaneously recovering human meshes. In contrast to existing scene reconstruction and SfM methods that lack metric scale information, our method estimates approximate metric scale by leveraging a human statistical model. Furthermore, it reconstructs multiple human meshes within the same world coordinate system alongside the scene point cloud, effectively capturing spatial relationships among individuals and their positions in the environment. We initialize the reconstruction of humans, scenes, and cameras using robust foundational models and jointly optimize these elements. This joint optimization synergistically improves the accuracy of each component. We compare our method to existing approaches on two challenging benchmarks, EgoHumans and EgoExo4D, demonstrating significant improvements in human localization accuracy within the world coordinate frame (reducing error from 3.51m to 1.04m in EgoHumans and from 2.9m to 0.56m in EgoExo4D). Notably, our results show that incorporating human data into the SfM pipeline improves camera pose estimation (e.g., increasing RRA@15 by 20.3% on EgoHumans). Additionally, qualitative results show that our approach improves overall scene reconstruction quality. Our code is available at: https://github.com/hongsukchoi/HSfM_RELEASE
Figures
Forward citations
Cited by 2 Pith papers
-
Humans as a Calibration Pattern: Dynamic 3D Scene Reconstruction from Unsynchronized and Uncalibrated Videos
Dynamic 3D scenes can be reconstructed from unsynchronized, uncalibrated multi-view videos by first aligning estimated human motion across views and then refining the alignment during neural field training.
-
UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting
UAV4D reconstructs 4D scenes from monocular drone video by fitting a single global scale to align human meshes with the background mesh, then renders with separate Gaussian splats.
Reference graph
Works this paper leans on
-
[1]
Sameer Agarwal, Yasutaka Furukawa, Noah Snavely, Ian Si- mon, Brian Curless, Steven M Seitz, and Richard Szeliski. Building rome in a day. ACM Communications, 2011. 3
work page 2011
-
[2]
Multi-hmr: Multi-person whole-body hu- man mesh recovery in a single shot
Fabien Baradel, Matthieu Armando, Salma Galaaoui, Ro- main Br ´egier, Philippe Weinzaepfel, Gr ´egory Rogez, and Thomas Lucas. Multi-hmr: Multi-person whole-body hu- man mesh recovery in a single shot. arXiv preprint arXiv:2402.14654, 2024. 3, 7
arXiv 2024
-
[3]
Key.net: Keypoint detection by hand- crafted and learned cnn filters
Axel Barroso-Laguna, Edgar Riba, Daniel Ponsa, and Krys- tian Mikolajczyk. Key.net: Keypoint detection by hand- crafted and learned cnn filters. In International Conference on Computer Vision (ICCV), 2019. 3
work page 2019
-
[4]
Surf: Speeded up robust features
Herbert Bay, Tinne Tuytelaars, and Luc Van Gool. Surf: Speeded up robust features. In European Conference on Computer Vision (ECCV), 2006. 3
work page 2006
-
[5]
3d model acquisition from extended image sequences
Paul Beardsley, Phil Torr, and Andrew Zisserman. 3d model acquisition from extended image sequences. In European Conference on Computer Vision (ECCV), 1996. 2
work page 1996
-
[6]
Multi-person 3D Pose Estimation in Crowded Scenes Based on Multi-View Geometry
Chao Chen, Georgios Pavlakos, Shihao Zou, Tony Tung, and Katerina Fragkiadaki. Multi-person 3d pose estimation in crowded scenes based on multi-view geometry. arXiv preprint arXiv:2007.10986, 2020. 3
work page Pith review arXiv 2007
-
[7]
Aspanformer: Detector-free image matching with adaptive span transformer
Hongkai Chen, Zixin Luo, Lei Zhou, Yurun Tian, Ming- min Zhen, Tian Fang, David McKinnon, Yanghai Tsin, and Long Quan. Aspanformer: Detector-free image matching with adaptive span transformer. In European Conference on Computer Vision (ECCV), 2022. 3
work page 2022
-
[8]
Pose2mesh: Graph convolutional network for 3d human pose and mesh recovery from a 2d human pose
Hongsuk Choi, Gyeongsik Moon, and Kyoung Mu Lee. Pose2mesh: Graph convolutional network for 3d human pose and mesh recovery from a 2d human pose. InEuropean Con- ference on Computer Vision (ECCV), 2020. 2
work page 2020
Show all 72 references
-
[9]
Sfm with mrfs: Discrete-continuous optimiza- tion for large-scale structure from motion
David Crandall, Andrew Owens, Noah Snavely, and Daniel Huttenlocher. Sfm with mrfs: Discrete-continuous optimiza- tion for large-scale structure from motion. IEEE Transac- tions on Pattern Analysis and Machine Intelligence (PAMI),
-
[10]
Hsfm: Hybrid structure-from-motion
Hainan Cui, Xiang Gao, Shuhan Shen, and Zhanyi Hu. Hsfm: Hybrid structure-from-motion. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), 2017. 3
2017
-
[11]
Superpoint: Self-supervised interest point detection and description
Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabi- novich. Superpoint: Self-supervised interest point detection and description. In CVPR Workshops, 2018. 3
2018
-
[12]
Fast and ro- bust multi-person 3d pose estimation from multiple views
Junting Dong, Hujun Bao, and Xiaowei Zhou. Fast and ro- bust multi-person 3d pose estimation from multiple views. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 7792–7801,
-
[13]
Mast3r- sfm: A fully-integrated solution for unconstrained structure- from-motion
Bardienus Duisterhof, Lojze Zust, Philippe Weinzaepfel, Vincent Leroy, Yohann Cabon, and Jerome Revaud. Mast3r- sfm: A fully-integrated solution for unconstrained structure- from-motion. arXiv preprint arXiv, 2024. 3, 7
2024
-
[14]
Sai Kumar Dwivedi, Yu Sun, Priyanka Patel, Yao Feng, and Michael J. Black. TokenHMR: Advancing human mesh re- covery with a tokenized pose representation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 7
2024
-
[15]
Automatic camera recovery for closed or open image sequences
Andrew W Fitzgibbon and Andrew Zisserman. Automatic camera recovery for closed or open image sequences. In Eu- ropean Conference on Computer Vision (ECCV), 1998. 2
1998
-
[16]
Humans in 4D: Reconstructing and tracking humans with transformers
Shubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa*, and Jitendra Malik*. Humans in 4D: Reconstructing and tracking humans with transformers. In International Conference on Computer Vision (ICCV), 2023. 2, 3, 4, 7, 8, 1
2023
-
[17]
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In Proceedings...
2024
-
[18]
Unsupervised multi-person 3d human pose estimation from 2d poses alone
Peter Hardy and Hansung Kim. Unsupervised multi-person 3d human pose estimation from 2d poses alone. arXiv preprint arXiv:2309.14865, 2023. 3
2023 arXiv
-
[19]
Multiple View Ge- ometry in Computer Vision
Richard Hartley and Andrew Zisserman. Multiple View Ge- ometry in Computer Vision. 2004. 2
2004
-
[20]
Huang, Hongwei Yi, Markus H ¨oschle, Matvey Safroshkin, Tsvetelina Alexiƒadis, Senya Polikovsky, Daniel Scharstein, and Michael J
Chun-Hao P. Huang, Hongwei Yi, Markus H ¨oschle, Matvey Safroshkin, Tsvetelina Alexiƒadis, Senya Polikovsky, Daniel Scharstein, and Michael J. Black. Capturing and inferring dense full-body human-scene contact. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recogni...
2022
-
[21]
Cotr: Correspondence transformer for matching across images
Wei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasac- chi, and Kwang Moo Yi. Cotr: Correspondence transformer for matching across images. In International Conference on Computer Vision (ICCV), 2021. 3
2021
-
[22]
A gener- alizable approach for multi-view 3d human pose regression
Amir Kadkhodamohammadi and Nicolas Padoy. A gener- alizable approach for multi-view 3d human pose regression. arXiv preprint arXiv:1804.10462, 2018. 3
2018 arXiv
-
[23]
Black, David W
Angjoo Kanazawa, Michael J. Black, David W. Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In Computer Vision and Pattern Recognition (CVPR), pages 7122–7131, 2018. 2, 3, 7
2018
-
[24]
Egohumans: An egocentric 3d multi-human benchmark
Rawal Khirodkar, Aayush Bansal, Lingni Ma, Richard New- combe, Minh V o, and Kris Kitani. Egohumans: An egocentric 3d multi-human benchmark. arXiv preprint arXiv:2305.16487, 2023. 2, 6, 1
2023 arXiv
-
[25]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015. 7
2015
-
[26]
Ground- ing image matching in 3d with mast3r
Vincent Leroy, Yohann Cabon, and Jerome Revaud. Ground- ing image matching in 3d with mast3r. In European Confer- ence on Computer Vision (ECCV), 2024. 2, 3, 7
2024
-
[27]
Relpose++: Recovering 6d poses from sparse-view observations
Amy Lin, Jason Y Zhang, Deva Ramanan, and Shubham Tul- siani. Relpose++: Recovering 6d poses from sparse-view observations. arXiv preprint arXiv:2305.04926, 2023. 6, 7, 4
2023 arXiv
-
[28]
Lightglue: Local feature matching at light speed
Philipp Lindenberger, Paul-Edouard Sarlin, and Marc Polle- feys. Lightglue: Local feature matching at light speed. In International Conference on Computer Vision (ICCV), 2023. 3
2023
-
[29]
4d human body capture from egocen- tric video via 3d scene grounding
Miao Liu, Dexin Yang, Yan Zhang, Zhaopeng Cui, James M Rehg, and Siyu Tang. 4d human body capture from egocen- tric video via 3d scene grounding. In 3DV, 2021. 7
2021
-
[30]
David G. Lowe. Distinctive image features from scale- invariant keypoints. International Journal of Computer Vi- sion (IJCV), 2004. 3
2004
-
[31]
Bag of tricks and a strong baseline for deep per- son re-identification
Hao Luo, Youzhi Gu, Xingyu Liao, Shenqi Lai, and Wei Jiang. Bag of tricks and a strong baseline for deep per- son re-identification. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition Work- shops, 2019. 1
2019
-
[32]
Yang, Shenlong Wang, Raquel Urtasun, and Antonio Torralba
Wei-Chiu Ma, Alexander J. Yang, Shenlong Wang, Raquel Urtasun, and Antonio Torralba. Virtual correspondence: Hu- mans as a cue for extreme-view geometry. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3
2022
-
[33]
Black, and Angjoo Kanazawa
Lea M ¨uller, Vickie Ye, Georgios Pavlakos, Michael J. Black, and Angjoo Kanazawa. Generative proxemics: A prior for 3D social interaction from images. 2024. 7
2024
-
[34]
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 1
2023 arXiv
-
[35]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. In Computer Vision and Pat- tern Recognition (CVPR), pages 10975–10985, 2019. 4, 7
2019
-
[36]
Human mesh recovery from multiple shots
Georgios Pavlakos, Jitendra Malik, and Angjoo Kanazawa. Human mesh recovery from multiple shots. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 1485–1495, 2022. 3
2022
-
[37]
Tracking people by predict- ing 3d appearance, location and pose
Jathushan Rajasegaran, Georgios Pavlakos, Angjoo Kanazawa, and Jitendra Malik. Tracking people by predict- ing 3d appearance, location and pose. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022. 1
2022
-
[38]
Sam 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junt- ing Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao- Yuan Wu, Ross Girshick, Piotr Doll´ar, and Christoph Feic...
2024 arXiv
-
[39]
R2d2: Reliable and re- peatable detector and descriptor
J ´erˆome Revaud, C ´esar Roberto de Souza, Martin Humen- berger, and Philippe Weinzaepfel. R2d2: Reliable and re- peatable detector and descriptor. In Advances in Neural In- formation Processing Systems (NeurIPS), 2019. 3
2019
-
[40]
Ethan Rublee, Vincent Rabaud, Kurt Konolige, and Gary R. Bradski. Orb: An efficient alternative to sift or surf. In In- ternational Conference on Computer Vision (ICCV) , 2011. 3
2011
-
[41]
Structure-from-motion revisited
Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-from-motion revisited. In Conference on Com- puter Vision and Pattern Recognition (CVPR), 2016. 2, 3
2016
-
[42]
Pixelwise view selection for un- structured multi-view stereo
Johannes Lutz Sch ¨onberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. Pixelwise view selection for un- structured multi-view stereo. In European Conference on Computer Vision (ECCV), 2016. 2, 3
2016
-
[43]
World-grounded human motion recovery via gravity-view coordinates
Zehong Shen, Huaijin Pi, Yan Xia, Zhi Cen, Sida Peng, Zechen Hu, Hujun Bao, Ruizhen Hu, and Xiaowei Zhou. World-grounded human motion recovery via gravity-view coordinates. In SIGGRAPH Asia 2024 Conference Papers , pages 1–11, 2024. 7
2024
-
[44]
Soyong Shin, Juyong Kim, Eni Halilaj, and Michael J. Black. WHAM: Reconstructing world-grounded humans with accu- rate 3D motion. In IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024. 1, 7
2024
-
[45]
Loftr: Detector-free local feature matching with transformers
Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. Loftr: Detector-free local feature matching with transformers. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 3
2021
-
[46]
Putting people in their place: Monocular regression of 3d people in depth
Yu Sun, Wu Liu, Qian Bao, Yili Fu, Tao Mei, and Michael J Black. Putting people in their place: Monocular regression of 3d people in depth. In Computer Vision and Pattern Recog- nition (CVPR), pages 13243–13252, 2022. 4, 7
2022
-
[47]
Yu Sun, Qian Bao, Wu Liu, Tao Mei, and Michael J. Black. TRACE: 5D Temporal Regression of Avatars with Dynamic Cameras in 3D Environments. In IEEE/CVF Conf. on Com- puter Vision and Pattern Recognition (CVPR), 2023. 7
2023
-
[48]
Shape and motion from image streams under orthography: a factorization method
Carlo Tomasi and Takeo Kanade. Shape and motion from image streams under orthography: a factorization method. International journal of computer vision , 9:137–154, 1992. 2
1992
-
[49]
McLauchlan, Richard I
Bill Triggs, Paul F. McLauchlan, Richard I. Hartley, and An- drew W. Fitzgibbon. Bundle adjustment—a modern synthe- sis. In Vision Algorithms: Theory and Practice, pages 298–
-
[50]
Glu- net: Global-local universal network for dense flow and corre- spondences
Prune Truong, Martin Danelljan, and Radu Timofte. Glu- net: Global-local universal network for dense flow and corre- spondences. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 3
2020
-
[51]
Learning accurate dense correspondences and when to trust them
Prune Truong, Martin Danelljan, Luc Van Gool, and Radu Timofte. Learning accurate dense correspondences and when to trust them. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 3
2021
-
[52]
Posediffusion: Solving pose estimation via diffusion-aided bundle adjustment
Jianyuan Wang, Christian Rupprecht, and David Novotny. Posediffusion: Solving pose estimation via diffusion-aided bundle adjustment. In International Conference on Com- puter Vision (ICCV), pages 9773–9783, 2023. 6, 7, 4
2023
-
[53]
Dust3r: Geometric 3d vi- sion made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vi- sion made easy. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20697–20709, 2024. 2, 3, 4, 5, 7, 8, 1
2024
-
[54]
Tram: Global trajectory and motion of 3d humans from in- the-wild videos
Yufu Wang, Ziyun Wang, Lingjie Liu, and Kostas Daniilidis. Tram: Global trajectory and motion of 3d humans from in- the-wild videos. arXiv preprint arXiv:2403.17346, 2024. 7
2024 arXiv
-
[55]
Multi-view multi-person 3d pose es- timation with uncalibrated camera networks
Yan Xu and Kris Kitani. Multi-view multi-person 3d pose es- timation with uncalibrated camera networks. In British Ma- chine Vision Conference (BMVC), 2022. 2, 3, 7
2022
-
[56]
Multi-view multi-person 3d pose es- timation with uncalibrated camera networks
Yan Xu and Kris Kitani. Multi-view multi-person 3d pose es- timation with uncalibrated camera networks. In British Ma- chine Vision Conference (BMVC), 2022. 2, 3
2022
-
[57]
DenseRaC: Joint 3D pose and shape estimation by dense render-and- compare
Yuanlu Xu, Song-Chun Zhu, and Tony Tung. DenseRaC: Joint 3D pose and shape estimation by dense render-and- compare. In International Conference on Computer Vision (ICCV), 2019. 1
2019
-
[58]
Wide-baseline multi-camera calibration using person re- identification
Yan Xu, Yu-Jhe Li, Xinshuo Weng, and Kris Kitani. Wide-baseline multi-camera calibration using person re- identification. In Conference on Computer Vision and Pat- tern Recognition (CVPR), 2021. 3
2021
-
[59]
Vit- pose: Simple vision transformer baselines for human pose estimation
Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. Vit- pose: Simple vision transformer baselines for human pose estimation. Advances in Neural Information Processing Sys- tems, 35:38571–38584, 2022. 2, 3, 5, 8, 1
2022
-
[60]
Mvsnet: Depth inference for unstructured multi-view stereo
Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan. Mvsnet: Depth inference for unstructured multi-view stereo. In Proceedings of the European conference on computer vi- sion (ECCV), 2018. 2
2018
-
[61]
Decoupling human and camera motion from videos in the wild
Vickie Ye, Georgios Pavlakos, Jitendra Malik, and Angjoo Kanazawa. Decoupling human and camera motion from videos in the wild. In Computer Vision and Pattern Recogni- tion (CVPR), 2023. 3, 5, 1, 2, 7
2023
-
[62]
Lift: Learned invariant feature transform
Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, and Pascal Fua. Lift: Learned invariant feature transform. In European Conference on Computer Vision (ECCV), 2016. 3
2016
-
[63]
Multi- view human body reconstruction from uncalibrated cam- eras
Tao Yu, Zerong Zheng, Kaiwen Guo, and Yebin Liu. Multi- view human body reconstruction from uncalibrated cam- eras. In Advances in Neural Information Processing Systems (NeurIPS), 2022. 3
2022
-
[64]
Huang, Donglai Xiang, Yufeng Zhou, Mengcheng Xu, Jingwei Huang, Chenxi Jiang, Tzu- Mao Xu, Deva Ramanan, and Michael J
Yuming Yuan, Chun-Hao P. Huang, Donglai Xiang, Yufeng Zhou, Mengcheng Xu, Jingwei Huang, Chenxi Jiang, Tzu- Mao Xu, Deva Ramanan, and Michael J. Black. Hum- man: Multi-modal 4d human dataset for versatile sensing and modeling. In Proceedings of the IEEE/CVF Conference on Compu...
2022
-
[65]
Glamr: Global occlusion-aware human mesh recov- ery with dynamic cameras
Ye Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani, and Jan Kautz. Glamr: Global occlusion-aware human mesh recov- ery with dynamic cameras. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 11038–11049, 2022. 7
2022
-
[66]
Monst3r: A simple approach for estimating geometry in the presence of motion
Junyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jam- pani, Trevor Darrell, Forrester Cole, Deqing Sun, and Ming- Hsuan Yang. Monst3r: A simple approach for estimating geometry in the presence of motion. ICLR, 2025. 8
2025
-
[67]
Ego- body: Human body shape and motion of interacting people from head-mounted devices
Siwei Zhang, Qianli Ma, Yan Zhang, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, and Siyu Tang. Ego- body: Human body shape and motion of interacting people from head-mounted devices. In European Conference on Computer Vision (ECCV), pages 180–200, 2022. 7
2022
-
[68]
Metric from human: Zero-shot monoc- ular metric depth estimation via test-time adaptation
Yizhou Zhao, Hengwei Bian, Kaihua Chen, Pengliang Ji, Liao Qu, Shao-yu Lin, Weichen Yu, Haoran Li, Hao Chen, Jun Shen, et al. Metric from human: Zero-shot monoc- ular metric depth estimation via test-time adaptation. In The Thirty-eighth Annual Conference on Neural Information...
2024
-
[69]
Wang, Bhiksha Raj, Min Xu, Jimei Yang, and Chun-Hao P
Yizhou Zhao, Tuanfeng Y . Wang, Bhiksha Raj, Min Xu, Jimei Yang, and Chun-Hao P. Huang. Synergistic global- space camera and human reconstruction from videos. 2024. 1, 7
2024
-
[70]
Pmatch: Paired masked image modeling for dense geometric matching
Shengjie Zhu and Xiaoming Liu. Pmatch: Paired masked image modeling for dense geometric matching. In Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,
-
[71]
Reconstructing People, Places, and Cameras
Yuliang Zou, Jimei Yang, Duygu Ceylan, Jianming Zhang, Federico Perazzi, and Jia-Bin Huang. Reducing footskate in human motion reconstruction with ground contact con- straints. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2020. 2 Reconst...
2020
-
[372]
Springer, Berlin, Heidelberg, 2000. 3
2000
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.