REVIEW 4 major objections 3 minor 81 references
Utilizing Uncertainty in 2D Pose Detectors for Probabilistic 3D Human Mesh Recovery
T0 review · 4 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A normalizing flow for 3D human mesh recovery can be trained to match the per-joint uncertainty encoded in 2D pose-detector heatmaps, and penalizing hypotheses that fall outside person masks removes implausible samples without sacrificing…
desk verdict Central accuracy claim is clean and well-supported; the plausible-diversity claims lean on a shared ViTPose visibility proxy and a manually filtered EMDB subset — worth fixing, not fatal. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a conditional normalizing flow over 132 SMPL-X body pose dimensions, built from RealNVP coupling layers, whose density can be evaluated exactly and whose latent space can be sampled to produce arbitrary numbers of hypotheses. Its condition vector concatenates an HRNet image context, an MLP embedding of the detector's most likely 2D pose, and the normalized bounding-box location from CLIFF, so the flow sees both image evidence and detector uncertainty context. The uncertainty-transfer mechanism is the joint-wise Maximum Mean Discrepancy between projected flow samples and heatmap samples, with a mixture of inverse multiquadratics kernels; because MMD is sample based, the full heatmap distribution is used without being simplified to a Gaussian. The mask mechanism is an L1 loss that moves outside-mask projected samples toward the closest corresponding heatmap sample, removing invalid diversity while leaving genuine ambiguity intact. The two introduced metrics, PercIn (percentage of generated joint samples inside the person mask) and MinDist (minimum distance to the mask), turn the invisible-joint failure mode into a measurable quantity.
What would settle it
Compute a calibration plot for ViTPose's maximum heatmap confidence against manually labeled joint visibility on the EMDB subset: if a substantial fraction of invisible joints receive confidence above the 0.5 threshold, then the same mislabeled visibility enters training (which joints get MMD supervision and mask loss) and evaluation (which joints count as visible for 2D consistency), so the reported plausibility and consistency improvements would be partly an artifact of the detector's errors rather than of the learned distribution.
Extended reading notes
Core claim
The central claim is that the posterior over 3D body poses is learned more faithfully when the model is forced to match the uncertainty of a 2D pose detector than when it only maximizes ground-truth likelihood. The method models SMPL pose parameters with a conditional normalizing flow built from RealNVP coupling layers, conditioned on image features, the most likely 2D pose from ViTPose, and crop location information, while shape and camera are regressed deterministically. During training, 2D projections of flow samples are compared, for highly articulated and uncertain joints, with samples drawn from ViTPose's heatmaps using the Maximum Mean Discrepancy with a mixture of inverse multiquadratics kernels, which needs no Gaussian simplification of the heatmaps. A segmentation-mask loss pulls projected samples that fall outside unioned person masks toward the closest heatmap sample, so hypotheses that place an occluded joint in a visible location are penalized. The paper reports best-of-100 errors of 46.2 mm MPJPE, 29.8 mm PA-MPJPE, and 54.4 mm PVE on 3DPW, and 63.6, 40.9, and 72.0 mm on EMDB, together with an increase in the percentage of plausible invisible-joint hypotheses to 91.4% and a minimum mask distance of 0.616 pixels on the EMDB subset.
Load-bearing premise
The load-bearing premise is that a 2D pose detector's heatmap confidence values reliably indicate both where a joint is and whether it is visible; if the detector is overconfident or miscalibrated, the MMD supervision, the mask loss, and the visibility-based evaluation are all biased in the same direction.
Editorial extensions
If this is right
- Because the supervision signal comes from heatmaps of a 2D detector trained on large-scale 2D pose data, the approach can learn meaningful 3D distributional structure without needing 3D annotations for every ambiguous pose in the training set.
- The learned flow assigns exact likelihoods to hypotheses, so it can serve as an image-conditioned prior for parametric body fitting and for downstream tasks that need uncertainty estimates.
- The mask-based metrics give a standard way to detect hallucinated visible joints: any multi-hypothesis model that scores well on MPJPE but poorly on PercIn/MinDist is generating physically implausible samples for occluded body parts.
- The mask loss removes only invalid diversity: on EMDB, plausibility improves (PercIn from 88.8 to 91.4, MinDist from 1.049 to 0.616 pixels) while distribution accuracy stays essentially unchanged.
Reading between the lines
- Beyond the paper's claims: the calibration gap the authors acknowledge implies a testable fix—replacing the 0.5 confidence threshold with explicitly learned visibility scores should remove the shared bias that currently enters both the MMD supervision and the evaluation protocol in the same direction.
- Beyond the paper's claims: because MMD is applied joint-wise, the model treats joints as independent; a pose-level extension that accounts for kinematic correlations could tighten the posterior without reducing the diversity that occlusion requires.
- Beyond the paper's claims: the mask-loss idea assumes unioned person masks cover all plausible placements of occluded joints; for object occlusions the paper already excludes such training examples, so extending it to general object occlusion would require object-aware or amodal masks.
- Beyond the paper's claims: the heatmap-distillation idea transfers to other ill-posed inverse problems, such as hand or animal mesh recovery, whenever a 2D detector with informative, reasonably calibrated heatmaps is available.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a probabilistic 3D human mesh recovery method based on normalizing flows. The key contribution is to supervise the learned distribution over SMPL pose parameters not only with ground-truth 3D poses and 2D keypoints but also with the distributions encoded in the heatmaps of a 2D pose detector (ViTPose). This is done by minimizing the Maximum Mean Discrepancy (MMD) between heatmap samples and projections of generated 3D hypotheses, applied only to uncertain, highly articulated joints. A second contribution is a mask-based loss that penalizes hypotheses whose joints project outside the person segmentation mask, together with two new evaluation metrics (PercIn and MinDist) that measure the plausibility of hypotheses for invisible joints. Experiments on 3DPW and EMDB report consistent improvements over a retrained ProHMR baseline and other probabilistic methods in terms of best-of-N accuracy, input consistency, diversity, and mask plausibility.
Significance. If the results are taken at face value, the paper makes a useful practical contribution: it demonstrates a concrete way to distill uncertainty information from a 2D pose detector into a probabilistic 3D human mesh model, and it draws attention to a real failure mode of existing multi-hypothesis methods, namely the generation of visible hypotheses for occluded joints. The introduced mask-based plausibility metrics are a reasonable step toward evaluating this failure mode. The experimental setup is mostly careful: the authors retrain ProHMR on the same data, use the same backbone, provide an ablation study on both 3DPW and EMDB, and release code. The main caveats are that the evaluation of invisible-joint plausibility shares the same ViTPose-confidence proxy that is used for supervision, that the mask metrics are computed on a manually filtered subset of EMDB, and that HuManiFlow is evaluated with an external checkpoint. These caveats temper the strength of the claims but do not invalidate the core accuracy results, which are measured against independent 3D ground truth.
major comments (4)
- [Sec. 5.1; Sec. 4.2–4.3] The evaluation protocol in Sec. 5.1 defines joint visibility from ViTPose maximum-confidence values with a threshold of 0.5, and the same confidence values are used during training to define uncertain joints for the MMD loss (Sec. 4.2) and to define invisible joints for the mask loss (Sec. 4.3). If ViTPose confidences are miscalibrated, as the authors acknowledge in Sec. D citing Gu et al. [20], then the supervision signal and the evaluation metrics for “invisible-joint plausibility” (PercIn/MinDist in Table 2) are biased in the same direction, so the reported improvements could partly reflect alignment with the proxy rather than genuine image consistency. This concern is not merely theoretical: the paper's own Limitations section states that ViTPose “sometimes tends to be overconfident.” To support the central claim of plausible diversity, please (i) report PercIn/MinDist and the 2DKP consistency metrics for a range of confidence thresholds or with an independent visibility estimator (e.g., [20,69]), and (ii) discuss how the threshold choice affects the ranking versus HuManiFlow and ProHMR†. Without this, the mask-based plausibility comparison is not robust to the choice of visibility proxy.
- [Table 1; Sec. 5.1] The distribution accuracy in Table 1 is reported as the minimum error over 100 hypotheses. While best-of-N is an established oracle metric for multi-hypothesis methods, it rewards methods that produce a heavy tail of samples that happen to include the ground truth, and the reported gains over competitors may be driven by a few lucky samples. Please additionally report the mean (or expected) MPJPE/PA-MPJPE/PVE over the 100 samples, which is a more faithful measure of the quality of the full distribution. The 2DKP error in Table 2 is already averaged over samples, so this addition would not be inconsistent with the paper's own emphasis on distribution quality.
- [Sec. 5.1; Table 2] The mask metrics PercIn/MinDist are evaluated only on a manually filtered subset of EMDB consisting of 1760 person masks. The filtering criteria are not specified beyond “manually filtering for quality” and selecting images with at least two invisible keypoints. Because the subset is chosen by the authors, it is unclear whether the comparison in Table 2 is balanced across methods or favorable to the proposed model. Please provide the complete filtering protocol, report the number of images/joints used for each method, and, if possible, also report the mask metrics on the full EMDB set and on 3DPW with automatically generated masks. This would make the plausibility evaluation reproducible and less dependent on subjective selection.
- [Tables 1 and 2] HuManiFlow is evaluated with the publicly available checkpoint (marked *) rather than being retrained on the same training data as the proposed method and ProHMR†. Since the paper's diversity and plausibility comparisons against HuManiFlow are central to the claim of maintaining high diversity, the reported differences could reflect training-data differences rather than algorithmic improvements. The authors state that they could not reproduce HuManiFlow training; in that case, the potential impact of different training data should be discussed explicitly, and the conclusions involving HuManiFlow should be tempered accordingly. A retrained HuManiFlow baseline on the same three datasets would be the cleanest fix.
minor comments (3)
- [Sec. 5.1] The description of the mask evaluation does not state whether the union of person masks or the single target-person mask is used for PercIn/MinDist. Since Sec. 4.3 uses the union of all person masks for training, please clarify which mask is used in evaluation.
- [Sec. 4.5] The choice of the uncertainty threshold (0.7) and the visibility threshold (0.5) is not justified or ablated. Given that the method's behavior depends on these thresholds, a small sensitivity analysis would improve confidence.
- [Eq. (7)] The loss weights in Eq. (7) are given only in the supplementary material; a brief mention in the main text would help the reader understand the relative importance of the new MMD and mask losses.
Circularity Check
Secondary diversity/plausibility evaluation shares its visibility proxy with the training supervision; central accuracy claims are independent.
-
self definitional
[Sec. 4.5 (Implementation Details) and Sec. 5.1 (Datasets and Evaluation Metrics); acknowledged in Sec. D (Limitations)]
"Furthermore, we use the maximum confidence values as proxy for the visibility of joints, and set the threshold to 0.5. ... We define a ground-truth joint to be invisible if its location is outside the image crop or its corresponding maximum confidence score is below 0.5. ... ViTPose confidence values with a threshold of 0.5 are used as proxy for the visibility of keypoints."
The same ViTPose maximum-confidence threshold selects which joints are treated as uncertain/invisible during training (gate for LMMD and mask-loss supervision, and definition of invisible joints in Sec. 4.5) and which joints count as visible or invisible in the evaluation metrics of Sec. 5.1 (2DKP error on 'visible' joints; 3DKP spread split into visible/invisible; diversity expected to be higher for invisible joints). Because the model is explicitly optimized via MMD against ViTPose heatmap samples to have high diversity for ViTPose-uncertain joints and low diversity for ViTPose-certain joints, the reported 'meaningful diversity' and consistency advantages are measured with the same proxy that generated the supervision. If ViTPose is miscalibrated, as the authors concede in Sec.
full rationale
The main distribution-accuracy claims (Table 1: MPJPE, PA-MPJPE, PVE) are computed against independent 3D ground-truth SMPL annotations and are not circular; the best-of-100 numbers do not reduce to the heatmap supervision signal. The mask-plausibility results are also a genuine generalization test: mask supervision is trained on synthetic BEDLAM/AGORA masks, while the PercIn/MinDist metrics are evaluated on EMDB with Mask R-CNN masks. The only exhibited coupling is the shared ViTPose confidence threshold for defining visibility in both training and evaluation, which the authors themselves flag as a calibration risk in Sec. D. This affects the secondary 'meaningful diversity for ambiguous body parts' claim but does not undermine the independent accuracy benchmarks, so the overall circularity score is low.
Assumptions & free parameters
free parameters (5)
- loss_weights =
λβ=5e-4, λ2D=1e-2, λNLL=1e-1, λorth=1e-1, λMMD=5e-2, λmask=1e-1
- uncertainty_threshold =
0.7
- visibility_threshold =
0.5
- num_heatmap_samples_n =
25
- mmd_bandwidths_B =
{0.05, 0.20, 0.90}
assumptions (6)
- domain assumption SMPL is a valid parametric model of the human body
- domain assumption Normalizing flows can express the conditional posterior of SMPL pose given an image
- standard math The change-of-variables formula is valid for the composed flows
- domain assumption ViTPose heatmaps encode meaningful joint occurrence probabilities
- domain assumption The union of person masks in an image covers all plausible locations of invisible joints
- domain assumption Training on BEDLAM, AGORA, and 3DPW generalizes to EMDB
Cite this review
Pith. "Pith review of Utilizing Uncertainty in 2D Pose Detectors for Probabilistic 3D Human Mesh Recovery." pith.science (2026). https://pith.science/paper/6Z5E23R2
@misc{pith2026241116289,
author = {Pith},
title = {Pith review of: Utilizing Uncertainty in 2D Pose Detectors for Probabilistic 3D Human Mesh Recovery},
year = {2026},
howpublished = {\url{https://pith.science/paper/6Z5E23R2}},
note = {Machine review of arXiv:2411.16289}
}
read the original abstract
Monocular 3D human pose and shape estimation is an inherently ill-posed problem due to depth ambiguities, occlusions, and truncations. Recent probabilistic approaches learn a distribution over plausible 3D human meshes by maximizing the likelihood of the ground-truth pose given an image. We show that this objective function alone is not sufficient to best capture the full distributions. Instead, we propose to additionally supervise the learned distributions by minimizing the distance to distributions encoded in heatmaps of a 2D pose detector. Moreover, we reveal that current methods often generate incorrect hypotheses for invisible joints which is not detected by the evaluation protocols. We demonstrate that person segmentation masks can be utilized during training to significantly decrease the number of invalid samples and introduce two metrics to evaluate it. Our normalizing flow-based approach predicts plausible 3D human mesh hypotheses that are consistent with the image evidence while maintaining high diversity for ambiguous body parts. Experiments on 3DPW and EMDB show that we outperform other state-of-the-art probabilistic methods. Code is available for research purposes at https://github.com/twehrbein/humr.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[20]
On the calibration of human pose estimation
Kerui Gu, Rongyu Chen, Xuanlong Yu, and Angela Yao. On the calibration of human pose estimation. InICML, 2024. 14
work page 2024
-
[1]
2d human pose estimation: New benchmark and state of the art analysis
Mykhaylo Andriluka, Leonid Pishchulin, Peter Gehler, and Bernt Schiele. 2d human pose estimation: New benchmark and state of the art analysis. In CVPR, 2014. 2
2014
-
[2]
Lynton Ardizzone, Till Bungert, Felix Draxler, Ullrich K¨othe, Jakob Kruse, Robert Schmier, and Peter Sorren- son. Framework for Easily Invertible Architectures (FrEIA), https://github.com/vislearn/FrEIA, August 2024. 12
work page 2024
-
[3]
Wirkert, Daniel Rahner, Eric W
Lynton Ardizzone, Jakob Kruse, Sebastian J. Wirkert, Daniel Rahner, Eric W. Pellegrini, Ralf S. Klessen, Lena Maier- Hein, Carsten Rother, and Ullrich K¨othe. Analyzing inverse problems with invertible neural networks. In ICLR, 2019. 4, 5
work page 2019
-
[4]
3d multi- bodies: Fitting sets of plausible 3d human models to am- biguous image data
Benjamin Biggs, S ´ebastien Ehrhadt, Hanbyul Joo, Benjamin Graham, Andrea Vedaldi, and David Novotny. 3d multi- bodies: Fitting sets of plausible 3d human models to am- biguous image data. In NeurIPS, 2020. 1, 3, 7
work page 2020
-
[5]
Christopher M. Bishop. Mixture density networks. Technical report, Aston University, 1994. 3
1994
-
[6]
Black, Priyanka Patel, Joachim Tesch, and Jin- long Yang
Michael J. Black, Priyanka Patel, Joachim Tesch, and Jin- long Yang. BEDLAM: A synthetic dataset of bodies exhibit- ing detailed lifelike animated motion. In CVPR, 2023. 6, 7, 12
work page 2023
-
[7]
Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J. Black. Keep it smpl: Automatic estimation of 3d human pose and shape from a single image. In ECCV, 2016. 2
work page 2016
Show all 81 references
-
[8]
Mhentropy: Entropy meets multiple hypotheses for pose and shape re- covery
Rongyu Chen, Linlin Yang, and Angela Yao. Mhentropy: Entropy meets multiple hypotheses for pose and shape re- covery. In ICCV, 2023. 1, 3
2023
-
[9]
Kiam Choo and D.J. Fleet. People tracking using hybrid monte carlo filtering. In ICCV, 2001. 3
2001
-
[10]
NICE: non-linear independent components estimation
Laurent Dinh, David Krueger, and Yoshua Bengio. NICE: non-linear independent components estimation. In ICLR,
-
[11]
Density estimation using real NVP
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. Density estimation using real NVP. In ICLR, 2017. 3, 4, 12
2017
-
[12]
Normalizing flows on the product space of so(3) manifolds for probabilis- tic human pose modeling
Olaf D ¨unkel, Tim Salzmann, and Florian Pfaff. Normalizing flows on the product space of so(3) manifolds for probabilis- tic human pose modeling. In CVPR, 2024. 3
2024
-
[13]
Black, and Dimitrios Tzionas
Sai Kumar Dwivedi, Cordelia Schmid, Hongwei Yi, Michael J. Black, and Dimitrios Tzionas. POCO: 3D pose and shape estimation using confidence. In 3DV, 2024. 3
2024
-
[14]
Sai Kumar Dwivedi, Yu Sun, Priyanka Patel, Yao Feng, and Michael J. Black. TokenHMR: Advancing human mesh re- covery with a tokenized pose representation. InCVPR, 2024. 3
2024
-
[15]
Learning analytical posterior probabil- ity for human mesh recovery
Qi Fang, Kang Chen, Yinghui Fan, Qing Shuai, Jiefeng Li, and Weidong Zhang. Learning analytical posterior probabil- ity for human mesh recovery. In CVPR, 2023. 3
2023
-
[16]
Distribution-aligned diffusion for human mesh recovery
Lin Geng Foo, Jia Gong, Hossein Rahmani, and Jun Liu. Distribution-aligned diffusion for human mesh recovery. In ICCV, 2023. 3
2023
-
[17]
Made: Masked autoencoder for distribution es- timation
Mathieu Germain, Karol Gregor, Iain Murray, and Hugo Larochelle. Made: Masked autoencoder for distribution es- timation. In ICML, 2015. 3
2015
-
[18]
Cubuk, Quoc V
Golnaz Ghiasi, Yin Cui, Aravind Srinivas, Rui Qian, Tsung- Yi Lin, Ekin D. Cubuk, Quoc V . Le, and Barret Zoph. Simple copy-paste is a strong data augmentation method for instance segmentation. In CVPR, 2021. 6
2021
-
[19]
Borgwardt, Malte J
Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Sch ¨olkopf, and Alexander Smola. A kernel two-sample test. Journal of Machine Learning Research , 13(25):723–773, 2012. 2
2012
-
[21]
B ˜alan, and Michael J
Peng Guan, Alexander Weiss, Alexandru O. B ˜alan, and Michael J. Black. Estimating human shape and pose from a single image. In ICCV, 2009. 2
2009
-
[22]
Instance-aware contrastive learning for oc- cluded human mesh reconstruction
Mi-Gyeong Gwon, Gi-Mun Um, Won-Sik Cheong, and Wonjun Kim. Instance-aware contrastive learning for oc- cluded human mesh reconstruction. In CVPR, 2024. 3
2024
-
[23]
Multilinear pose and body shape estimation of dressed subjects from image sets
Nils Hasler, Hanno Ackermann, Bodo Rosenhahn, Thorsten Thorm¨ahlen, and Hans-Peter Seidel. Multilinear pose and body shape estimation of dressed subjects from image sets. In CVPR, 2010. 2
2010
-
[24]
Girshick
Kaiming He, Georgia Gkioxari, Piotr Doll ´ar, and Ross B. Girshick. Mask r-cnn. In ICCV, 2017. 6
2017
-
[25]
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In CVPR,
-
[26]
Denoising diffu- sion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. In NeurIPS, 2020. 3
2020
-
[27]
Diffpose: Multi- hypothesis human pose estimation using diffusion models
Karl Holmquist and Bastian Wandt. Diffpose: Multi- hypothesis human pose estimation using diffusion models. In ICCV, 2023. 1, 2, 3, 4, 13, 14
2023
-
[28]
Ehsan Jahangiri and Alan L. Yuille. Generating multiple di- verse hypotheses for human 3d pose consistent with 2d joint detections. In ICCV Workshops, 2017. 1
2017
-
[29]
Black, David W
Angjoo Kanazawa, Michael J. Black, David W. Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In CVPR, 2018. 2, 3, 4
2018
-
[30]
EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild
Manuel Kaufmann, Jie Song, Chen Guo, Kaiyue Shen, Tian- jian Jiang, Chengcheng Tang, Juan Jos ´e Z ´arate, and Otmar Hilliges. EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild. In ICCV, 2023. 2, 6, 7, 12
2023
-
[31]
Sapiens: Foundation for human vision mod- els
Rawal Khirodkar, Timur Bagautdinov, Julieta Martinez, Su Zhaoen, Austin James, Peter Selednik, Stuart Anderson, and Shunsuke Saito. Sapiens: Foundation for human vision mod- els. In ECCV, 2024. 6
2024
-
[32]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 12
2015
-
[33]
Kingma and Prafulla Dhariwal
Durk P. Kingma and Prafulla Dhariwal. Glow: Generative flow with invertible 1x1 convolutions. In NeurIPS, 2018. 3
2018
-
[34]
Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling
Durk P. Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling. Improved variational in- ference with inverse autoregressive flow. In NeurIPS, 2016. 2, 3
2016
-
[35]
Beyond weak perspective for monocular 3d human pose estimation
Imry Kissos, Lior Fritz, Matan Goldman, Omer Meir, Ed- uard Oks, and Mark Kliger. Beyond weak perspective for monocular 3d human pose estimation. In ECCV Workshops,
-
[36]
Huang, Otmar Hilliges, and Michael J
Muhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, and Michael J. Black. PARE: Part attention regressor for 3D human body estimation. In ICCV, 2021. 3
2021
-
[37]
Black, and Kostas Daniilidis
Nikos Kolotouros, Georgios Pavlakos, Michael J. Black, and Kostas Daniilidis. Learning to reconstruct 3d human pose and shape via model-fitting in the loop. In ICCV, 2019. 3
2019
-
[38]
Probabilistic modeling for human mesh recovery
Nikos Kolotouros, Georgios Pavlakos, Dinesh Jayaraman, and Kostas Daniilidis. Probabilistic modeling for human mesh recovery. In ICCV, 2021. 1, 2, 3, 4, 5, 6, 7, 8, 12, 13, 16
2021
-
[39]
Black, and Peter V
Christoph Lassner, Javier Romero, Martin Kiefel, Federica Bogo, Michael J. Black, and Peter V . Gehler. Unite the peo- ple: Closing the loop between 3d and 2d human representa- tions. In CVPR, 2017. 2
2017
-
[40]
Determination of 3d human body postures from a single view
Hsi-Jian Lee and Zen Chen. Determination of 3d human body postures from a single view. Computer Vision, Graph- ics, and Image Processing, 30(2):148–168, 1985. 3
1985
-
[41]
Proposal maps driven mcmc for estimating human body pose in static images
Mun Wai Lee and Isaac Cohen. Proposal maps driven mcmc for estimating human body pose in static images. In CVPR,
-
[42]
Generating multiple hypotheses for 3d human pose estimation with mixture density network
Chen Li and Gim Hee Lee. Generating multiple hypotheses for 3d human pose estimation with mixture density network. In CVPR, 2019. 1, 3
2019
-
[43]
Crowdpose: Efficient crowded scenes pose estimation and a new benchmark
Jiefeng Li, Can Wang, Hao Zhu, Yihuan Mao, Hao-Shu Fang, and Cewu Lu. Crowdpose: Efficient crowded scenes pose estimation and a new benchmark. In CVPR, 2019. 2
2019
-
[44]
Hybrik: A hybrid analytical-neural inverse kinematics solution for 3d human pose and shape estimation
Jiefeng Li, Chao Xu, Zhicun Chen, Siyuan Bian, Lixin Yang, and Cewu Lu. Hybrik: A hybrid analytical-neural inverse kinematics solution for 3d human pose and shape estimation. In CVPR, 2021. 3
2021
-
[45]
Rosin, and Shi-Min Hu
Ruilong Li, Xin Dong, Zixi Cai, Dingcheng Yang, Haozhi Huang, Song-Hai Zhang, Paul L. Rosin, and Shi-Min Hu. Pose2seg: Human instance segmentation without detection. In CVPR, 2019. 2
2019
-
[46]
Mhformer: Multi-hypothesis transformer for 3d human pose estimation
Wenhao Li, Hong Liu, Hao Tang, Pichao Wang, and Luc Van Gool. Mhformer: Multi-hypothesis transformer for 3d human pose estimation. In CVPR, 2022. 1
2022
-
[47]
Cliff: Carrying location information in full frames into human pose and shape estimation
Zhihao Li, Jianzhuang Liu, Zhensong Zhang, Songcen Xu, and Youliang Yan. Cliff: Carrying location information in full frames into human pose and shape estimation. In ECCV,
-
[48]
Belongie, Lubomir D
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C. Lawrence Zitnick. Microsoft COCO: common objects in context. In ECCV, 2014. 2, 5, 6
2014
-
[49]
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. SMPL: A skinned multi- person linear model. In ACM Trans. Graphics (Proc. SIG- GRAPH Asia). ACM, 2015. 2, 3, 12
2015
-
[50]
Oikarinen, Daniel C
Tuomas P. Oikarinen, Daniel C. Hannah, and Sohrob Kaze- rounian. Graphmdn: Leveraging graph structure and deep learning to solve inverse problems. In IJCNN, 2021. 1, 3
2021
-
[51]
Masked autoregressive flow for density estimation
George Papamakarios, Theo Pavlakou, and Iain Murray. Masked autoregressive flow for density estimation. In NeurIPS, 2017. 3
2017
-
[52]
Huang, Joachim Tesch, David T
Priyanka Patel, Chun-Hao P. Huang, Joachim Tesch, David T. Hoffmann, Shashank Tripathi, and Michael J. Black. AGORA: Avatars in geography optimized for regres- sion analysis. In CVPR, 2021. 6
2021
-
[53]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3d hands, face, and body from a single image. In CVPR, 2019. 2, 12
2019
-
[54]
Coarse-to-fine volumetric predic- tion for single-image 3D human pose
Georgios Pavlakos, Xiaowei Zhou, Konstantinos G Derpa- nis, and Kostas Daniilidis. Coarse-to-fine volumetric predic- tion for single-image 3D human pose. In CVPR, 2017. 12
2017
-
[55]
Sam 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junt- ing Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao- Yuan Wu, Ross Girshick, Piotr Doll´ar, and Christoph Feic...
2024 arXiv
-
[56]
Variational inference with normalizing flows
Danilo Rezende and Shakir Mohamed. Variational inference with normalizing flows. InProceedings of Machine Learning Research (PMLR), 2015. 2
2015
-
[57]
Hi- erarchical kinematic probability distributions for 3d human shape and pose estimation from images in the wild
Akash Sengupta, Ignas Budvytis, and Roberto Cipolla. Hi- erarchical kinematic probability distributions for 3d human shape and pose estimation from images in the wild. InICCV,
-
[58]
Prob- abilistic 3d human shape and pose estimation from multiple unconstrained images in the wild
Akash Sengupta, Ignas Budvytis, and Roberto Cipolla. Prob- abilistic 3d human shape and pose estimation from multiple unconstrained images in the wild. In CVPR, 2021. 1, 3, 7
2021
-
[59]
Hu- maniflow: Ancestor-conditioned normalising flows on so(3) manifolds for human pose and shape distribution estimation
Akash Sengupta, Ignas Budvytis, and Roberto Cipolla. Hu- maniflow: Ancestor-conditioned normalising flows on so(3) manifolds for human pose and shape distribution estimation. In CVPR, 2023. 1, 3, 6, 7, 8, 12, 13, 14, 16
2023
-
[60]
Monocular 3d human pose estimation by generation and ordinal ranking
Saurabh Sharma, Pavan Teja Varigonda, Prashast Bindal, Abhishek Sharma, and Arjun Jain. Monocular 3d human pose estimation by generation and ordinal ranking. In ICCV,
-
[61]
Com- bined discriminative and generative articulated pose and non-rigid shape estimation
Leonid Sigal, Alexandru Balan, and Michael Black. Com- bined discriminative and generative articulated pose and non-rigid shape estimation. In NeurIPS, 2007. 2
2007
-
[62]
Simo-Serra, A
E. Simo-Serra, A. Ramisa, G. Aleny `a, C. Torras, and F. Moreno-Noguer. Single image 3d human pose estimation from noisy observations. In CVPR, 2012. 3
2012
-
[63]
Covariance scaled sampling for monocular 3d body tracking
Cristian Sminchisescu and Bill Triggs. Covariance scaled sampling for monocular 3d body tracking. In CVPR, 2001. 3
2001
-
[64]
Hyperdynamics im- portance sampling
Cristian Sminchisescu and Bill Triggs. Hyperdynamics im- portance sampling. In ECCV, 2002. 3
2002
-
[65]
Kinematic jump pro- cesses for monocular 3d human tracking
Cristian Sminchisescu and Bill Triggs. Kinematic jump pro- cesses for monocular 3d human tracking. In CVPR, 2003. 3
2003
-
[66]
Posturehmr: Posture transformation for 3d hu- man mesh recovery
Yu-Pei Song, Xiao Wu, Zhaoquan Yuan, Jian-Jun Qiao, and Qiang Peng. Posturehmr: Posture transformation for 3d hu- man mesh recovery. In CVPR, 2024. 3
2024
-
[67]
Virtualpose: Learning generalizable 3d hu- man pose models from virtual data
Jiajun Su, Chunyu Wang, Xiaoxuan Ma, Wenjun Zeng, and Yizhou Wang. Virtualpose: Learning generalizable 3d hu- man pose models from virtual data. In ECCV, 2022. 12
2022
-
[68]
Deep high-resolution representation learning for human pose esti- mation
Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang. Deep high-resolution representation learning for human pose esti- mation. In CVPR, 2019. 2, 3, 4, 6 10
2019
-
[69]
Rethinking visibility in human pose estimation: Occluded pose reasoning via transformers
Pengzhan Sun, Kerui Gu, Yunsong Wang, Linlin Yang, and Angela Yao. Rethinking visibility in human pose estimation: Occluded pose reasoning via transformers. In WACV, 2024. 14
2024
-
[70]
Wasserstein auto-encoders
Ilya Tolstikhin, Olivier Bousquet, Sylvain Gelly, and Bern- hard Sch¨olkopf. Wasserstein auto-encoders. In ICLR, 2018. 5
2018
-
[71]
Black, Bodo Rosenhahn, and Gerard Pons-Moll
Timo von Marcard, Roberto Henschel, Michael J. Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering ac- curate 3D human pose in the wild using IMUs and a moving camera. In ECCV, 2018. 2, 6, 7, 12
2018
-
[72]
Refit: Recurrent fitting network for 3d human recovery
Yufu Wang and Kostas Daniilidis. Refit: Recurrent fitting network for 3d human recovery. In ICCV, 2023. 3
2023
-
[73]
Personalized 3d human pose and shape re- finement
Tom Wehrbein, Bodo Rosenhahn, Iain Matthews, and Carsten Stoll. Personalized 3d human pose and shape re- finement. In ICCV Workshops, 2023. 3
2023
-
[74]
Probabilistic monocular 3d human pose estima- tion with normalizing flows
Tom Wehrbein, Marco Rudolph, Bodo Rosenhahn, and Bas- tian Wandt. Probabilistic monocular 3d human pose estima- tion with normalizing flows. In ICCV, 2021. 1, 2, 3
2021
-
[75]
Learning likelihoods with conditional normal- izing flows
Christina Winkler, Daniel Worrall, Emiel Hoogeboom, and Max Welling. Learning likelihoods with conditional normal- izing flows. arXiv preprint arXiv:1912.00042, 2019. 3
1912 arXiv
-
[76]
Large-scale datasets for going deeper in image understanding
Jiahong Wu, He Zheng, Bo Zhao, Yixin Li, Baoming Yan, Rui Liang, Wenjia Wang, Shipei Zhou, Guosen Lin, Yanwei Fu, Yizhou Wang, and Yonggang Wang. Large-scale datasets for going deeper in image understanding. In 2019 IEEE International Conference on Multimedia and Expo (ICME) ,
2019
-
[77]
Scorehypo: Probabilistic human mesh estimation with hypothesis scoring
Yuan Xu, Xiaoxuan Ma, Jiajun Su, Wentao Zhu, Yu Qiao, and Yizhou Wang. Scorehypo: Probabilistic human mesh estimation with hypothesis scoring. In CVPR, 2024. 1, 3, 7, 12
2024
-
[78]
ViTPose: Simple vision transformer baselines for human pose estimation
Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. ViTPose: Simple vision transformer baselines for human pose estimation. In NeurIPS, 2022. 2, 3, 4, 5, 12, 13, 16
2022
-
[79]
Pymaf: 3d human pose and shape regression with pyramidal mesh alignment feedback loop
Hongwen Zhang, Yating Tian, Xinchi Zhou, Wanli Ouyang, Yebin Liu, Limin Wang, and Zhenan Sun. Pymaf: 3d human pose and shape regression with pyramidal mesh alignment feedback loop. In ICCV, 2021. 3
2021
-
[80]
Probabilistic human mesh recovery in 3d scenes from egocentric views
Siwei Zhang, Qianli Ma, Yan Zhang, Sadegh Aliakbarian, Darren Cosker, and Siyu Tang. Probabilistic human mesh recovery in 3d scenes from egocentric views. InICCV, 2023. 3, 5
2023
-
[81]
On the continuity of rotation representations in neural networks
Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. On the continuity of rotation representations in neural networks. In CVPR, 2019. 5 11 A. Implementation Details Network and training. The normalizing flow (NF) consists of eight RealNVP [11] coupling layers, each pa...
2019
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.