REVIEW 4 major objections 7 minor 1 cited by
EHPE: A Segmented Architecture for Enhanced Hand Pose Estimation
T0 review · 4 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The central claim of this paper is that estimating the five fingertip joints and the wrist first, then using them as structural anchors for the remaining joints, reduces hand-pose error to 5.6/5.7 mm PA-MPJPE/PA-MPVPE on FreiHAND and…
desk verdict Solid segmented-architecture hand pose paper, but the state-of-the-art claim on InterHand is not established as written. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the segmented two-stage pipeline driven by the kinematic assumption that fingertip and wrist positions nearly determine the other joints. In the first stage, an Hourglass network plus residual refinement modules form a 2.5D heatmap (with depth dimension $d=8$), and soft-argmax extracts the six anchor coordinates $(x,y,z)$. In the second stage, the Structural Prior Based Inference (SPI) module runs two layers of graph attention with dynamic edge weights over the 21-joint graph, while a Feature Enhancement Module (FEM) processes image features with self- and cross-attention; the outputs are combined as $\hat{c}_i = \omega_G\, \mathrm{SPI}_i + \omega_E\, \mathrm{FEM}_i$ with learnable weights. The dynamic graph attention is what converts the anchors into a structural prior, and the frozen two-stage training keeps the anchor estimates from being contaminated by distal-joint error.
What would settle it
Replace the TW-stage's predicted fingertip and wrist coordinates with ground-truth coordinates at inference and measure the remaining joints' error; if the residual error of the middle and proximal finger joints does not drop well below the full model's error, the anchors are not the error source the paper identifies, and the reported gains must come from added model capacity or the second training stage.
Extended reading notes
Core claim
The paper's central claim is that the error of the distal phalanx tip (TIP) joints is the main source of overall hand-pose error, and that estimating the five TIP joints and the wrist first, then conditioning the rest of the pose on them, reduces error for every joint category. The TW-stage produces a 2.5D heatmap for the six anchor joints and extracts their coordinates with soft-argmax. The PG-stage then builds a joint feature vector from each anchor's geometry and sampled image features, passes it through a structural-prior branch with two graph-attention layers using dynamic edge weights and a visual branch with self- and cross-attention, and fuses both branches with learnable weights. The authors verify the design with ablations that show estimating extra joints in the first stage hurts accuracy, that removing either branch raises error, and that dynamic edge weights beat fixed ones. On its own terms, the paper establishes that this segmented anchoring is what produces the reported state-of-the-art results.
Load-bearing premise
The pipeline rests on the kinematic premise that knowing the five fingertip positions relative to the wrist almost determines the other fifteen joints; if middle-joint flexion can vary independently while the fingertip stays put, the structural-prior branch is not carrying the information the paper attributes to it.
Editorial extensions
If this is right
- On FreiHAND, EHPE with the FastViT-MA36 backbone reaches PA-MPJPE/PA-MPVPE of 5.6/5.7 mm, a 1.0 mm PA-MPJPE improvement over the FastViT baseline and the best numbers among the methods compared.
- On InterHand2.6M, EHPE reports 5.73/5.87 mm MPJPE/MPVPE, ahead of EANet's 5.88/6.04 mm and other listed single- and two-hand methods.
- Ablations indicate that putting extra joints such as DIP or PIP into the first stage raises error to 5.8–6.5 mm, so the TIP-and-wrist pairing is doing the structural work rather than any arbitrary splitting of joints.
- Removing either the structural-prior branch or the visual branch degrades accuracy to 6.0–6.6 mm, and fixing the graph edge weights costs 0.9 mm, so the reported gain depends on both branches and on dynamic edge weighting.
- At inference the two stages operate end-to-end, so the segmented prior adds accuracy without requiring a separate optimization loop at run time.
Reading between the lines
- Editorial inference: if TIP and wrist truly anchor the remaining joints, the same anchoring should transfer to hand mesh regression, where distal phalanx vertices dominate error; conditioning mesh parameters on the six anchor joints is a natural extension the paper does not test.
- Editorial inference: the measured per-joint error hierarchy (TIP 100%, DIP roughly 80%, PIP roughly 72%, MCP roughly 59%, wrist roughly 40%) suggests that a joint-category reweighted loss might capture part of the gain without a two-stage architecture; this is a testable alternative explanation.
- Editorial inference: the kinematic premise predicts that the model's advantage grows with poses that have large middle-joint flexion or fingertip occlusion; a synthetic hand with controllable joint angles could isolate whether the structural-prior branch or added capacity produces the improvement.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes EHPE, a two-stage monocular 3D hand pose estimation architecture. In the first stage (TW-stage), the model predicts the five distal phalanx tips and the wrist using a 2.5D heatmap and soft-argmax. In the second stage (PG-stage), these predictions are used as structural anchors: one branch feeds them into a dynamic graph-attention module that estimates the remaining joints, while a second branch uses a self/cross-attention feature-enhancement module; the two outputs are fused with learned weights. The paper reports state-of-the-art results on FreiHAND (PA-MPJPE 5.6 mm, PA-MPVPE 5.7 mm) and InterHand2.6M (MPJPE 5.73 mm), and the ablations show that removing the TW-stage guidance, using fixed edge weights, or dropping either branch degrades accuracy. Code is released.
Significance. Assuming the reported numbers survive a controlled comparison, the paper makes a useful empirical point: treating TIP and wrist as explicit structural anchors and segmenting the estimation process can improve hand pose accuracy. The ablations in Tables 3-6 are internally consistent and support the core design choices, and the public code release makes the central experiments checkable. The contribution is primarily empirical, with no circular derivation; tuning loss weights per dataset is standard practice. The significance is moderated by the evaluation-provenance gaps in the InterHand2.6M comparison and the absence of uncertainty estimates, which currently make the headline state-of-the-art claim stronger than the evidence supports.
major comments (4)
- [Table 2; Section 4.3.1] The claim of state-of-the-art performance on InterHand2.6M is not yet supported. Table 2 lists EHPE(Ours) without identifying the backbone, even though Section 4.1 states that EHPE is trained with ResNet-50, HRNet, and FastViT-MA36; the closest baseline, EANet, uses ResNet-50. If the reported EHPE row uses a stronger backbone, the 0.15 mm MPJPE margin over EANet may reflect backbone capacity rather than the proposed architecture. The paper must state the backbone and evaluation protocol (which split of InterHand2.6M, cropping, alignment) and, ideally, rerun the comparison with matched backbones and multiple seeds before the state-of-the-art claim can be accepted.
- [Tables 1 and 2; Section 4.3.1] The 'surpasses all compared methods' statement is overstated. In Table 1, EHPE with FastViT-MA36 ties HaMeR on PA-MPVPE (5.7 vs 5.7), and several FreiHAND margins are 0.1 mm; in Table 2, the winning MPJPE margin over EANet is 0.15 mm. No confidence intervals, standard deviations, or seed information are reported, so these differences lie within plausible run-to-run noise. Reporting multiple runs or bootstrap intervals, or at minimum softening the wording, is necessary to support the central claim.
- [Section 1, second observation] The kinematic premise that 'once the positions of TIP relative to the wrist are determined, the relative positions of other joints can almost be locked' is overly strong. Each finger retains independent flexion and abduction degrees of freedom, so a fingertip position relative to the wrist does not uniquely determine the DIP, PIP, and MCP joint angles. The paper should either provide biomechanical or empirical support for this claim or reframe it as a heuristic; otherwise the method's motivation is overstated. This does not invalidate the ablations, but it weakens the stated rationale.
- [Section 4.4] The ablation studies are said to use 'lightweight model versions' but the exact configuration is never specified. Since the main results use ResNet-50, HRNet, and FastViT-MA36, it is unclear whether the ablation conclusions transfer to the full models. Please define the lightweight backbone, feature dimensions, and training schedule, or repeat at least the key ablation with the main backbone.
minor comments (7)
- [Eq. (1)] The notation uses i and j both as matrix indices and as the joint index; renaming the joint index to k would remove the ambiguity.
- [Eq. (2)] The symbol (x, y, d) is overloaded as both coordinate values and summation indices; the intended averaging over heatmap locations should be clarified.
- [Eq. (7)] The dimensions of q_SA, k_SA, and v_SA are given as (hw + hw + 1) x 512, which is unusual, and d_k is never defined; please clarify the attention formulation.
- [Eq. (11)] The shapes of omega_G (21x21) and omega_E (21x1) do not obviously multiply with the SPI and FEM outputs, which are expected to be 3D joint coordinates; the output dimensions of each branch should be made explicit.
- [Section 4.3.1] The term 'EABlock' appears in the discussion of InterHand2.6M results but is never defined; this is likely a typo for the SPI/FEM modules and should be corrected.
- [Throughout] There are numerous typographical errors, including 'freamwork', 'Comprised', 'in-replaceable', 'stricture prior', and 'ere' in the Figure 5 caption; these should be corrected in a revision.
- [Section 4.2] The description of InterHand2.6M training and test sizes (1.36M training, 849K test) does not match the standard split conventions for that dataset; please specify which subset and protocol are actually used.
Circularity Check
No significant circularity: the EHPE derivation is an empirical architecture proposal evaluated on external benchmarks, with no load-bearing step that reduces to its own inputs.
full rationale
The paper's central claim is that first estimating TIP and wrist joints and then using them as structural priors improves full hand pose estimation. This claim is supported by controlled ablations (Tables 3, 4, 5, 6) and by comparisons on two external benchmarks (FreiHAND and InterHand2.6M), not by a definitional identity or by fitting a parameter to the quantity being predicted. The TW-stage directly supervises TIP/wrist heatmaps with ground-truth heatmaps (Eqs. 1-3), and the PG-stage estimates remaining joints from those predicted anchors plus image features; the final output is a weighted fusion (Eq. 11) and the improvement over the TW-only and PG-only variants is measured, not assumed. Loss weights are hyperparameters tuned per dataset, which is standard practice and does not make any reported metric a fitted constant. The reused FEM module is explicitly credited to prior work [32] (EANet), and the author's own earlier publications cited in the reference list ([53, 54, 55]) concern bokeh rendering, image compression, and demosaicing; they are not load-bearing for the hand-pose arguments. No uniqueness theorem or prior-work-derived ansatz is invoked to forbid alternatives. The paper's kinematic premise that finger joints are nearly determined by TIP and wrist positions is a debatable anatomical assumption, and the InterHand2.6M comparison in Table 2 is incompletely specified (backbone not stated, few baselines, no uncertainty estimates), but those are evaluation-provenance or correctness concerns, not circularity: the reported numbers are externally benchmarked and the architecture could in principle fail on the stated benchmarks. Therefore, no step in the derivation chain is circular by construction, and the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (8)
- Heatmap loss weight lambda_H =
3
- Euclidean distance loss weight lambda_ED =
1e-2
- Regularization loss weight lambda_R =
1e-2
- Position loss weight lambda_P =
2e-2
- Edge loss weight lambda_E =
2e-1
- Number of graph attention heads K =
8
- Number of graph attention layers =
2
- Discrete depth resolution d =
8
assumptions (4)
- domain assumption Hand kinematic lock: once TIP positions relative to wrist are known, DIP/PIP/MCP positions are almost locked
- domain assumption Error accumulation order TIP > DIP > PIP > MCP > W is a stable cross-method phenomenon
- domain assumption FreiHAND and InterHand2.6M 3D annotations are treated as ground truth
- domain assumption A single 21-joint graph is sufficient for the two-hand InterHand2.6M setting
Cite this review
Pith. "Pith review of EHPE: A Segmented Architecture for Enhanced Hand Pose Estimation." pith.science (2026). https://pith.science/paper/KP7K5IVP
@misc{pith2026250709560,
author = {Pith},
title = {Pith review of: EHPE: A Segmented Architecture for Enhanced Hand Pose Estimation},
year = {2026},
howpublished = {\url{https://pith.science/paper/KP7K5IVP}},
note = {Machine review of arXiv:2507.09560}
}
read the original abstract
3D hand pose estimation has garnered great attention in recent years due to its critical applications in human-computer interaction, virtual reality, and related fields. The accurate estimation of hand joints is essential for high-quality hand pose estimation. However, existing methods neglect the importance of Distal Phalanx Tip (TIP) and Wrist in predicting hand joints overall and often fail to account for the phenomenon of error accumulation for distal joints in gesture estimation, which can cause certain joints to incur larger errors, resulting in misalignments and artifacts in the pose estimation and degrading the overall reconstruction quality. To address this challenge, we propose a novel segmented architecture for enhanced hand pose estimation (EHPE). We perform local extraction of TIP and wrist, thus alleviating the effect of error accumulation on TIP prediction and further reduce the predictive errors for all joints on this basis. EHPE consists of two key stages: In the TIP and Wrist Joints Extraction stage (TW-stage), the positions of the TIP and wrist joints are estimated to provide an initial accurate joint configuration; In the Prior Guided Joints Estimation stage (PG-stage), a dual-branch interaction network is employed to refine the positions of the remaining joints. Extensive experiments on two widely used benchmarks demonstrate that EHPE achieves state-of-the-arts performance. Code is available at https://github.com/SereinNout/EHPE.
Figures
Forward citations
Cited by 1 Pith paper
-
MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark
Multilingual LLMs show a reproducible 'Illusion of Cultural Alignment': they can be fluent in a language while lacking the culture's factual knowledge, and confidence, sampling, and retrieval do not fix it.
Reference graph
Works this paper leans on
-
[1]
Seungryul Baek, Kwang In Kim, and Tae-Kyun Kim. 2019. Pushing the envelope for rgb-based dense 3d hand pose estimation via neural rendering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 1067–1076
work page 2019
-
[2]
Yujun Cai, Liuhao Ge, Jianfei Cai, Nadia Magnenat Thalmann, and Junsong Yuan
-
[3]
Yujun Cai, Liuhao Ge, Jianfei Cai, and Junsong Yuan. 2018. Weakly-supervised 3d hand pose estimation from monocular rgb images. In Proceedings of the European conference on computer vision (ECCV) . 666–682
work page 2018
-
[4]
Jiayi Chen, Mi Yan, Jiazhao Zhang, Yinzhen Xu, Xiaolong Li, Yijia Weng, Li Yi, Shuran Song, and He Wang. 2023. Tracking and reconstructing hand object interactions from point cloud sequences in the wild. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 304–312
work page 2023
-
[5]
Xingyu Chen, Yufeng Liu, Yajiao Dong, Xiong Zhang, Chongyang Ma, Yanmin Xiong, Yuan Zhang, and Xiaoyan Guo. 2022. Mobrecon: Mobile-friendly hand mesh reconstruction from monocular image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 20544–20554
work page 2022
-
[6]
Xiaoming Deng, Dexin Zuo, Yinda Zhang, Zhaopeng Cui, Jian Cheng, Ping Tan, Liang Chang, Marc Pollefeys, Sean Fanello, and Hongan Wang. 2022. Recur- rent 3d hand pose estimation using cascaded pose-guided 3d alignments. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 1 (2022), 932–945
work page 2022
-
[7]
Liuhao Ge, Yujun Cai, Junwu Weng, and Junsong Yuan. 2018. Hand pointnet: 3d hand pose estimation using point sets. In Proceedings of the IEEE conference on computer vision and pattern recognition . 8417–8426
work page 2018
-
[8]
Liuhao Ge, Hui Liang, Junsong Yuan, and Daniel Thalmann. 2017. 3D Convo- lutional Neural Networks for Efficient and Robust Hand Pose Estimation from Single Depth Images. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). doi:10.1109/cvpr.2017.602
Show all 65 references
-
[9]
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2020. Generative adversarial networks. Commun. ACM 63, 11 (2020), 139–144
2020
-
[10]
Mustafa Haiderbhai, Sergio Ledesma, Sing Chun Lee, Matthias Seibold, Phillipp Fürnstahl, Nassir Navab, and Pascal Fallavollita. 2020. Pix2xray: Converting RGB images into X-rays using generative adversarial networks. International journal of computer assisted radiology and sur...
2020
-
[11]
Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. 2022. Key- point transformer: Solving joint identification in challenging hands and object interactions for accurate 3d pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...
2022
-
[12]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
2016
-
[13]
Yiming He and Wei Hu. 2021. 3D hand pose estimation via regularized graph representation learning. In CAAI International Conference on Artificial Intelligence. Springer, 540–552
2021
-
[14]
Guoguang Hua, Lihong Li, and Shiguang Liu. 2020. Multipath affinage stacked—hourglass networks for human pose estimation. Frontiers of Computer Science 14 (2020), 1–12
2020
-
[15]
Weiting Huang, Pengfei Ren, Jingyu Wang, Qi Qi, and Haifeng Sun. 2020. Awr: Adaptive weighting regression for 3d hand pose estimation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 11061–11068
2020
-
[16]
Changlong Jiang, Yang Xiao, Cunlin Wu, Mingyang Zhang, Jinghong Zheng, Zhiguo Cao, and Joey Tianyi Zhou. 2023. A2j-transformer: Anchor-to-joint transformer network for 3d interacting hand pose estimation from a single rgb image. In Proceedings of the IEEE/CVF conference on com...
2023
-
[17]
Jianping Jiang, Jiahe Li, Baowen Zhang, Xiaoming Deng, and Boxin Shi. 2024. Evhandpose: Event-based 3d hand pose estimation with sparse supervision. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)
2024
-
[18]
Evangelos Kazakos, Christophoros Nikou, and Ioannis A Kakadiaris. 2018. On the fusion of RGB and depth information for hand pose estimation. In 2018 25th IEEE International Conference on Image Processing (ICIP) . IEEE, 868–872
2018
-
[19]
Leyla Khaleghi, Alireza Sepas-Moghaddam, Joshua Marshall, and Ali Etemad
-
[20]
Kingma and Jimmy Ba
DiederikP. Kingma and Jimmy Ba. 2014. Adam: A Method for Stochastic Opti- mization. arXiv: Learning,arXiv: Learning (Dec 2014)
2014
-
[21]
Thomas N Kipf and Max Welling. 2016. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907 (2016)
2016 arXiv
-
[22]
Chuankun Li, Shuai Li, Yanbo Gao, Xiang Zhang, and Wanqing Li. 2021. A two-stream neural network for pose-based hand gesture recognition. IEEE Trans- actions on Cognitive and Developmental Systems 14, 4 (2021), 1594–1603
2021
-
[23]
Lijun Li, Li’an Zhuo, Bang Zhang, Liefeng Bo, and Chen Chen. 2023. DiffHand: End-to-End Hand Mesh Reconstruction via Diffusion Models. arXiv preprint arXiv:2305.13705 (2023). MM2025, October 27-0ctober 31,2025, Dublin, lreland Zheng et al
2023 arXiv
-
[24]
Mengcheng Li, Liang An, Hongwen Zhang, Lianpeng Wu, Feng Chen, Tao Yu, and Yebin Liu. 2022. Interacting attention graph for single image two-hand reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2761–2770
2022
-
[25]
Hui Liang, Junsong Yuan, Jun Lee, Liuhao Ge, and Daniel Thalmann. 2019. Hough Forest With Optimized Leaves for Global Hand Pose Estimation With Arbitrary Postures. IEEE Transactions on Cybernetics 49, 2 (Feb 2019), 527–541. doi:10.1109/ tcyb.2017.2779800
2019
-
[26]
Jun Lv, Wenqiang Xu, Lixin Yang, Sucheng Qian, Chongzhao Mao, and Cewu Lu
-
[27]
Hao Meng, Sheng Jin, Wentao Liu, Chen Qian, Mengxiang Lin, Wanli Ouyang, and Ping Luo. 2022. 3d interacting hand pose estimation by hand de-occlusion and removal. In European Conference on Computer Vision . Springer, 380–397
2022
-
[28]
Gyeongsik Moon, Ju Yong Chang, and Kyoung Mu Lee. 2018. V2v-posenet: Voxel- to-voxel prediction network for accurate 3d hand and human pose estimation from a single depth map. In Proceedings of the IEEE conference on computer vision and pattern Recognition. 5079–5088
2018
-
[29]
Gyeongsik Moon, Shoou-I Yu, He Wen, Takaaki Shiratori, and Kyoung Mu Lee
-
[30]
Franziska Mueller, Florian Bernard, Oleksandr Sotnychenko, Dushyant Mehta, Srinath Sridhar, Dan Casas, and Christian Theobalt. 2018. Ganerated hands for real-time 3d hand tracking from monocular rgb. In Proceedings of the IEEE conference on computer vision and pattern recognit...
2018
-
[31]
Franziska Mueller, Dushyant Mehta, Oleksandr Sotnychenko, Srinath Sridhar, Dan Casas, and Christian Theobalt. 2017. Real-time hand tracking under occlusion from an egocentric rgb-d sensor. InProceedings of the IEEE international conference on computer vision. 1154–1163
2017
-
[32]
JoonKyu Park, Daniel Sungho Jung, Gyeongsik Moon, and Kyoung Mu Lee
-
[33]
6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image
Interhand2. 6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image. InComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XX 16 . Springer, 548–564
2020
-
[34]
Jinwei Ren, Jianke Zhu, and Jialiang Zhang. 2023. End-to-end weakly-supervised single-stage multiple 3D hand mesh reconstruction from a single RGB image. Computer Vision and Image Understanding 232 (2023), 103706
2023
-
[35]
Pengfei Ren, Chao Wen, Xiaozheng Zheng, Zhou Xue, Haifeng Sun, Qi Qi, Jingyu Wang, and Jianxin Liao. 2023. Decoupled iterative refinement framework for interacting hands reconstruction from a single rgb image. In Proceedings of the IEEE/CVF international conference on computer...
2023
-
[36]
Javier Romero, Dimitrios Tzionas, and Michael J. Black. 2017. Embodied Hands: Modeling and Capturing Hands and Bodies Together. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia) 36, 6 (Nov. 2017)
2017
-
[37]
Adrian Spurr, Umar Iqbal, Pavlo Molchanov, Otmar Hilliges, and Jan Kautz. 2020. Weakly supervised 3d hand pose estimation via biomechanical constraints. In European conference on computer vision . Springer, 211–228
2020
-
[38]
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. 2024. Reconstructing hands in 3d with transform- ers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 9826–9836
2024
-
[39]
Xiao Tang, Tianyu Wang, and Chi-Wing Fu. 2021. Towards accurate alignment in real-time 3d hand-mesh reconstruction. In Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision . 11698–11707
2021
-
[40]
Denis Tome, Chris Russell, and Lourdes Agapito. 2017. Lifting from the deep: Convolutional 3d pose estimation from a single image. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2500–2509
2017
-
[41]
Jonathan Tompson, Murphy Stein, Yann Lecun, and Ken Perlin. 2014. Real-time continuous pose recovery of human hands using convolutional networks. ACM Transactions on Graphics (ToG) 33, 5 (2014), 1–10
2014
-
[42]
Zhigang Tu, Zhisheng Huang, Yujin Chen, Di Kang, Linchao Bao, Bisheng Yang, and Junsong Yuan. 2023. Consistent 3d hand reconstruction in video via self- supervised learning. IEEE Transactions on Pattern Analysis and Machine Intelli- gence 45, 8 (2023), 9469–9485
2023
-
[43]
Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang. 2019. Deep high-resolution representation learning for human pose estimation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5693–5703
2019
-
[44]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)
2017
-
[45]
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. 2017. Graph attention networks. arXiv preprint arXiv:1710.10903 (2017)
2017 arXiv
-
[46]
Chengde Wan, Thomas Probst, Luc Van Gool, and Angela Yao. 2018. Dense 3d regression for hand pose estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5147–5156
2018
-
[47]
Tingyu Wang, Zihao Yang, Quan Chen, Yaoqi Sun, and Chenggang Yan. 2024. Rethinking pooling for multi-granularity features in aerial-view geo-localization. IEEE Signal Processing Letters (2024)
2024
-
[48]
Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel, and Anurag Ranjan. 2023. FastViT: A fast hybrid vision transformer using structural reparam- eterization. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 5785–5795
2023
-
[49]
Huan Yao, Changxing Ding, Xuanda Xu, and Zhifeng Lin. 2024. Decoupling Heterogeneous Features for Robust 3D Interacting Hand Poses Estimation. In Proceedings of the 32nd ACM International Conference on Multimedia . 5338–5346
2024
-
[50]
Baowen Zhang, Yangang Wang, Xiaoming Deng, Yinda Zhang, Ping Tan, Cuixia Ma, and Hongan Wang. 2021. Interacting Two-Hand 3D Pose and Shape Recon- struction From Single Color Image. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV). doi:10.1109/iccv48922.2021.01116
2021
-
[51]
Pengfei Zhang and Deying Kong. 2024. Handformer2T: A Lightweight Regression- based Model for Interacting Hands Pose Estimation from A Single RGB Image. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 6248–6257
2024
-
[52]
Xiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang, and Wen Zheng. 2019. End- to-end hand mesh recovery from a monocular rgb image. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 2354–2364
2019
-
[53]
Ying Wu, John Lin, and Thomas S Huang. 2005. Analyzing and capturing articu- lated hand motion in image sequences. IEEE transactions on pattern analysis and machine intelligence 27, 12 (2005), 1910–1922
2005
-
[54]
Bolun Zheng, Yaowu Chen, Xiang Tian, Fan Zhou, and Xuesong Liu. 2019. Implicit dual-domain convolutional network for robust color image compression artifact reduction. IEEE Transactions on Circuits and Systems for Video Technology 30, 11 (2019), 3982–3994
2019
-
[55]
Bolun Zheng, Haoran Li, Quan Chen, Tingyu Wang, Xiaofei Zhou, Zhenghui Hu, and Chenggang Yan. 2024. Quad bayer joint demosaicing and denoising based on dual encoder network with joint residual learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38....
2024
-
[56]
Zhishan Zhou, Shihao Zhou, Zhi Lv, Minqiang Zou, Yao Tang, and Jiajun Liang
-
[57]
Christian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan Russell, Max Argus, and Thomas Brox. 2019. Freihand: A dataset for markerless capture of hand pose and shape from single rgb images. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 813–822
2019
-
[58]
Bolun Zheng, Quan Chen, Shanxin Yuan, Xiaofei Zhou, Hua Zhang, Jiyong Zhang, Chenggang Yan, and Gregory Slabaugh. 2022. Constrained predictive filters for single image bokeh rendering. IEEE Transactions on Computational Imaging 8 (2022), 346–357
2022
-
[64]
Binghui Zuo, Zimeng Zhao, Wenqian Sun, Wei Xie, Zhou Xue, and Yangang Wang
-
[65]
In Proceedings of the IEEE/CVF international conference on computer vision
Reconstructing interacting hands with interaction prior from monocular images. In Proceedings of the IEEE/CVF international conference on computer vision . 9054–9064. Received 23 April 2025; revised 19 June 2025; accepted 4 July 2025
2025
-
[2020]
IEEE transactions on pattern analysis and machine intelligence 43, 11 (2020), 3739–3753
3D hand pose estimation using synthetic data and weakly labeled RGB images. IEEE transactions on pattern analysis and machine intelligence 43, 11 (2020), 3739–3753
2020
-
[2021]
arXiv preprint arXiv:2102.09244 (2021)
Handtailor: Towards high-precision monocular 3d hand recovery. arXiv preprint arXiv:2102.09244 (2021)
2021 arXiv
-
[2022]
IEEE Transactions on Artificial Intelligence 4, 4 (2022), 896–909
Multiview video-based 3-d hand pose estimation. IEEE Transactions on Artificial Intelligence 4, 4 (2022), 896–909
2022
-
[2023]
In Proceedings of the IEEE/CVF International Conference on Computer Vision
Extract-and-adaptation network for 3D interacting hand mesh recovery. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 4200– 4209
-
[2024]
InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
A Simple Baseline for Efficient Hand Mesh Reconstruction. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1367–1376
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.