REVIEW 2 major objections 5 minor 300 references
Foundation models help hand–object interaction only when we name the prior they inject and the uncertainty it reduces—not when we merely say a method “uses a large model.”
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 08:48 UTC pith:FNALSTCS
load-bearing objection Solid first map of foundation-model priors for HOI; the taxonomy is the real contribution and the §1 boundary holds up well enough. the 2 major comments →
Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper’s central claim is that it is the first systematic review of foundation-model priors for hand–object interaction: methods should be classified by which cross-domain prior they introduce, where that prior enters a common HOI pipeline, and which of five uncertainties (shape, spatial, physical, semantic, dynamic) it reduces—via a taxonomy of eight sub-priors in geometric, semantic, and visual families spanning six HOI tasks plus embodied transfer to robots.
What carries the argument
The eight-sub-prior taxonomy (shape retrieval, shape reconstruction, spatial reconstruction; semantic grounding, language reasoning; visual representation, image generation, video generation), paired with a shared injection vocabulary—prior source, injected representation, injection operator, target, task, and residual limitation—that turns “uses a large model” into a comparable account of knowledge flow.
Load-bearing premise
The whole map stands on a sharp definition: only an explicitly named, large-scale general-purpose pretrained model that contributes cross-domain knowledge counts as a foundation prior, while in-domain assets and task-specific initializations do not unless such a model retrieves, generates, or scores them.
What would settle it
Re-label every method in the survey’s main table under the paper’s own definition and check whether primary-versus-auxiliary prior tags and the claim of being first systematic still hold; if many borderline systems flip families or if an earlier survey already organized HOI the same way by prior source, injection site, and uncertainty reduced, the central framing fails.
If this is right
- New HOI papers would report prior source, injection operator, and which uncertainty is targeted instead of only architecture and dataset names.
- Benchmarks would pair geometry scores with contact, physical plausibility, and task or object-state success under one protocol.
- Multi-prior systems would need confidence, routing, and conflict handling when retrieval, reconstruction, grounding, and generation disagree.
- Robot learning from human video would be judged on preserved contact, intent, and object-state change—not trajectory similarity alone.
- Long-horizon egocentric HOI would aim for world-frame camera, hand, object, and contact state across grasp, use, release, and re-grasp.
Where Pith is reading between the lines
- If the injection taxonomy sticks, ablation studies should disable one prior family at a time and measure the matching uncertainty, not only end-task error.
- The same prior-versus-uncertainty grid could grade non-HOI manipulation stacks (tool use, bimanual assembly) without rewriting the task list.
- Live repositories tied to this taxonomy will matter only if each new method is forced to declare primary prior and mitigated uncertainty at ingest.
- Embodied memory after grasp occlusion is a natural stress test: pre-contact HOI state must stay queryable when the hand hides the object.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey reviews hand-object interaction (HOI) methods that exploit foundation-model priors. It organizes six HOI tasks (three reconstruction, three generation) and proposes a taxonomy of eight sub-priors in geometric, semantic, and visual families. A stated inclusion rule treats a method as foundation-model-prior only when an explicitly identified large-scale general-purpose pretrained model contributes cross-domain knowledge. The paper maps how priors are represented and injected into HOI pipelines (Table 3; Figs. 3–6), separates non-foundation baselines (§2), reviews HOI-derived embodied transfer to robot learning (§6), and summarizes datasets, metrics, and open challenges, with a live repository for ongoing coverage.
Significance. If the taxonomy and inclusion boundary hold, the paper fills a clear gap relative to prior HOI surveys (Table 1): it reframes a fragmented “uses large models” literature by prior source, injection operator, and residual uncertainty (shape, spatial, physical, semantic, dynamic). The shared pipeline abstraction, primary/auxiliary labeling in Table 3, and embodied-transfer routes (pretraining, skill transfer, data engines) are practically useful for both vision and robotics readers. The live repository and explicit non-foundation baseline section are concrete strengths that support cumulative work beyond a static review.
major comments (2)
- [§1, Table 3] §1 and Table 3: The inclusion rule is clear, but the operational rule for assigning the primary family (and thus table grouping and Unc.↓ tags) in multi-prior systems is under-specified. EasyHOI, ArtHOI, MCC-HO, Jiang et al., CHOIR, and GraG combine geometric, semantic, and visual components; “contributes most directly to the HOI solution” is not a reproducible criterion. Please add a short decision procedure (e.g., which module produces the task-defining output, ablation/ablation-proxy, or dependency order) and apply it consistently, or mark ambiguous cases explicitly. Without this, the eight-sub-prior taxonomy remains useful as a map but weaker as a systematic classification.
- [§7.2, §8.2] §7.2 and §8.2: The survey correctly criticizes geometry-dominated evaluation and calls for interaction correctness, yet the manuscript does not tabulate which reviewed methods report contact/physical/embodied metrics versus only MPJPE/Chamfer/FID. A compact coverage matrix (method × metric family, or benchmark × reported interaction metrics) would make the “blind spots” claim evidence-based and better support the future-benchmark recommendations. This is load-bearing for the evaluation narrative, not merely a wish-list item.
minor comments (5)
- [Figure 7] Figure 7 labels routes as “Sec. 4.2 / 4.3 / 4.4” while the body places embodied transfer in §6; renumber for consistency.
- [Table 3] Table 3 mixes venue tags (CVPR 2026, arXiv 2026, etc.) for very recent work; a footnote on inclusion cutoff date and how preprints versus accepted versions are handled would help readers and the live repo stay aligned.
- [§3.2] §3.2 argues shape retrieval is sparse as a primary prior; consider one sentence on when retrieval should be preferred over reconstruction (asset coverage, topology stability) to balance the critical tone.
- [§2.3.3, §5.4] Occasional notation/typo issues (e.g., “SLAM [53 COPSfM [54]” in §2.3.3; “Pl¨ucker” encoding) should be cleaned in copy-edit.
- [Abstract, §1] Keywords and abstract claim “first systematic review”; keep that claim tied to the prior-injection lens (as in Table 1) so it is not read as first HOI survey overall.
Circularity Check
No circularity: a bibliographic taxonomy survey with no fitted predictions or self-justifying derivation chain.
full rationale
This paper is a literature survey. It does not derive a quantitative prediction, fit parameters to data and re-label them as forecasts, or invoke a uniqueness theorem that collapses to author-only prior work. Its central contribution is an organizational cut (foundation-model prior vs. non-foundation-prior; eight sub-priors; six HOI tasks plus embodied transfer) applied to external methods. That cut is definitional taxonomy, not a circular derivation: inclusion criteria are stated explicitly in §1, methods are tabulated with primary/auxiliary prior tags drawn from named external foundation models (DUSt3R, SAM, DINOv2, Stable Diffusion, etc.), and residual uncertainties are framed as open evaluation gaps rather than results forced by the inputs. Author-adjacent systems (e.g., GeoHand, HandOS) and the live repository appear as map entries or infrastructure, which is normal survey practice and does not make the taxonomy or “first systematic review” claim true by construction. Under the stated circularity criteria there is nothing to reduce: no Eq. X = Eq. Y by fit, no load-bearing self-cited uniqueness, no ansatz smuggled in as a theorem. Score 0 is the honest finding.
Axiom & Free-Parameter Ledger
axioms (4)
- ad hoc to paper A method is a foundation-model-prior method only if an explicitly identified large-scale general-purpose pretrained model contributes cross-domain knowledge via predictions, representations, parameters, adaptation, or distillation (§1).
- domain assumption HOI difficulty under visual uncertainty factors into five residual uncertainties—shape, spatial, physical, semantic, and dynamic—that foundation priors can target (Intro; Fig. 3).
- domain assumption Six tasks (R1–R3, G1–G3) plus embodied transfer suffice to cover the foundation-era HOI literature for this review (§2.1, §6).
- domain assumption Cited concurrent methods’ reported mechanisms can be trusted enough to assign primary/auxiliary prior tags without re-running experiments (Table 3; §§3–6).
invented entities (2)
-
Eight foundation-model sub-priors in three families (G-Ret, G-Rec, G-Spa; S-Gnd, S-Lng; V-Rep, V-Img, V-Vid)
no independent evidence
-
Shared injection-operator vocabulary (retrieve/align/regularize, initialize, token fusion, region conditioning, condition/fuse, weight init/fine-tune, adapter/ControlNet, score-guided regularization,
no independent evidence
read the original abstract
Hand-object interaction (HOI) modeling remains challenging because it requires joint reasoning about hand articulation, object geometry, contact, semantics, and dynamics under severe visual uncertainty. Foundation models introduce transferable prior knowledge learned from large-scale cross-domain data, offering new ways to address these challenges beyond task-specific data and models. However, the rapidly growing literature remains fragmented, and existing studies typically describe these methods simply as ``using large models'' without systematically characterizing what knowledge is introduced, where it enters the HOI pipeline, or which HOI uncertainty it helps reduce. This survey presents the first systematic review of foundation-model priors for HOI. We organize the literature into six HOI tasks spanning reconstruction and generation. More importantly, we establish a taxonomy of eight foundation-model sub-priors grouped into geometric, semantic, and visual families. Geometric priors encompass shape retrieval, shape reconstruction, and spatial reconstruction; semantic priors include semantic grounding and language reasoning; and visual priors cover visual representation, image generation, and video generation. Based on this taxonomy, we systematically analyze how different priors are represented, injected, and adapted across HOI pipelines and tasks. Beyond how foundation models empower HOI, we further examine how HOI-derived knowledge is used in robot learning, including human-data pretraining, human-to-robot skill transfer, and HOI-to-robot data generation. Finally, we summarize datasets and evaluation protocols, and discuss limitations and future directions toward more generalizable HOI systems. To support long-term progress, we curate a live repository that continuously aggregates emerging methods and benchmarks.
Figures
Reference graph
Works this paper leans on
-
[1]
H+O: unified egocentric recognition of 3D hand- object poses and interactions
Bugra Tekin, Federica Bogo, and Marc Pollefeys. H+O: unified egocentric recognition of 3D hand- object poses and interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019. doi: 10.1109/CVPR.2019.00464
arXiv 2019
-
[2]
Learning joint reconstruction of hands and manipulated objects
Yana Hasson, Gul Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J Black, Ivan Laptev, and Cordelia Schmid. Learning joint reconstruction of hands and manipulated objects. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019
2019
-
[3]
CPF: Learning a contact potential field to model the hand-object interaction
Lixin Yang, Xinyu Zhan, Kailin Li, Wenqiang Xu, Jiefeng Li, and Cewu Lu. CPF: Learning a contact potential field to model the hand-object interaction. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021
2021
-
[4]
HOLD: Category-agnostic 3D re- construction of interacting hands and objects from video
Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen, Muhammed Kocabas, Michael J Black, and Otmar Hilliges. HOLD: Category-agnostic 3D re- construction of interacting hands and objects from video. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024. 23 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Gener...
2024
-
[5]
What’s in your hands? 3D reconstruction of generic objects in hands
Yufei Ye, Abhinav Gupta, and Shubham Tulsiani. What’s in your hands? 3D reconstruction of generic objects in hands. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[6]
Grasping field: Learning implicit representations for human grasps
Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael J Black, Krikamol Muandet, and Siyu Tang. Grasping field: Learning implicit representations for human grasps. InProc. Int. Conf. 3D Vis. (3DV), 2020
2020
-
[7]
DUSt3R: Geomet- ric 3D vision made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. DUSt3R: Geomet- ric 3D vision made easy. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[8]
Contactpose: A dataset of grasps with object con- tact and hand pose
Samarth Brahmbhatt, Chengcheng Tang, Christo- pher D Twigg, Charles C Kemp, and James Hays. Contactpose: A dataset of grasps with object con- tact and hand pose. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020
2020
-
[9]
S2contact: Graph-based network for 3D hand-object contact estimation with semi-supervised learning
Tze Ho Elden Tse, Zhongqun Zhang, Kwang In Kim, Ales Leonardis, Feng Zheng, and Hyung Jin Chang. S2contact: Graph-based network for 3D hand-object contact estimation with semi-supervised learning. In Proc. Eur. Conf. Comput. Vis. (ECCV), 2022. doi: 10.1007/978-3-031-19769-7\ 33
-
[10]
Contactopt: Optimizing contact to improve grasps
Patrick Grady, Chengcheng Tang, Christopher D Twigg, Minh V o, Samarth Brahmbhatt, and Charles C Kemp. Contactopt: Optimizing contact to improve grasps. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2021
2021
-
[11]
Deep- SimHO: Stable pose estimation for hand-object in- teraction via physics simulation
Rong Wang, Wei Mao, and Hongdong Li. Deep- SimHO: Stable pose estimation for hand-object in- teraction via physics simulation. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023
2023
-
[12]
D- grasp: Physically plausible dynamic grasp synthesis for hand-object interactions
Sammy Christen, Muhammed Kocabas, Emre Aksan, Jemin Hwangbo, Jie Song, and Otmar Hilliges. D- grasp: Physically plausible dynamic grasp synthesis for hand-object interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[13]
Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jian- wei Yang, Hang Su, et al. Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024
2024
-
[14]
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023
2023
-
[15]
SemGrasp: Semantic grasp generation via language aligned discretization
Kailin Li, Jingbo Wang, Lixin Yang, Cewu Lu, and Bo Dai. SemGrasp: Semantic grasp generation via language aligned discretization. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024
2024
-
[16]
Text2grasp: Synthesis of grasps by text prompts for object grasping parts
Xiaoyun Chang and Yi Sun. Text2grasp: Synthesis of grasps by text prompts for object grasping parts. In Proc. Int. Symp. Neural Netw. (ISNN), 2025
2025
-
[17]
Towards unconstrained joint hand-object re- construction from RGB videos
Yana Hasson, G¨ul Varol, Cordelia Schmid, and Ivan Laptev. Towards unconstrained joint hand-object re- construction from RGB videos. InProc. Int. Conf. 3D Vis. (3DV), 2021
2021
-
[18]
Ayce Idil Aytekin, Xu Chen, Zhengyang Shen, Thabo Beeler, Helge Rhodin, Rishabh Dabral, and Christian Theobalt. Grasp in gaussians: Fast monocular recon- struction of dynamic hand-object interactions.arXiv preprint arXiv:2604.12929, 2026
Pith/arXiv arXiv 2026
-
[19]
EasyHOI: Unleashing the power of large models for reconstructing hand-object inter- actions in the wild
Yumeng Liu, Xiaoxiao Long, Zemin Yang, Yuan Liu, Marc Habermann, Christian Theobalt, Yuexin Ma, and Wenping Wang. EasyHOI: Unleashing the power of large models for reconstructing hand-object inter- actions in the wild. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[20]
MagicHOI: Leveraging 3D priors for accurate hand-object recon- struction from short monocular video clips
Shibo Wang, Haonan He, Maria Parelli, Christoph Gebhardt, Zicong Fan, and Jie Song. MagicHOI: Leveraging 3D priors for accurate hand-object recon- struction from short monocular video clips. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025
2025
-
[21]
Diffusion-guided reconstruction of everyday hand-object interaction clips
Yufei Ye, Poorvi Hebbar, Abhinav Gupta, and Shub- ham Tulsiani. Diffusion-guided reconstruction of everyday hand-object interaction clips. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023
2023
-
[22]
Hand-object interaction image gen- eration
Hezhen Hu, Weilun Wang, Wengang Zhou, and Houqiang Li. Hand-object interaction image gen- eration. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 35, 2022
2022
-
[23]
HOIDiffusion: Generating realistic 3D hand-object interaction data
Mengqi Zhang, Yang Fu, Zheng Ding, Sifei Liu, Zhuowen Tu, and Xiaolong Wang. HOIDiffusion: Generating realistic 3D hand-object interaction data. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2024
2024
-
[24]
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127, 2023. 24 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction...
Pith/arXiv arXiv 2023
-
[25]
Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Efficient 3D mesh generation from a single image with sparse- view large reconstruction models.arXiv preprint arXiv:2404.07191, 2024
Pith/arXiv arXiv 2024
-
[26]
OpenShape: Scaling up 3D shape representation towards open-world understanding
Minghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu, Xuanlin Li, Shizhong Han, Hong Cai, Fatih Porikli, and Hao Su. OpenShape: Scaling up 3D shape representation towards open-world understanding. In Adv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023
2023
-
[27]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. InAdv. Neu- ral Inf. Process. Syst. (NeurIPS), volume 36, 2023
2023
-
[28]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A versatile vision-language model for understanding, localiza- tion, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023
Pith/arXiv arXiv 2023
-
[29]
DINOv2: Learning robust vi- sual features without supervision.arXiv preprint arXiv:2304.07193, 2023
Maxime Oquab, Timoth´ee Darcet, Th´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fer- nandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. DINOv2: Learning robust vi- sual features without supervision.arXiv preprint arXiv:2304.07193, 2023
Pith/arXiv arXiv 2023
-
[30]
High- resolution image synthesis with latent diffusion mod- els
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High- resolution image synthesis with latent diffusion mod- els. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[31]
Objaverse: A universe of annotated 3D objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3D objects. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2023
2023
-
[32]
Learning trans- ferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sas- try, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning trans- ferable visual models from natural language supervi- sion. InProc. Int. Conf. Mach. Learn. (ICML), 2021
2021
-
[33]
Reconstructing hand-held objects in 3D from images and videos
Jane Wu, Georgios Pavlakos, Georgia Gkioxari, and Jitendra Malik. Reconstructing hand-held objects in 3D from images and videos. InProc. Int. Conf. 3D Vis. (3DV), 2026
2026
-
[34]
Ghost: Fast category- agnostic hand-object interaction reconstruction from RGB videos using gaussian splatting
Ahmed Tawfik Aboukhadra, Marcel Rogge, Nadia Robertini, Abdalla Arafa, Jameel Malik, Ahmed El- hayek, and Didier Stricker. Ghost: Fast category- agnostic hand-object interaction reconstruction from RGB videos using gaussian splatting. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026
2026
-
[35]
Hand-held object reconstruction from RGB video with dynamic interaction
Shijian Jiang, Qi Ye, Rengan Xie, Yuchi Huo, and Jim- ing Chen. Hand-held object reconstruction from RGB video with dynamic interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[36]
Reconstructing hands in 3D with transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstructing hands in 3D with transformers. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[37]
Wilor: End- to-end 3D hand localization and reconstruction in-the- wild
Rolandos Alexandros Potamias, Jinglei Zhang, Jiankang Deng, and Stefanos Zafeiriou. Wilor: End- to-end 3D hand localization and reconstruction in-the- wild. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[38]
ViTPose: Simple vision transformer baselines for human pose estimation
Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. ViTPose: Simple vision transformer baselines for human pose estimation. InAdv. Neural Inf. Process. Syst. (NeurIPS), 2022
2022
-
[39]
Latent action pretraining from videos
Seonghyeon Ye, Joel Jang, Byeongguk Jeon, Se June Joo, Jianwei Yang, Baolin Peng, Ajay Mandlekar, Reuben Tan, Yu-Wei Chao, Bill Yuchen Lin, et al. Latent action pretraining from videos. InProc. Int. Conf. Learn. Represent. (ICLR), 2025
2025
-
[40]
Ruihan Yang, Qinxi Yu, Yecheng Wu, Rui Yan, Borui Li, An-Chieh Cheng, Xueyan Zou, Yunhao Fang, Xuxin Cheng, Ri-Zhao Qiu, et al. Egovla: Learning vision-language-action models from egocentric hu- man videos.arXiv preprint arXiv:2507.12440, 2025
Pith/arXiv arXiv 2025
-
[41]
Dexmv: Imitation learning for dexterous manipula- tion from human videos
Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang, Ruihan Yang, Yang Fu, and Xiaolong Wang. Dexmv: Imitation learning for dexterous manipula- tion from human videos. InProc. Eur. Conf. Comput. Vis. (ECCV), 2022
2022
-
[42]
Yangcen Liu, Shuo Cheng, Xinchen Yin, Woo Chul Shin, Alfred Cueva, Yiran Yang, Zhenyang Chen, Chuye Zhang, and Danfei Xu. Egoengine: From ego- centric human videos to high-fidelity dexterous robot demonstrations.arXiv preprint arXiv:2606.12604, 2026
Pith/arXiv arXiv 2026
-
[43]
Takehiko Ohkawa, Ryosuke Furuta, and Yoichi Sato. Efficient annotation and learning for 3D hand pose 25 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer estimation: A survey.Int. J. Comput. Vis., 131(12), 2023
2023
-
[44]
A survey of deep learning methods and datasets for hand pose estimation from hand-object interaction images.Comput
Taeyun Woo, Wonjung Park, Woohyun Jeong, and Jinah Park. A survey of deep learning methods and datasets for hand pose estimation from hand-object interaction images.Comput. Graph., 116, 2023
2023
-
[45]
Advances in vision-based deep learning methods for interacting hands reconstruction: A survey.Comput
Yu Miao and Yue Liu. Advances in vision-based deep learning methods for interacting hands reconstruction: A survey.Comput. Graph., 124, 2024
2024
-
[46]
An overview of learning-based dexterous grasping: recent advances and future directions.Artif
Xu Song, Yongyao Li, Yunfan Zhang, Yufei Liu, and Lei Jiang. An overview of learning-based dexterous grasping: recent advances and future directions.Artif. Intell. Rev., 58(10), 2025
2025
-
[47]
Arthoi: Taming foundation models for monocular 4D reconstruction of hand-articulated- object interactions
Zikai Wang, Zhilu Zhang, Yiqing Wang, Hui Li, and Wangmeng Zuo. Arthoi: Taming foundation models for monocular 4D reconstruction of hand-articulated- object interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026
2026
-
[48]
Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: modeling and capturing hands and bodies together.ACM Trans. Graph., 36 (6), 2017. doi: 10.1145/3130800.3130883
arXiv 2017
-
[49]
Nimble: a non-rigid hand model with bones and muscles.ACM Trans
Yuwei Li, Longwen Zhang, Zesong Qiu, Yingwenqi Jiang, Nianyi Li, Yuexin Ma, Yuyao Zhang, Lan Xu, and Jingyi Yu. Nimble: a non-rigid hand model with bones and muscles.ACM Trans. Graph., 41(4), 2022
2022
-
[50]
HOnnotate: A method for 3D an- notation of hand and object poses
Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. HOnnotate: A method for 3D an- notation of hand and object poses. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020
2020
-
[51]
AlignSDF: Pose-aligned signed distance fields for hand-object reconstruction
Zerui Chen, Yana Hasson, Cordelia Schmid, and Ivan Laptev. AlignSDF: Pose-aligned signed distance fields for hand-object reconstruction. InProc. Eur. Conf. Comput. Vis. (ECCV), 2022
2022
-
[52]
HandNeRF: Learn- ing to reconstruct hand-object interaction scene from a single RGB image
Hongsuk Choi, Nikhil Chavan-Dafle, Jiacheng Yuan, V olkan Isler, and Hyunsoo Park. HandNeRF: Learn- ing to reconstruct hand-object interaction scene from a single RGB image. InProc. IEEE Int. Conf. Robot. Autom. (ICRA), 2024
2024
-
[53]
Simultaneous localization and mapping: part i.IEEE Robot
Hugh Durrant-Whyte and Tim Bailey. Simultaneous localization and mapping: part i.IEEE Robot. Autom. Mag., 13(2), 2006
2006
-
[54]
Seitz, and Richard Szeliski
Noah Snavely, Steven M. Seitz, and Richard Szeliski. Photo tourism: Exploring photo collections in 3D. In ACM SIGGRAPH 2006 Papers. ACM, 2006
2006
-
[55]
LatentHOI: On the generalizable hand object motion generation with latent hand diffusion
Muchen Li, Sammy Christen, Chengde Wan, Yu- jun Cai, Renjie Liao, Leonid Sigal, and Shugao Ma. LatentHOI: On the generalizable hand object motion generation with latent hand diffusion. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[56]
3D hand pose estimation in everyday egocentric images
Aditya Prakash, Ruisen Tu, Matthew Chang, and Saurabh Gupta. 3D hand pose estimation in everyday egocentric images. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024
2024
-
[57]
DDF-HO: Hand-held object recon- struction via conditional directed distance field
Chenyangguang Zhang, Yan Di, Ruida Zhang, Guangyao Zhai, Fabian Manhardt, Federico Tombari, and Xiangyang Ji. DDF-HO: Hand-held object recon- struction via conditional directed distance field. In Adv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023
2023
-
[58]
Ncrf: neural contact radiance fields for free-viewpoint rendering of hand-object interaction
Zhongqun Zhang, Jifei Song, Eduardo P´erez-Pellitero, Yiren Zhou, Hyung Jin Chang, and Ale ˇs Leonardis. Ncrf: neural contact radiance fields for free-viewpoint rendering of hand-object interaction. InProc. Int. Conf. 3D Vis. (3DV), 2024
2024
-
[59]
Gan- hand: Predicting human grasp affordances in multi- object scenes
Enric Corona, Albert Pumarola, Guillem Alenya, Francesc Moreno-Noguer, and Gr´egory Rogez. Gan- hand: Predicting human grasp affordances in multi- object scenes. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020
2020
-
[60]
Artigrasp: Physically plausible synthesis of bi-manual dexterous grasping and articulation
Hui Zhang, Sammy Christen, Zicong Fan, Luocheng Zheng, Jemin Hwangbo, Jie Song, and Otmar Hilliges. Artigrasp: Physically plausible synthesis of bi-manual dexterous grasping and articulation. InProc. Int. Conf. 3D Vis. (3DV), 2024
2024
-
[61]
Deepsdf: Learning continuous signed distance functions for shape representation
Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2019
2019
-
[62]
A skeleton-driven neural occupancy representation for articulated hands
Korrawe Karunratanakul, Adrian Spurr, Zicong Fan, Otmar Hilliges, and Siyu Tang. A skeleton-driven neural occupancy representation for articulated hands. InProc. Int. Conf. 3D Vis. (3DV), 2021
2021
-
[63]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020. doi: 10.1007/978-3-030-58452-8 \ 24. 26 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and ...
-
[64]
3D gaussian splat- ting for real-time radiance field rendering.ACM Trans
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk¨uhler, and George Drettakis. 3D gaussian splat- ting for real-time radiance field rendering.ACM Trans. Graph., 42(4), 2023
2023
-
[65]
Affordpose: A large-scale dataset of hand- object interactions with affordance-driven hand pose
Juntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu, and Jian Liu. Affordpose: A large-scale dataset of hand- object interactions with affordance-driven hand pose. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023
2023
-
[66]
A versatile and differentiable hand-object interaction representation
Th´eo Morales, Omid Taheri, and Gerard Lacey. A versatile and differentiable hand-object interaction representation. InProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), 2025
2025
-
[67]
Black, and Dima Damen
Siddhant Bansal, Zhifan Zhu, Shashank Tripathi, Ji- ahe Zhao, Michael J. Black, and Dima Damen. To- wards in-the-wild egocentric 3D hand-object pose estimation. InProc. Eur. Conf. Comput. Vis. (ECCV), 2026
2026
-
[68]
HOI4D: A 4D egocentric dataset for category-level human-object interaction
Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4D egocentric dataset for category-level human-object interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2022
2022
-
[69]
Arctic: A dataset for dexterous bimanual hand-object manipulation
Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J Black, and Otmar Hilliges. Arctic: A dataset for dexterous bimanual hand-object manipulation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023
2023
-
[70]
Handoccnet: Occlusion-robust 3D hand mesh estimation network
JoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hong- suk Choi, and Kyoung Mu Lee. Handoccnet: Occlusion-robust 3D hand mesh estimation network. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2022
2022
-
[71]
Mobrecon: Mobile-friendly hand mesh reconstruction from monocular image
Xingyu Chen, Yufeng Liu, Yajiao Dong, Xiong Zhang, Chongyang Ma, Yanmin Xiong, Yuan Zhang, and Xiaoyan Guo. Mobrecon: Mobile-friendly hand mesh reconstruction from monocular image. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[72]
A simple baseline for ef- ficient hand mesh reconstruction
Zhishan Zhou, Shihao Zhou, Zhi Lv, Minqiang Zou, Yao Tang, and Jiajun Liang. A simple baseline for ef- ficient hand mesh reconstruction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[73]
Model-based 3D hand reconstruction via self- supervised learning
Yujin Chen, Zhigang Tu, Di Kang, Linchao Bao, Ying Zhang, Xuefei Zhe, Ruizhi Chen, and Junsong Yuan. Model-based 3D hand reconstruction via self- supervised learning. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2021
2021
-
[74]
Keypoint fusion for RGB-D based 3D hand pose estimation
Xingyu Liu, Pengfei Ren, Yuanyuan Gao, Jingyu Wang, Haifeng Sun, Qi Qi, Zirui Zhuang, and Jianxin Liao. Keypoint fusion for RGB-D based 3D hand pose estimation. InProc. AAAI Conf. Artif. Intell., 2024
2024
-
[75]
Hope-net: A graph-based model for hand-object pose estimation
Bardia Doosti, Shujon Naha, Majid Mirbagheri, and David J Crandall. Hope-net: A graph-based model for hand-object pose estimation. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020
2020
-
[76]
Keypoint transformer: Solv- ing joint identification in challenging hands and ob- ject interactions for accurate 3D pose estimation
Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Keypoint transformer: Solv- ing joint identification in challenging hands and ob- ject interactions for accurate 3D pose estimation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[77]
Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision
Ahmed Tawfik Aboukhadra, Jameel Malik, Ahmed El- hayek, Nadia Robertini, and Didier Stricker. Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision. InProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), 2023
2023
-
[78]
HOISDF: Constraining 3D hand- object pose estimation with global signed distance fields
Haozhe Qi, Chen Zhao, Mathieu Salzmann, and Alexander Mathis. HOISDF: Constraining 3D hand- object pose estimation with global signed distance fields. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[79]
Contactart: Learning 3D interaction priors for category-level ar- ticulated object and hand poses estimation
Zehao Zhu, Jiashun Wang, Yuzhe Qin, Deqing Sun, Varun Jampani, and Xiaolong Wang. Contactart: Learning 3D interaction priors for category-level ar- ticulated object and hand poses estimation. InProc. Int. Conf. 3D Vis. (3DV), 2024
2024
-
[80]
gSDF: Geometry-driven signed dis- tance functions for 3D hand-object reconstruction
Zerui Chen, Shizhe Chen, Cordelia Schmid, and Ivan Laptev. gSDF: Geometry-driven signed dis- tance functions for 3D hand-object reconstruction. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023
2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.