Pith. sign in

REVIEW 2 major objections 5 minor 300 references

Foundation models help hand–object interaction only when we name the prior they inject and the uncertainty it reduces—not when we merely say a method “uses a large model.”

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 08:48 UTC pith:FNALSTCS

load-bearing objection Solid first map of foundation-model priors for HOI; the taxonomy is the real contribution and the §1 boundary holds up well enough. the 2 major comments →

arxiv 2607.28394 v1 pith:FNALSTCS submitted 2026-07-30 cs.CV

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer

classification cs.CV
keywords Hand-Object InteractionHOI ReconstructionHOI GenerationFoundation-Model PriorsEmbodied TransferContact and AffordanceGeometric PriorsSemantic Priors
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Modeling how hands grasp and move objects is hard because shape, pose, contact, meaning, and motion are all partly hidden. This survey argues that large pretrained models help only if we treat them as sources of specific prior knowledge, not as generic black boxes. It organizes six reconstruction and generation tasks under a shared pipeline, then classifies eight kinds of foundation-model prior—geometric (retrieve, reconstruct, or locate shape in space), semantic (ground regions or reason in language), and visual (reuse features or generate images and video). For each, it tracks what is injected, how it enters the pipeline, and which of five residual uncertainties it is meant to shrink. It also follows the same HOI signals into robot learning: pretraining on human video, retargeting skills, and building robot training data. The point for a sympathetic reader is practical: without that map, the field stays a pile of “we used CLIP/SAM/diffusion” papers and cannot compose, trust, or evaluate the next systems.

Core claim

The paper’s central claim is that it is the first systematic review of foundation-model priors for hand–object interaction: methods should be classified by which cross-domain prior they introduce, where that prior enters a common HOI pipeline, and which of five uncertainties (shape, spatial, physical, semantic, dynamic) it reduces—via a taxonomy of eight sub-priors in geometric, semantic, and visual families spanning six HOI tasks plus embodied transfer to robots.

What carries the argument

The eight-sub-prior taxonomy (shape retrieval, shape reconstruction, spatial reconstruction; semantic grounding, language reasoning; visual representation, image generation, video generation), paired with a shared injection vocabulary—prior source, injected representation, injection operator, target, task, and residual limitation—that turns “uses a large model” into a comparable account of knowledge flow.

Load-bearing premise

The whole map stands on a sharp definition: only an explicitly named, large-scale general-purpose pretrained model that contributes cross-domain knowledge counts as a foundation prior, while in-domain assets and task-specific initializations do not unless such a model retrieves, generates, or scores them.

What would settle it

Re-label every method in the survey’s main table under the paper’s own definition and check whether primary-versus-auxiliary prior tags and the claim of being first systematic still hold; if many borderline systems flip families or if an earlier survey already organized HOI the same way by prior source, injection site, and uncertainty reduced, the central framing fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • New HOI papers would report prior source, injection operator, and which uncertainty is targeted instead of only architecture and dataset names.
  • Benchmarks would pair geometry scores with contact, physical plausibility, and task or object-state success under one protocol.
  • Multi-prior systems would need confidence, routing, and conflict handling when retrieval, reconstruction, grounding, and generation disagree.
  • Robot learning from human video would be judged on preserved contact, intent, and object-state change—not trajectory similarity alone.
  • Long-horizon egocentric HOI would aim for world-frame camera, hand, object, and contact state across grasp, use, release, and re-grasp.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the injection taxonomy sticks, ablation studies should disable one prior family at a time and measure the matching uncertainty, not only end-task error.
  • The same prior-versus-uncertainty grid could grade non-HOI manipulation stacks (tool use, bimanual assembly) without rewriting the task list.
  • Live repositories tied to this taxonomy will matter only if each new method is forced to declare primary prior and mitigated uncertainty at ingest.
  • Embodied memory after grasp occlusion is a natural stress test: pre-contact HOI state must stay queryable when the hand hides the object.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This survey reviews hand-object interaction (HOI) methods that exploit foundation-model priors. It organizes six HOI tasks (three reconstruction, three generation) and proposes a taxonomy of eight sub-priors in geometric, semantic, and visual families. A stated inclusion rule treats a method as foundation-model-prior only when an explicitly identified large-scale general-purpose pretrained model contributes cross-domain knowledge. The paper maps how priors are represented and injected into HOI pipelines (Table 3; Figs. 3–6), separates non-foundation baselines (§2), reviews HOI-derived embodied transfer to robot learning (§6), and summarizes datasets, metrics, and open challenges, with a live repository for ongoing coverage.

Significance. If the taxonomy and inclusion boundary hold, the paper fills a clear gap relative to prior HOI surveys (Table 1): it reframes a fragmented “uses large models” literature by prior source, injection operator, and residual uncertainty (shape, spatial, physical, semantic, dynamic). The shared pipeline abstraction, primary/auxiliary labeling in Table 3, and embodied-transfer routes (pretraining, skill transfer, data engines) are practically useful for both vision and robotics readers. The live repository and explicit non-foundation baseline section are concrete strengths that support cumulative work beyond a static review.

major comments (2)
  1. [§1, Table 3] §1 and Table 3: The inclusion rule is clear, but the operational rule for assigning the primary family (and thus table grouping and Unc.↓ tags) in multi-prior systems is under-specified. EasyHOI, ArtHOI, MCC-HO, Jiang et al., CHOIR, and GraG combine geometric, semantic, and visual components; “contributes most directly to the HOI solution” is not a reproducible criterion. Please add a short decision procedure (e.g., which module produces the task-defining output, ablation/ablation-proxy, or dependency order) and apply it consistently, or mark ambiguous cases explicitly. Without this, the eight-sub-prior taxonomy remains useful as a map but weaker as a systematic classification.
  2. [§7.2, §8.2] §7.2 and §8.2: The survey correctly criticizes geometry-dominated evaluation and calls for interaction correctness, yet the manuscript does not tabulate which reviewed methods report contact/physical/embodied metrics versus only MPJPE/Chamfer/FID. A compact coverage matrix (method × metric family, or benchmark × reported interaction metrics) would make the “blind spots” claim evidence-based and better support the future-benchmark recommendations. This is load-bearing for the evaluation narrative, not merely a wish-list item.
minor comments (5)
  1. [Figure 7] Figure 7 labels routes as “Sec. 4.2 / 4.3 / 4.4” while the body places embodied transfer in §6; renumber for consistency.
  2. [Table 3] Table 3 mixes venue tags (CVPR 2026, arXiv 2026, etc.) for very recent work; a footnote on inclusion cutoff date and how preprints versus accepted versions are handled would help readers and the live repo stay aligned.
  3. [§3.2] §3.2 argues shape retrieval is sparse as a primary prior; consider one sentence on when retrieval should be preferred over reconstruction (asset coverage, topology stability) to balance the critical tone.
  4. [§2.3.3, §5.4] Occasional notation/typo issues (e.g., “SLAM [53 COPSfM [54]” in §2.3.3; “Pl¨ucker” encoding) should be cleaned in copy-edit.
  5. [Abstract, §1] Keywords and abstract claim “first systematic review”; keep that claim tied to the prior-injection lens (as in Table 1) so it is not read as first HOI survey overall.

Circularity Check

0 steps flagged

No circularity: a bibliographic taxonomy survey with no fitted predictions or self-justifying derivation chain.

full rationale

This paper is a literature survey. It does not derive a quantitative prediction, fit parameters to data and re-label them as forecasts, or invoke a uniqueness theorem that collapses to author-only prior work. Its central contribution is an organizational cut (foundation-model prior vs. non-foundation-prior; eight sub-priors; six HOI tasks plus embodied transfer) applied to external methods. That cut is definitional taxonomy, not a circular derivation: inclusion criteria are stated explicitly in §1, methods are tabulated with primary/auxiliary prior tags drawn from named external foundation models (DUSt3R, SAM, DINOv2, Stable Diffusion, etc.), and residual uncertainties are framed as open evaluation gaps rather than results forced by the inputs. Author-adjacent systems (e.g., GeoHand, HandOS) and the live repository appear as map entries or infrastructure, which is normal survey practice and does not make the taxonomy or “first systematic review” claim true by construction. Under the stated circularity criteria there is nothing to reduce: no Eq. X = Eq. Y by fit, no load-bearing self-cited uniqueness, no ansatz smuggled in as a theorem. Score 0 is the honest finding.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 2 invented entities

A survey’s load-bearing commitments are definitional and scope choices, not fitted constants. The central map stands on (1) a crisp inclusion boundary for “foundation-model prior,” (2) the claim that HOI residual difficulty factors into five named uncertainties, and (3) the invented but operational taxonomy of eight sub-priors and injection operators. No numerical free parameters underwrite the claims.

axioms (4)
  • ad hoc to paper A method is a foundation-model-prior method only if an explicitly identified large-scale general-purpose pretrained model contributes cross-domain knowledge via predictions, representations, parameters, adaptation, or distillation (§1).
    This boundary excludes HOI-domain assets and task-specialized pretraining unless a foundation model mediates them; the entire taxonomy partitions on it.
  • domain assumption HOI difficulty under visual uncertainty factors into five residual uncertainties—shape, spatial, physical, semantic, and dynamic—that foundation priors can target (Intro; Fig. 3).
    Standard in HOI literature but treated as an exhaustive organizing frame for prior interventions.
  • domain assumption Six tasks (R1–R3, G1–G3) plus embodied transfer suffice to cover the foundation-era HOI literature for this review (§2.1, §6).
    Contact/affordance are demoted to attributes; editing/reenactment folded into G3—scope choices that shape inclusion.
  • domain assumption Cited concurrent methods’ reported mechanisms can be trusted enough to assign primary/auxiliary prior tags without re-running experiments (Table 3; §§3–6).
    Normal survey epistemology; errors in source papers propagate into the taxonomy.
invented entities (2)
  • Eight foundation-model sub-priors in three families (G-Ret, G-Rec, G-Spa; S-Gnd, S-Lng; V-Rep, V-Img, V-Vid) no independent evidence
    purpose: Provide a stable vocabulary for what knowledge enters HOI pipelines and which uncertainty it reduces.
    The grouping is the paper’s main conceptual product; methods are forced into primary-family labels under it.
  • Shared injection-operator vocabulary (retrieve/align/regularize, initialize, token fusion, region conditioning, condition/fuse, weight init/fine-tune, adapter/ControlNet, score-guided regularization, no independent evidence
    purpose: Describe how priors enter backbone vs. task head without equating all “uses of large models.”
    Operators are analytical constructs over existing systems, not independently measured mechanisms.

pith-pipeline@v1.2.0-daily-grok45 · 58684 in / 3207 out tokens · 70661 ms · 2026-07-31T08:48:52.592954+00:00 · methodology

0 comments
read the original abstract

Hand-object interaction (HOI) modeling remains challenging because it requires joint reasoning about hand articulation, object geometry, contact, semantics, and dynamics under severe visual uncertainty. Foundation models introduce transferable prior knowledge learned from large-scale cross-domain data, offering new ways to address these challenges beyond task-specific data and models. However, the rapidly growing literature remains fragmented, and existing studies typically describe these methods simply as ``using large models'' without systematically characterizing what knowledge is introduced, where it enters the HOI pipeline, or which HOI uncertainty it helps reduce. This survey presents the first systematic review of foundation-model priors for HOI. We organize the literature into six HOI tasks spanning reconstruction and generation. More importantly, we establish a taxonomy of eight foundation-model sub-priors grouped into geometric, semantic, and visual families. Geometric priors encompass shape retrieval, shape reconstruction, and spatial reconstruction; semantic priors include semantic grounding and language reasoning; and visual priors cover visual representation, image generation, and video generation. Based on this taxonomy, we systematically analyze how different priors are represented, injected, and adapted across HOI pipelines and tasks. Beyond how foundation models empower HOI, we further examine how HOI-derived knowledge is used in robot learning, including human-data pretraining, human-to-robot skill transfer, and HOI-to-robot data generation. Finally, we summarize datasets and evaluation protocols, and discuss limitations and future directions toward more generalizable HOI systems. To support long-term progress, we curate a live repository that continuously aggregates emerging methods and benchmarks.

Figures

Figures reproduced from arXiv: 2607.28394 by Jiaolong Yang, Junzhi Yu, Lei Zhang, Luping Xiao, Shiyang Liu, Weiquan Lin, Xingyu Chen, Xu Tang, Yu Deng.

Figure 1
Figure 1. Figure 1: Overview of this survey. The center organizes six HOI tasks into reconstruction (pose estimation, object reconstruc [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Taxonomy roadmap of this survey. Three foundation-model prior families are decomposed into section-level [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Residual HOI uncertainties and corresponding [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Injection mechanisms of geometric priors for HOI. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Injection mechanisms of semantic priors for HOI. [PITH_FULL_IMAGE:figures/full_fig_p013_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Injection mechanisms of visual priors for HOI. [PITH_FULL_IMAGE:figures/full_fig_p015_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Five routes by which HOI evidence becomes robot [PITH_FULL_IMAGE:figures/full_fig_p018_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

300 extracted references · 2 canonical work pages

  1. [1]

    H+O: unified egocentric recognition of 3D hand- object poses and interactions

    Bugra Tekin, Federica Bogo, and Marc Pollefeys. H+O: unified egocentric recognition of 3D hand- object poses and interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019. doi: 10.1109/CVPR.2019.00464

  2. [2]

    Learning joint reconstruction of hands and manipulated objects

    Yana Hasson, Gul Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J Black, Ivan Laptev, and Cordelia Schmid. Learning joint reconstruction of hands and manipulated objects. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019

  3. [3]

    CPF: Learning a contact potential field to model the hand-object interaction

    Lixin Yang, Xinyu Zhan, Kailin Li, Wenqiang Xu, Jiefeng Li, and Cewu Lu. CPF: Learning a contact potential field to model the hand-object interaction. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021

  4. [4]

    HOLD: Category-agnostic 3D re- construction of interacting hands and objects from video

    Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen, Muhammed Kocabas, Michael J Black, and Otmar Hilliges. HOLD: Category-agnostic 3D re- construction of interacting hands and objects from video. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024. 23 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Gener...

  5. [5]

    What’s in your hands? 3D reconstruction of generic objects in hands

    Yufei Ye, Abhinav Gupta, and Shubham Tulsiani. What’s in your hands? 3D reconstruction of generic objects in hands. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  6. [6]

    Grasping field: Learning implicit representations for human grasps

    Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael J Black, Krikamol Muandet, and Siyu Tang. Grasping field: Learning implicit representations for human grasps. InProc. Int. Conf. 3D Vis. (3DV), 2020

  7. [7]

    DUSt3R: Geomet- ric 3D vision made easy

    Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. DUSt3R: Geomet- ric 3D vision made easy. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  8. [8]

    Contactpose: A dataset of grasps with object con- tact and hand pose

    Samarth Brahmbhatt, Chengcheng Tang, Christo- pher D Twigg, Charles C Kemp, and James Hays. Contactpose: A dataset of grasps with object con- tact and hand pose. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020

  9. [9]

    S2contact: Graph-based network for 3D hand-object contact estimation with semi-supervised learning

    Tze Ho Elden Tse, Zhongqun Zhang, Kwang In Kim, Ales Leonardis, Feng Zheng, and Hyung Jin Chang. S2contact: Graph-based network for 3D hand-object contact estimation with semi-supervised learning. In Proc. Eur. Conf. Comput. Vis. (ECCV), 2022. doi: 10.1007/978-3-031-19769-7\ 33

  10. [10]

    Contactopt: Optimizing contact to improve grasps

    Patrick Grady, Chengcheng Tang, Christopher D Twigg, Minh V o, Samarth Brahmbhatt, and Charles C Kemp. Contactopt: Optimizing contact to improve grasps. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2021

  11. [11]

    Deep- SimHO: Stable pose estimation for hand-object in- teraction via physics simulation

    Rong Wang, Wei Mao, and Hongdong Li. Deep- SimHO: Stable pose estimation for hand-object in- teraction via physics simulation. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023

  12. [12]

    D- grasp: Physically plausible dynamic grasp synthesis for hand-object interactions

    Sammy Christen, Muhammed Kocabas, Emre Aksan, Jemin Hwangbo, Jie Song, and Otmar Hilliges. D- grasp: Physically plausible dynamic grasp synthesis for hand-object interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  13. [13]

    Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jian- wei Yang, Hang Su, et al. Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024

  14. [14]

    Segment anything

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023

  15. [15]

    SemGrasp: Semantic grasp generation via language aligned discretization

    Kailin Li, Jingbo Wang, Lixin Yang, Cewu Lu, and Bo Dai. SemGrasp: Semantic grasp generation via language aligned discretization. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024

  16. [16]

    Text2grasp: Synthesis of grasps by text prompts for object grasping parts

    Xiaoyun Chang and Yi Sun. Text2grasp: Synthesis of grasps by text prompts for object grasping parts. In Proc. Int. Symp. Neural Netw. (ISNN), 2025

  17. [17]

    Towards unconstrained joint hand-object re- construction from RGB videos

    Yana Hasson, G¨ul Varol, Cordelia Schmid, and Ivan Laptev. Towards unconstrained joint hand-object re- construction from RGB videos. InProc. Int. Conf. 3D Vis. (3DV), 2021

  18. [18]

    Grasp in gaussians: Fast monocular recon- struction of dynamic hand-object interactions.arXiv preprint arXiv:2604.12929, 2026

    Ayce Idil Aytekin, Xu Chen, Zhengyang Shen, Thabo Beeler, Helge Rhodin, Rishabh Dabral, and Christian Theobalt. Grasp in gaussians: Fast monocular recon- struction of dynamic hand-object interactions.arXiv preprint arXiv:2604.12929, 2026

  19. [19]

    EasyHOI: Unleashing the power of large models for reconstructing hand-object inter- actions in the wild

    Yumeng Liu, Xiaoxiao Long, Zemin Yang, Yuan Liu, Marc Habermann, Christian Theobalt, Yuexin Ma, and Wenping Wang. EasyHOI: Unleashing the power of large models for reconstructing hand-object inter- actions in the wild. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  20. [20]

    MagicHOI: Leveraging 3D priors for accurate hand-object recon- struction from short monocular video clips

    Shibo Wang, Haonan He, Maria Parelli, Christoph Gebhardt, Zicong Fan, and Jie Song. MagicHOI: Leveraging 3D priors for accurate hand-object recon- struction from short monocular video clips. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025

  21. [21]

    Diffusion-guided reconstruction of everyday hand-object interaction clips

    Yufei Ye, Poorvi Hebbar, Abhinav Gupta, and Shub- ham Tulsiani. Diffusion-guided reconstruction of everyday hand-object interaction clips. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023

  22. [22]

    Hand-object interaction image gen- eration

    Hezhen Hu, Weilun Wang, Wengang Zhou, and Houqiang Li. Hand-object interaction image gen- eration. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 35, 2022

  23. [23]

    HOIDiffusion: Generating realistic 3D hand-object interaction data

    Mengqi Zhang, Yang Fu, Zheng Ding, Sifei Liu, Zhuowen Tu, and Xiaolong Wang. HOIDiffusion: Generating realistic 3D hand-object interaction data. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2024

  24. [24]

    Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127, 2023

    Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127, 2023. 24 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction...

  25. [25]

    Instantmesh: Efficient 3D mesh generation from a single image with sparse- view large reconstruction models.arXiv preprint arXiv:2404.07191, 2024

    Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Efficient 3D mesh generation from a single image with sparse- view large reconstruction models.arXiv preprint arXiv:2404.07191, 2024

  26. [26]

    OpenShape: Scaling up 3D shape representation towards open-world understanding

    Minghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu, Xuanlin Li, Shizhong Han, Hong Cai, Fatih Porikli, and Hao Su. OpenShape: Scaling up 3D shape representation towards open-world understanding. In Adv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023

  27. [27]

    Visual instruction tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. InAdv. Neu- ral Inf. Process. Syst. (NeurIPS), volume 36, 2023

  28. [28]

    Qwen-VL: A versatile vision-language model for understanding, localiza- tion, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A versatile vision-language model for understanding, localiza- tion, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023

  29. [29]

    DINOv2: Learning robust vi- sual features without supervision.arXiv preprint arXiv:2304.07193, 2023

    Maxime Oquab, Timoth´ee Darcet, Th´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fer- nandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. DINOv2: Learning robust vi- sual features without supervision.arXiv preprint arXiv:2304.07193, 2023

  30. [30]

    High- resolution image synthesis with latent diffusion mod- els

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High- resolution image synthesis with latent diffusion mod- els. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  31. [31]

    Objaverse: A universe of annotated 3D objects

    Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3D objects. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2023

  32. [32]

    Learning trans- ferable visual models from natural language supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sas- try, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning trans- ferable visual models from natural language supervi- sion. InProc. Int. Conf. Mach. Learn. (ICML), 2021

  33. [33]

    Reconstructing hand-held objects in 3D from images and videos

    Jane Wu, Georgios Pavlakos, Georgia Gkioxari, and Jitendra Malik. Reconstructing hand-held objects in 3D from images and videos. InProc. Int. Conf. 3D Vis. (3DV), 2026

  34. [34]

    Ghost: Fast category- agnostic hand-object interaction reconstruction from RGB videos using gaussian splatting

    Ahmed Tawfik Aboukhadra, Marcel Rogge, Nadia Robertini, Abdalla Arafa, Jameel Malik, Ahmed El- hayek, and Didier Stricker. Ghost: Fast category- agnostic hand-object interaction reconstruction from RGB videos using gaussian splatting. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026

  35. [35]

    Hand-held object reconstruction from RGB video with dynamic interaction

    Shijian Jiang, Qi Ye, Rengan Xie, Yuchi Huo, and Jim- ing Chen. Hand-held object reconstruction from RGB video with dynamic interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  36. [36]

    Reconstructing hands in 3D with transformers

    Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstructing hands in 3D with transformers. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  37. [37]

    Wilor: End- to-end 3D hand localization and reconstruction in-the- wild

    Rolandos Alexandros Potamias, Jinglei Zhang, Jiankang Deng, and Stefanos Zafeiriou. Wilor: End- to-end 3D hand localization and reconstruction in-the- wild. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  38. [38]

    ViTPose: Simple vision transformer baselines for human pose estimation

    Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. ViTPose: Simple vision transformer baselines for human pose estimation. InAdv. Neural Inf. Process. Syst. (NeurIPS), 2022

  39. [39]

    Latent action pretraining from videos

    Seonghyeon Ye, Joel Jang, Byeongguk Jeon, Se June Joo, Jianwei Yang, Baolin Peng, Ajay Mandlekar, Reuben Tan, Yu-Wei Chao, Bill Yuchen Lin, et al. Latent action pretraining from videos. InProc. Int. Conf. Learn. Represent. (ICLR), 2025

  40. [40]

    Egovla: Learning vision-language-action models from egocentric hu- man videos.arXiv preprint arXiv:2507.12440, 2025

    Ruihan Yang, Qinxi Yu, Yecheng Wu, Rui Yan, Borui Li, An-Chieh Cheng, Xueyan Zou, Yunhao Fang, Xuxin Cheng, Ri-Zhao Qiu, et al. Egovla: Learning vision-language-action models from egocentric hu- man videos.arXiv preprint arXiv:2507.12440, 2025

  41. [41]

    Dexmv: Imitation learning for dexterous manipula- tion from human videos

    Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang, Ruihan Yang, Yang Fu, and Xiaolong Wang. Dexmv: Imitation learning for dexterous manipula- tion from human videos. InProc. Eur. Conf. Comput. Vis. (ECCV), 2022

  42. [42]

    Egoengine: From ego- centric human videos to high-fidelity dexterous robot demonstrations.arXiv preprint arXiv:2606.12604, 2026

    Yangcen Liu, Shuo Cheng, Xinchen Yin, Woo Chul Shin, Alfred Cueva, Yiran Yang, Zhenyang Chen, Chuye Zhang, and Danfei Xu. Egoengine: From ego- centric human videos to high-fidelity dexterous robot demonstrations.arXiv preprint arXiv:2606.12604, 2026

  43. [43]

    Takehiko Ohkawa, Ryosuke Furuta, and Yoichi Sato. Efficient annotation and learning for 3D hand pose 25 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer estimation: A survey.Int. J. Comput. Vis., 131(12), 2023

  44. [44]

    A survey of deep learning methods and datasets for hand pose estimation from hand-object interaction images.Comput

    Taeyun Woo, Wonjung Park, Woohyun Jeong, and Jinah Park. A survey of deep learning methods and datasets for hand pose estimation from hand-object interaction images.Comput. Graph., 116, 2023

  45. [45]

    Advances in vision-based deep learning methods for interacting hands reconstruction: A survey.Comput

    Yu Miao and Yue Liu. Advances in vision-based deep learning methods for interacting hands reconstruction: A survey.Comput. Graph., 124, 2024

  46. [46]

    An overview of learning-based dexterous grasping: recent advances and future directions.Artif

    Xu Song, Yongyao Li, Yunfan Zhang, Yufei Liu, and Lei Jiang. An overview of learning-based dexterous grasping: recent advances and future directions.Artif. Intell. Rev., 58(10), 2025

  47. [47]

    Arthoi: Taming foundation models for monocular 4D reconstruction of hand-articulated- object interactions

    Zikai Wang, Zhilu Zhang, Yiqing Wang, Hui Li, and Wangmeng Zuo. Arthoi: Taming foundation models for monocular 4D reconstruction of hand-articulated- object interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026

  48. [48]

    Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: modeling and capturing hands and bodies together.ACM Trans. Graph., 36 (6), 2017. doi: 10.1145/3130800.3130883

  49. [49]

    Nimble: a non-rigid hand model with bones and muscles.ACM Trans

    Yuwei Li, Longwen Zhang, Zesong Qiu, Yingwenqi Jiang, Nianyi Li, Yuexin Ma, Yuyao Zhang, Lan Xu, and Jingyi Yu. Nimble: a non-rigid hand model with bones and muscles.ACM Trans. Graph., 41(4), 2022

  50. [50]

    HOnnotate: A method for 3D an- notation of hand and object poses

    Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. HOnnotate: A method for 3D an- notation of hand and object poses. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020

  51. [51]

    AlignSDF: Pose-aligned signed distance fields for hand-object reconstruction

    Zerui Chen, Yana Hasson, Cordelia Schmid, and Ivan Laptev. AlignSDF: Pose-aligned signed distance fields for hand-object reconstruction. InProc. Eur. Conf. Comput. Vis. (ECCV), 2022

  52. [52]

    HandNeRF: Learn- ing to reconstruct hand-object interaction scene from a single RGB image

    Hongsuk Choi, Nikhil Chavan-Dafle, Jiacheng Yuan, V olkan Isler, and Hyunsoo Park. HandNeRF: Learn- ing to reconstruct hand-object interaction scene from a single RGB image. InProc. IEEE Int. Conf. Robot. Autom. (ICRA), 2024

  53. [53]

    Simultaneous localization and mapping: part i.IEEE Robot

    Hugh Durrant-Whyte and Tim Bailey. Simultaneous localization and mapping: part i.IEEE Robot. Autom. Mag., 13(2), 2006

  54. [54]

    Seitz, and Richard Szeliski

    Noah Snavely, Steven M. Seitz, and Richard Szeliski. Photo tourism: Exploring photo collections in 3D. In ACM SIGGRAPH 2006 Papers. ACM, 2006

  55. [55]

    LatentHOI: On the generalizable hand object motion generation with latent hand diffusion

    Muchen Li, Sammy Christen, Chengde Wan, Yu- jun Cai, Renjie Liao, Leonid Sigal, and Shugao Ma. LatentHOI: On the generalizable hand object motion generation with latent hand diffusion. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  56. [56]

    3D hand pose estimation in everyday egocentric images

    Aditya Prakash, Ruisen Tu, Matthew Chang, and Saurabh Gupta. 3D hand pose estimation in everyday egocentric images. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024

  57. [57]

    DDF-HO: Hand-held object recon- struction via conditional directed distance field

    Chenyangguang Zhang, Yan Di, Ruida Zhang, Guangyao Zhai, Fabian Manhardt, Federico Tombari, and Xiangyang Ji. DDF-HO: Hand-held object recon- struction via conditional directed distance field. In Adv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023

  58. [58]

    Ncrf: neural contact radiance fields for free-viewpoint rendering of hand-object interaction

    Zhongqun Zhang, Jifei Song, Eduardo P´erez-Pellitero, Yiren Zhou, Hyung Jin Chang, and Ale ˇs Leonardis. Ncrf: neural contact radiance fields for free-viewpoint rendering of hand-object interaction. InProc. Int. Conf. 3D Vis. (3DV), 2024

  59. [59]

    Gan- hand: Predicting human grasp affordances in multi- object scenes

    Enric Corona, Albert Pumarola, Guillem Alenya, Francesc Moreno-Noguer, and Gr´egory Rogez. Gan- hand: Predicting human grasp affordances in multi- object scenes. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020

  60. [60]

    Artigrasp: Physically plausible synthesis of bi-manual dexterous grasping and articulation

    Hui Zhang, Sammy Christen, Zicong Fan, Luocheng Zheng, Jemin Hwangbo, Jie Song, and Otmar Hilliges. Artigrasp: Physically plausible synthesis of bi-manual dexterous grasping and articulation. InProc. Int. Conf. 3D Vis. (3DV), 2024

  61. [61]

    Deepsdf: Learning continuous signed distance functions for shape representation

    Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2019

  62. [62]

    A skeleton-driven neural occupancy representation for articulated hands

    Korrawe Karunratanakul, Adrian Spurr, Zicong Fan, Otmar Hilliges, and Siyu Tang. A skeleton-driven neural occupancy representation for articulated hands. InProc. Int. Conf. 3D Vis. (3DV), 2021

  63. [63]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020. doi: 10.1007/978-3-030-58452-8 \ 24. 26 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and ...

  64. [64]

    3D gaussian splat- ting for real-time radiance field rendering.ACM Trans

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk¨uhler, and George Drettakis. 3D gaussian splat- ting for real-time radiance field rendering.ACM Trans. Graph., 42(4), 2023

  65. [65]

    Affordpose: A large-scale dataset of hand- object interactions with affordance-driven hand pose

    Juntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu, and Jian Liu. Affordpose: A large-scale dataset of hand- object interactions with affordance-driven hand pose. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023

  66. [66]

    A versatile and differentiable hand-object interaction representation

    Th´eo Morales, Omid Taheri, and Gerard Lacey. A versatile and differentiable hand-object interaction representation. InProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), 2025

  67. [67]

    Black, and Dima Damen

    Siddhant Bansal, Zhifan Zhu, Shashank Tripathi, Ji- ahe Zhao, Michael J. Black, and Dima Damen. To- wards in-the-wild egocentric 3D hand-object pose estimation. InProc. Eur. Conf. Comput. Vis. (ECCV), 2026

  68. [68]

    HOI4D: A 4D egocentric dataset for category-level human-object interaction

    Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4D egocentric dataset for category-level human-object interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2022

  69. [69]

    Arctic: A dataset for dexterous bimanual hand-object manipulation

    Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J Black, and Otmar Hilliges. Arctic: A dataset for dexterous bimanual hand-object manipulation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023

  70. [70]

    Handoccnet: Occlusion-robust 3D hand mesh estimation network

    JoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hong- suk Choi, and Kyoung Mu Lee. Handoccnet: Occlusion-robust 3D hand mesh estimation network. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2022

  71. [71]

    Mobrecon: Mobile-friendly hand mesh reconstruction from monocular image

    Xingyu Chen, Yufeng Liu, Yajiao Dong, Xiong Zhang, Chongyang Ma, Yanmin Xiong, Yuan Zhang, and Xiaoyan Guo. Mobrecon: Mobile-friendly hand mesh reconstruction from monocular image. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  72. [72]

    A simple baseline for ef- ficient hand mesh reconstruction

    Zhishan Zhou, Shihao Zhou, Zhi Lv, Minqiang Zou, Yao Tang, and Jiajun Liang. A simple baseline for ef- ficient hand mesh reconstruction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  73. [73]

    Model-based 3D hand reconstruction via self- supervised learning

    Yujin Chen, Zhigang Tu, Di Kang, Linchao Bao, Ying Zhang, Xuefei Zhe, Ruizhi Chen, and Junsong Yuan. Model-based 3D hand reconstruction via self- supervised learning. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2021

  74. [74]

    Keypoint fusion for RGB-D based 3D hand pose estimation

    Xingyu Liu, Pengfei Ren, Yuanyuan Gao, Jingyu Wang, Haifeng Sun, Qi Qi, Zirui Zhuang, and Jianxin Liao. Keypoint fusion for RGB-D based 3D hand pose estimation. InProc. AAAI Conf. Artif. Intell., 2024

  75. [75]

    Hope-net: A graph-based model for hand-object pose estimation

    Bardia Doosti, Shujon Naha, Majid Mirbagheri, and David J Crandall. Hope-net: A graph-based model for hand-object pose estimation. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020

  76. [76]

    Keypoint transformer: Solv- ing joint identification in challenging hands and ob- ject interactions for accurate 3D pose estimation

    Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Keypoint transformer: Solv- ing joint identification in challenging hands and ob- ject interactions for accurate 3D pose estimation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  77. [77]

    Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision

    Ahmed Tawfik Aboukhadra, Jameel Malik, Ahmed El- hayek, Nadia Robertini, and Didier Stricker. Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision. InProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), 2023

  78. [78]

    HOISDF: Constraining 3D hand- object pose estimation with global signed distance fields

    Haozhe Qi, Chen Zhao, Mathieu Salzmann, and Alexander Mathis. HOISDF: Constraining 3D hand- object pose estimation with global signed distance fields. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  79. [79]

    Contactart: Learning 3D interaction priors for category-level ar- ticulated object and hand poses estimation

    Zehao Zhu, Jiashun Wang, Yuzhe Qin, Deqing Sun, Varun Jampani, and Xiaolong Wang. Contactart: Learning 3D interaction priors for category-level ar- ticulated object and hand poses estimation. InProc. Int. Conf. 3D Vis. (3DV), 2024

  80. [80]

    gSDF: Geometry-driven signed dis- tance functions for 3D hand-object reconstruction

    Zerui Chen, Shizhe Chen, Cordelia Schmid, and Ivan Laptev. gSDF: Geometry-driven signed dis- tance functions for 3D hand-object reconstruction. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023

Showing first 80 references.