Pith. sign in

REVIEW 3 major objections 3 minor 292 references

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer

T0 review · 3 major / 3 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read This survey argues that the scattered field of foundation-model-assisted hand-object interaction is unified by a taxonomy of eight types of prior knowledge in three families, and that every method can be understood by which prior it injects

desk verdict A useful organizing taxonomy for HOI+foundation models, but the uncertainty-mitigation claims are inferred rather than evidenced. read the letter →

arxiv 2607.28394 v2 pith:FNALSTCS submitted 2026-07-30 cs.CV

classification cs.CV
keywords hand-objectinteractionfoundationmodelsHOIreconstructiongenerationembodiedtransfertaxonomycomputervisionsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a systematic survey arguing that hand-object interaction (HOI) research in the foundation-model era is best understood not by architecture or dataset, but by what cross-domain knowledge a pretrained model contributes and where it enters the pipeline. It proposes a taxonomy of eight foundation-model priors in three families—geometric (shape retrieval, shape reconstruction, spatial reconstruction), semantic (semantic grounding, language reasoning), and visual (visual representation, image generation, video generation). It claims this taxonomy can organize methods across six HOI tasks in reconstruction and generation, and that tracing prior injection explains which of five recurring uncertainties each method reduces. It further claims that HOI-derived knowledge transfers to robot learning through five routes, and that current evaluation metrics measure geometry but not interaction correctness. A sympathetic reader would care because the survey gives the field a common vocabulary for comparing methods and designing integrated HOI systems.

What carries the argument

The central organizing device is the eight-sub-prior taxonomy combined with a pipeline abstraction: each method is characterized by its prior source, the representation it injects, and the injection operator it uses. This vocabulary—initialization, regularization, conditioning, token fusion, score-guided regularization, retargeting—describes how foundation-model knowledge enters the HOI backbone and task head. The taxonomy is linked to a five-uncertainty model (shape, spatial, physical, semantic, dynamic), so each prior family is mapped to the specific ambiguities it mitigates. For embodied transfer, the machinery is a five-route diagram tracing HOI evidence through a transferred signal and

What would settle it

A concrete test would be to re-run the survey's taxonomic tables under a broader inclusion rule that also counts task-specialized pretrained initializations as foundation priors; if the qualitative mapping between prior families and the five uncertainties still holds, the boundary is not load-bearing, but if the taxonomy's cells become crowded and the uncertainty mapping blurs, the paper's organizational claim is weakened. Alternatively, a single benchmark that jointly reports geometry, contact agreement, physical plausibility, and task success across the six tasks could settle whether current

Watch

Extended reading notes

Core claim

The paper's central claim is that the apparent fragmentation of HOI-plus-foundation-model work reflects a missing organizing dimension: what knowledge is introduced, where it enters, and which uncertainty it reduces. The authors define a deliberately narrow boundary—a method counts as foundation-model-prior only when an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge through predictions, representations, transferred parameters, adaptation, or distillation—and on that basis catalog eight sub-priors in three families. They argue this taxonomy covers six HOI tasks (pose estimation, hand-held object reconstruction, dynamic reconstruction, gr

Load-bearing premise

The load-bearing premise is the paper's stipulative boundary for what counts as a foundation-model prior: a method is included only when an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge through predictions, representations, transferred parameters, adaptation, or distillation, and task-specialized pretraining does not count.

Editorial extensions

If this is right

  • HOI methods should be characterized and compared by the foundation-model priors they exploit, not only by architecture or training data.
  • Geometric, semantic, and visual priors are complementary, so multi-prior systems are the natural next step for reducing shape, spatial, physical, semantic, and dynamic uncertainty together.
  • Evaluation that reports only geometry or image fidelity is insufficient; contact agreement, physical plausibility, and functional task success must be reported alongside.
  • Embodied transfer is best viewed as a downstream consumer of HOI reconstruction and generation outputs, with concrete transfer routes from human evidence to robot policy.
  • The taxonomy provides a shared vocabulary that can make future HOI papers comparable and can guide the design of integrated, verifiable HOI systems.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy's gatekeeping boundary is an author choice rather than an empirical result; re-running the survey's tables with a broader definition of foundation priors (for example, including task-specialized pretrained initializations) would shift coverage and primary/auxiliary assignments.
  • The eight sub-priors are not cleanly orthogonal in practice: the same vision-language model counts as a semantic grounding prior when used for localization and as a language reasoning prior when used for intent inference, suggesting the taxonomy is a lens for reading methods rather than a unique partition.
  • A concrete extension the paper leaves implicit is a joint benchmark that measures geometry, contact, physical plausibility, and task success together; such a benchmark would directly test the survey's claim that current metrics have blind spots.
  • The emerging line of action-conditioned video generation (world models) is identified as an inference pattern rather than an injection operator; editorially, this could mature into a fourth visual sub-prior or a new task category as interactive rollouts grow.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. This survey organizes the hand-object interaction (HOI) literature around eight foundation-model sub-priors grouped into geometric, semantic, and visual families, and maps how these priors are represented, injected, and adapted across six HOI reconstruction/generation tasks and embodied transfer to robot learning. The paper claims to be the first systematic review of foundation-model priors for HOI, and supports the taxonomy with representative methods in Tables 3–5, a dataset/evaluation summary, and a live repository.

Significance. If the taxonomy and the uncertainty-mitigation mapping are accepted, the survey would provide a useful organizing principle for a rapidly growing but fragmented field. The manuscript is internally consistent, the gatekeeping definition of 'foundation-model prior' is explicit and applied transparently (e.g., excluding ViTPose initialization in HaMeR), and the paper repeatedly identifies evaluation blind spots. These are strengths. However, the central analytical output — the mapping from each prior to the HOI uncertainties it 'helps reduce' — is largely inferential and not backed by the surveyed papers' evidence, which weakens the strongest claim.

major comments (3)
  1. [Abstract, Table 3, Secs. 3.3/5.3/7.2.3] The abstract claims the survey reveals 'which HOI uncertainty it helps reduce,' and Table 3's Unc.↓ column operationalizes this. Yet the cited papers rarely ablate the foundation-model component against the listed uncertainty. The paper itself concedes: Sec. 3.3 states a reconstructed shape 'should not be interpreted as interaction evidence by itself'; Sec. 5.3 says 'image realism is not interaction correctness'; Sec. 7.2.3 notes physical metrics 'depend strongly on mesh quality, friction, contact modeling, and simulator settings.' The Unc.↓ assignments thus appear to be authors' inference, not literature evidence. Please either (a) reclassify the Unc.↓ column as a hypothesized mechanism with a clear caveat, or (b) for each Table 3 row, cite an ablation/experiment from the original paper that supports the assignment. This is load-bearing because the abstract frames the entire survey arou
  2. [Sec. 1] The manuscript claims 'the first systematic review' of foundation-model priors for HOI, but no search/selection methodology is provided: no databases, query terms, inclusion/exclusion criteria, screening process, or date cutoff are documented. Without this, the 'systematic' claim cannot be audited and the survey cannot be distinguished from an author-selected narrative review. Please add a methodology paragraph (or appendix) documenting the protocol, or soften the claim to 'first literature survey' / 'comprehensive review'.
  3. [Table 3 and Table 4] The representative method list and dataset tables rely on a large number of arXiv preprints and very recent 2026 venue entries (e.g., GeoHand arXiv:2605.17354, ScaleHP arXiv:2606.25619, several CVPR 2026 entries). Given the survey's cutoff is not stated, the 'first systematic' claim and the balanced coverage of the field are difficult to assess. State the literature cutoff date, and mark entries that are preprints or not yet peer-reviewed at that date. This is especially relevant because several Table 3 exemplars are from the authors' own group, and the selection criteria for 'representative' methods are not specified.
minor comments (3)
  1. [Fig. 7] The figure caption and labels refer to 'Sec. 4.2: Human-Data Pretraining', 'Sec. 4.3: Human-to-Robot Skill Transfer', and 'Sec. 4.4: HOI-to-Robot Data Engines', but the corresponding sections in the text are 6.2, 6.3, and 6.4. Update the figure numbering.
  2. [Sec. 6.2.1] The phrase 'frame-aligned action chunks' and '6DoF object trajectories' appear without prior definition in the HOI taxonomy; consider adding these to the interaction-representation list in Sec. 2.2.2 for terminological consistency.
  3. [Throughout] The text frequently uses 'systematically analyze' and 'systematic coverage' without a clear definition of systematicity. Align these terms with the (proposed) methodology, or use more neutral phrasing such as 'structured analysis.'

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the taxonomy is stipulative and the uncertainty mapping is interpretive, but no result reduces to its inputs by construction.

full rationale

This is a survey, not a derivation. Its central deliverable is a taxonomy of foundation-model priors and a qualitative mapping from priors to five uncertainties. The boundary in Sec. 1 is explicitly stipulative: "we use a deliberately narrow boundary: a method is considered a foundation-model-prior method only when an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge to the HOI pipeline." Classification decisions are therefore authorial choices, not circular deductions. The Unc.↓ column in Table 3 is interpretive attribution rather than a fitted/predicted quantity; the paper itself qualifies the evidence base in several places: Sec. 3.3 says "the reconstructed shape should not be interpreted as interaction evidence by itself," Sec. 5.3 says "image realism is not interaction correctness," Sec. 7.2.3 says penetration and force-closure measures "depend strongly on mesh quality, friction, contact modeling, and simulator settings," and Sec. 8.4 lists prior reliability as an open problem. These caveats weaken the empirical support for the mapping, but they do not make it circular: no quantity is fitted and then presented as a prediction, and no equation or definition forces the survey's conclusions. The self-citations (GeoHand, HandOS, and ScaleHP in Table 3, with author overlap; MoGe-2 also has overlapping authors) are visible, but these are externally falsifiable method papers used as exemplars, the taxonomy does not depend on them, and per rule 4 self-citation alone is not circularity. The central claim—that the fragmented HOI literature can be organized by eight foundation-model sub-priors in three families—is a transparent analytic frame with independent content.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey rests on three declared or implicit premises: a stipulative boundary for what counts as a foundation-model prior; a five-way decomposition of HOI uncertainty; and an assumption that each method can be tagged with a single primary sub-prior. These are author choices rather than demonstrated empirical facts; they gate every subsequent table and figure. No free parameters or invented physical entities appear.

assumptions (3)
  • ad hoc to paper A method counts as foundation-model-prior only if an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge through predictions, representations, transferred parameters, adaptation, or distillation.
    Stipulative boundary defined in Sec. 1 ('To make this question precise...'); the entire selection and classification in Tables 3-5 and Figs. 4-6 depend on this inclusion rule, and it excludes e.g. HaMeR's ViTPose initialization.
  • domain assumption HOI failure modes decompose into five uncertainties: shape, spatial, physical, semantic, dynamic.
    Introduced in Sec. 1 and used by Fig. 3 and all prior-family sections to explain which priors mitigate which uncertainties; the mapping loses force if the decomposition is not exhaustive or orthogonal.
  • domain assumption Each method can be assigned a unique primary sub-prior and optional auxiliary sub-priors with a single injection operator.
    Table 3 tags each method with P/A and operator; the survey's analyses in Secs. 3-5 assume these assignments are unambiguously correct, but many methods combine multiple priors and the primary/auxiliary choice is judgment-based.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer." pith.science (2026). https://pith.science/paper/FNALSTCS

@misc{pith2026260728394,
  author       = {Pith},
  title        = {Pith review of: Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FNALSTCS}},
  note         = {Machine review of arXiv:2607.28394}
}
read the original abstract

Hand-object interaction (HOI) modeling remains challenging because it requires joint reasoning about hand articulation, object geometry, contact, semantics, and dynamics under severe visual uncertainty. Foundation models introduce transferable prior knowledge learned from large-scale cross-domain data, offering new ways to address these challenges beyond task-specific data and models. However, the rapidly growing literature remains fragmented, and existing studies typically describe these methods simply as ``using large models'' without systematically characterizing what knowledge is introduced, where it enters the HOI pipeline, or which HOI uncertainty it helps reduce. This survey presents the first systematic review of foundation-model priors for HOI. We organize the literature into six HOI tasks spanning reconstruction and generation. More importantly, we establish a taxonomy of eight foundation-model sub-priors grouped into geometric, semantic, and visual families. Geometric priors encompass shape retrieval, shape reconstruction, and spatial reconstruction; semantic priors include semantic grounding and language reasoning; and visual priors cover visual representation, image generation, and video generation. Based on this taxonomy, we systematically analyze how different priors are represented, injected, and adapted across HOI pipelines and tasks. Beyond how foundation models empower HOI, we further examine how HOI-derived knowledge is used in robot learning, including human-data pretraining, human-to-robot skill transfer, and HOI-to-robot data generation. Finally, we summarize datasets and evaluation protocols, and discuss limitations and future directions toward more generalizable HOI systems. To support long-term progress, we curate a live repository that continuously aggregates emerging methods and benchmarks.

Figures

Figures reproduced from arXiv: 2607.28394 by the authors.

Figure 1
Figure 1. Overview of this survey. The center organizes six HOI tasks into reconstruction (pose estimation, object reconstruc [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Taxonomy roadmap of this survey. Three foundation-model prior families are decomposed into section-level [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Residual HOI uncertainties and corresponding [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Injection mechanisms of geometric priors for HOI. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Injection mechanisms of semantic priors for HOI. [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 6
Figure 6. Figure 6: Injection mechanisms of visual priors for HOI. [PITH_FULL_IMAGE:figures/full_fig_p015_6.png]
Figure 7
Figure 7. Figure 7: Five routes by which HOI evidence becomes robot [PITH_FULL_IMAGE:figures/full_fig_p018_7.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

292 extracted references · 66 linked inside Pith

  1. [1]

    H+O: unified egocentric recognition of 3D hand- object poses and interactions

    Bugra Tekin, Federica Bogo, and Marc Pollefeys. H+O: unified egocentric recognition of 3D hand- object poses and interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019. doi: 10.1109/CVPR.2019.00464

  2. [2]

    Learning joint reconstruction of hands and manipulated objects

    Yana Hasson, Gul Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J Black, Ivan Laptev, and Cordelia Schmid. Learning joint reconstruction of hands and manipulated objects. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019

  3. [3]

    CPF: Learning a contact potential field to model the hand-object interaction

    Lixin Yang, Xinyu Zhan, Kailin Li, Wenqiang Xu, Jiefeng Li, and Cewu Lu. CPF: Learning a contact potential field to model the hand-object interaction. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021

  4. [4]

    HOLD: Category-agnostic 3D re- construction of interacting hands and objects from video

    Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen, Muhammed Kocabas, Michael J Black, and Otmar Hilliges. HOLD: Category-agnostic 3D re- construction of interacting hands and objects from video. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  5. [5]

    What’s in your hands? 3D reconstruction of generic objects in hands

    Yufei Ye, Abhinav Gupta, and Shubham Tulsiani. What’s in your hands? 3D reconstruction of generic objects in hands. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  6. [6]

    Grasping field: Learning implicit representations for human grasps

    Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael J Black, Krikamol Muandet, and Siyu Tang. Grasping field: Learning implicit representations for human grasps. InProc. Int. Conf. 3D Vis. (3DV), 2020

  7. [7]

    DUSt3R: Geomet- ric 3D vision made easy

    Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. DUSt3R: Geomet- ric 3D vision made easy. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  8. [8]

    Samarth Brahmbhatt, Chengcheng Tang, Christo- pher D Twigg, Charles C Kemp, and James Hays. 23 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer Contactpose: A dataset of grasps with object con- tact and hand pose. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020

Show all 292 references
  1. [9]

    S2contact: Graph-based network for 3D hand-object contact estimation with semi-supervised learning

    Tze Ho Elden Tse, Zhongqun Zhang, Kwang In Kim, Ales Leonardis, Feng Zheng, and Hyung Jin Chang. S2contact: Graph-based network for 3D hand-object contact estimation with semi-supervised learning. In Proc. Eur. Conf. Comput. Vis. (ECCV), 2022. doi: 10.1007/978-3-031-19769-7\ 33

  2. [10]

    Contactopt: Optimizing contact to improve grasps

    Patrick Grady, Chengcheng Tang, Christopher D Twigg, Minh V o, Samarth Brahmbhatt, and Charles C Kemp. Contactopt: Optimizing contact to improve grasps. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2021

  3. [11]

    Deep- SimHO: Stable pose estimation for hand-object in- teraction via physics simulation

    Rong Wang, Wei Mao, and Hongdong Li. Deep- SimHO: Stable pose estimation for hand-object in- teraction via physics simulation. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023

  4. [12]

    D- grasp: Physically plausible dynamic grasp synthesis for hand-object interactions

    Sammy Christen, Muhammed Kocabas, Emre Aksan, Jemin Hwangbo, Jie Song, and Otmar Hilliges. D- grasp: Physically plausible dynamic grasp synthesis for hand-object interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  5. [13]

    Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jian- wei Yang, Hang Su, et al. Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024

  6. [14]

    Segment anything

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023

  7. [15]

    SemGrasp: Semantic grasp generation via language aligned discretization

    Kailin Li, Jingbo Wang, Lixin Yang, Cewu Lu, and Bo Dai. SemGrasp: Semantic grasp generation via language aligned discretization. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024

  8. [16]

    Text2grasp: Synthesis of grasps by text prompts for object grasping parts

    Xiaoyun Chang and Yi Sun. Text2grasp: Synthesis of grasps by text prompts for object grasping parts. In Proc. Int. Symp. Neural Netw. (ISNN), 2025

  9. [17]

    Towards unconstrained joint hand-object re- construction from RGB videos

    Yana Hasson, G¨ul Varol, Cordelia Schmid, and Ivan Laptev. Towards unconstrained joint hand-object re- construction from RGB videos. InProc. Int. Conf. 3D Vis. (3DV), 2021

  10. [18]

    Grasp in gaussians: Fast monocular recon- struction of dynamic hand-object interactions.arXiv preprint arXiv:2604.12929, 2026

    Ayce Idil Aytekin, Xu Chen, Zhengyang Shen, Thabo Beeler, Helge Rhodin, Rishabh Dabral, and Christian Theobalt. Grasp in gaussians: Fast monocular recon- struction of dynamic hand-object interactions.arXiv preprint arXiv:2604.12929, 2026

  11. [19]

    EasyHOI: Unleashing the power of large models for reconstructing hand-object inter- actions in the wild

    Yumeng Liu, Xiaoxiao Long, Zemin Yang, Yuan Liu, Marc Habermann, Christian Theobalt, Yuexin Ma, and Wenping Wang. EasyHOI: Unleashing the power of large models for reconstructing hand-object inter- actions in the wild. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  12. [20]

    MagicHOI: Leveraging 3D priors for accurate hand-object recon- struction from short monocular video clips

    Shibo Wang, Haonan He, Maria Parelli, Christoph Gebhardt, Zicong Fan, and Jie Song. MagicHOI: Leveraging 3D priors for accurate hand-object recon- struction from short monocular video clips. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025

  13. [21]

    Diffusion-guided reconstruction of everyday hand-object interaction clips

    Yufei Ye, Poorvi Hebbar, Abhinav Gupta, and Shub- ham Tulsiani. Diffusion-guided reconstruction of everyday hand-object interaction clips. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023

  14. [22]

    Hand-object interaction image gen- eration

    Hezhen Hu, Weilun Wang, Wengang Zhou, and Houqiang Li. Hand-object interaction image gen- eration. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 35, 2022

  15. [23]

    HOIDiffusion: Generating realistic 3D hand-object interaction data

    Mengqi Zhang, Yang Fu, Zheng Ding, Sifei Liu, Zhuowen Tu, and Xiaolong Wang. HOIDiffusion: Generating realistic 3D hand-object interaction data. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2024

  16. [24]

    Svimo: Synchro- nized diffusion for video and motion generation in hand-object interaction scenarios

    Lingwei Dang, Ruizhi Shao, Hongwen Zhang, Wei Min, Yebin Liu, and Qingyao Wu. Svimo: Synchro- nized diffusion for video and motion generation in hand-object interaction scenarios. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 38, 2025

  17. [25]

    Instantmesh: Efficient 3D mesh generation from a single image with sparse- view large reconstruction models.arXiv preprint arXiv:2404.07191, 2024

    Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Efficient 3D mesh generation from a single image with sparse- view large reconstruction models.arXiv preprint arXiv:2404.07191, 2024

  18. [26]

    OpenShape: Scaling up 3D shape representation towards open-world understanding

    Minghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu, Xuanlin Li, Shizhong Han, Hong Cai, Fatih Porikli, and Hao Su. OpenShape: Scaling up 3D shape representation towards open-world understanding. In Adv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023

  19. [27]

    Visual instruction tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. InAdv. Neu- ral Inf. Process. Syst. (NeurIPS), volume 36, 2023. 24 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer

  20. [28]

    Qwen-VL: A versatile vision-language model for understanding, localiza- tion, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A versatile vision-language model for understanding, localiza- tion, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023

  21. [29]

    DINOv2: Learning robust vi- sual features without supervision.arXiv preprint arXiv:2304.07193, 2023

    Maxime Oquab, Timoth´ee Darcet, Th´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fer- nandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. DINOv2: Learning robust vi- sual features without supervision.arXiv preprint arXiv:2304.07193, 2023

  22. [30]

    High- resolution image synthesis with latent diffusion mod- els

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High- resolution image synthesis with latent diffusion mod- els. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  23. [31]

    Cogvideox: Text-to-video diffusion models with an expert transformer

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. InProc. Int. Conf. Learn. Repre- sent. (ICLR), 2025

  24. [32]

    Objaverse: A universe of annotated 3D objects

    Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3D objects. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2023

  25. [33]

    Learning trans- ferable visual models from natural language supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sas- try, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning trans- ferable visual models from natural language supervi- sion. InProc. Int. C...

  26. [34]

    Reconstructing hand-held objects in 3D from images and videos

    Jane Wu, Georgios Pavlakos, Georgia Gkioxari, and Jitendra Malik. Reconstructing hand-held objects in 3D from images and videos. InProc. Int. Conf. 3D Vis. (3DV), 2026

  27. [35]

    Ghost: Fast category- agnostic hand-object interaction reconstruction from RGB videos using gaussian splatting

    Ahmed Tawfik Aboukhadra, Marcel Rogge, Nadia Robertini, Abdalla Arafa, Jameel Malik, Ahmed El- hayek, and Didier Stricker. Ghost: Fast category- agnostic hand-object interaction reconstruction from RGB videos using gaussian splatting. InProc. IEEE/CVF Conf. Comput. Vis. Patter...

  28. [36]

    Hand-held object reconstruction from RGB video with dynamic interaction

    Shijian Jiang, Qi Ye, Rengan Xie, Yuchi Huo, and Jim- ing Chen. Hand-held object reconstruction from RGB video with dynamic interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  29. [37]

    Reconstructing hands in 3D with transformers

    Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstructing hands in 3D with transformers. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  30. [38]

    Wilor: End- to-end 3D hand localization and reconstruction in-the- wild

    Rolandos Alexandros Potamias, Jinglei Zhang, Jiankang Deng, and Stefanos Zafeiriou. Wilor: End- to-end 3D hand localization and reconstruction in-the- wild. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  31. [39]

    ViTPose: Simple vision transformer baselines for human pose estimation

    Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. ViTPose: Simple vision transformer baselines for human pose estimation. InAdv. Neural Inf. Process. Syst. (NeurIPS), 2022

  32. [40]

    Latent action pretraining from videos

    Seonghyeon Ye, Joel Jang, Byeongguk Jeon, Se June Joo, Jianwei Yang, Baolin Peng, Ajay Mandlekar, Reuben Tan, Yu-Wei Chao, Bill Yuchen Lin, et al. Latent action pretraining from videos. InProc. Int. Conf. Learn. Represent. (ICLR), 2025

  33. [41]

    Egovla: Learning vision-language-action models from egocentric hu- man videos.arXiv preprint arXiv:2507.12440, 2025

    Ruihan Yang, Qinxi Yu, Yecheng Wu, Rui Yan, Borui Li, An-Chieh Cheng, Xueyan Zou, Yunhao Fang, Xuxin Cheng, Ri-Zhao Qiu, et al. Egovla: Learning vision-language-action models from egocentric hu- man videos.arXiv preprint arXiv:2507.12440, 2025

  34. [42]

    Dexmv: Imitation learning for dexterous manipula- tion from human videos

    Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang, Ruihan Yang, Yang Fu, and Xiaolong Wang. Dexmv: Imitation learning for dexterous manipula- tion from human videos. InProc. Eur. Conf. Comput. Vis. (ECCV), 2022

  35. [43]

    Egoengine: From ego- centric human videos to high-fidelity dexterous robot demonstrations.arXiv preprint arXiv:2606.12604, 2026

    Yangcen Liu, Shuo Cheng, Xinchen Yin, Woo Chul Shin, Alfred Cueva, Yiran Yang, Zhenyang Chen, Chuye Zhang, and Danfei Xu. Egoengine: From ego- centric human videos to high-fidelity dexterous robot demonstrations.arXiv preprint arXiv:2606.12604, 2026

  36. [44]

    Efficient annotation and learning for 3D hand pose estimation: A survey.Int

    Takehiko Ohkawa, Ryosuke Furuta, and Yoichi Sato. Efficient annotation and learning for 3D hand pose estimation: A survey.Int. J. Comput. Vis., 131(12), 2023

  37. [45]

    A survey of deep learning methods and datasets for hand pose estimation from hand-object interaction images.Comput

    Taeyun Woo, Wonjung Park, Woohyun Jeong, and Jinah Park. A survey of deep learning methods and datasets for hand pose estimation from hand-object interaction images.Comput. Graph., 116, 2023

  38. [46]

    Advances in vision-based deep learning methods for interacting hands reconstruction: A survey.Comput

    Yu Miao and Yue Liu. Advances in vision-based deep learning methods for interacting hands reconstruction: A survey.Comput. Graph., 124, 2024. 25 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer

  39. [47]

    An overview of learning-based dexterous grasping: recent advances and future directions.Artif

    Xu Song, Yongyao Li, Yunfan Zhang, Yufei Liu, and Lei Jiang. An overview of learning-based dexterous grasping: recent advances and future directions.Artif. Intell. Rev., 58(10), 2025

  40. [48]

    Arthoi: Taming foundation models for monocular 4D reconstruction of hand-articulated- object interactions

    Zikai Wang, Zhilu Zhang, Yiqing Wang, Hui Li, and Wangmeng Zuo. Arthoi: Taming foundation models for monocular 4D reconstruction of hand-articulated- object interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026

  41. [49]

    Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: modeling and capturing hands and bodies together.ACM Trans. Graph., 36 (6), 2017. doi: 10.1145/3130800.3130883

  42. [50]

    Nimble: a non-rigid hand model with bones and muscles.ACM Trans

    Yuwei Li, Longwen Zhang, Zesong Qiu, Yingwenqi Jiang, Nianyi Li, Yuexin Ma, Yuyao Zhang, Lan Xu, and Jingyi Yu. Nimble: a non-rigid hand model with bones and muscles.ACM Trans. Graph., 41(4), 2022

  43. [51]

    HOnnotate: A method for 3D an- notation of hand and object poses

    Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. HOnnotate: A method for 3D an- notation of hand and object poses. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020

  44. [52]

    AlignSDF: Pose-aligned signed distance fields for hand-object reconstruction

    Zerui Chen, Yana Hasson, Cordelia Schmid, and Ivan Laptev. AlignSDF: Pose-aligned signed distance fields for hand-object reconstruction. InProc. Eur. Conf. Comput. Vis. (ECCV), 2022

  45. [53]

    HandNeRF: Learn- ing to reconstruct hand-object interaction scene from a single RGB image

    Hongsuk Choi, Nikhil Chavan-Dafle, Jiacheng Yuan, V olkan Isler, and Hyunsoo Park. HandNeRF: Learn- ing to reconstruct hand-object interaction scene from a single RGB image. InProc. IEEE Int. Conf. Robot. Autom. (ICRA), 2024

  46. [54]

    Simultaneous localization and mapping: part i.IEEE Robot

    Hugh Durrant-Whyte and Tim Bailey. Simultaneous localization and mapping: part i.IEEE Robot. Autom. Mag., 13(2), 2006

  47. [55]

    Seitz, and Richard Szeliski

    Noah Snavely, Steven M. Seitz, and Richard Szeliski. Photo tourism: Exploring photo collections in 3D. In ACM SIGGRAPH 2006 Papers. ACM, 2006

  48. [56]

    LatentHOI: On the generalizable hand object motion generation with latent hand diffusion

    Muchen Li, Sammy Christen, Chengde Wan, Yu- jun Cai, Renjie Liao, Leonid Sigal, and Shugao Ma. LatentHOI: On the generalizable hand object motion generation with latent hand diffusion. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  49. [57]

    3D hand pose estimation in everyday egocentric images

    Aditya Prakash, Ruisen Tu, Matthew Chang, and Saurabh Gupta. 3D hand pose estimation in everyday egocentric images. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024

  50. [58]

    DDF-HO: Hand-held object recon- struction via conditional directed distance field

    Chenyangguang Zhang, Yan Di, Ruida Zhang, Guangyao Zhai, Fabian Manhardt, Federico Tombari, and Xiangyang Ji. DDF-HO: Hand-held object recon- struction via conditional directed distance field. In Adv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023

  51. [59]

    Ncrf: neural contact radiance fields for free-viewpoint rendering of hand-object interaction

    Zhongqun Zhang, Jifei Song, Eduardo P´erez-Pellitero, Yiren Zhou, Hyung Jin Chang, and Ale ˇs Leonardis. Ncrf: neural contact radiance fields for free-viewpoint rendering of hand-object interaction. InProc. Int. Conf. 3D Vis. (3DV), 2024

  52. [60]

    Gan- hand: Predicting human grasp affordances in multi- object scenes

    Enric Corona, Albert Pumarola, Guillem Alenya, Francesc Moreno-Noguer, and Gr´egory Rogez. Gan- hand: Predicting human grasp affordances in multi- object scenes. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020

  53. [61]

    Artigrasp: Physically plausible synthesis of bi-manual dexterous grasping and articulation

    Hui Zhang, Sammy Christen, Zicong Fan, Luocheng Zheng, Jemin Hwangbo, Jie Song, and Otmar Hilliges. Artigrasp: Physically plausible synthesis of bi-manual dexterous grasping and articulation. InProc. Int. Conf. 3D Vis. (3DV), 2024

  54. [62]

    Deepsdf: Learning continuous signed distance functions for shape representation

    Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2019

  55. [63]

    A skeleton-driven neural occupancy representation for articulated hands

    Korrawe Karunratanakul, Adrian Spurr, Zicong Fan, Otmar Hilliges, and Siyu Tang. A skeleton-driven neural occupancy representation for articulated hands. InProc. Int. Conf. 3D Vis. (3DV), 2021

  56. [64]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020. doi: 10.1007/978-3-030-58452-8 \ 24

  57. [65]

    3D gaussian splat- ting for real-time radiance field rendering.ACM Trans

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk¨uhler, and George Drettakis. 3D gaussian splat- ting for real-time radiance field rendering.ACM Trans. Graph., 42(4), 2023

  58. [66]

    Affordpose: A large-scale dataset of hand- object interactions with affordance-driven hand pose

    Juntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu, and Jian Liu. Affordpose: A large-scale dataset of hand- object interactions with affordance-driven hand pose. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023

  59. [67]

    Th´eo Morales, Omid Taheri, and Gerard Lacey. A versatile and differentiable hand-object interaction 26 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer representation. InProc. IEEE/CVF Winter Conf. Appl. Comput. ...

  60. [68]

    Black, and Dima Damen

    Siddhant Bansal, Zhifan Zhu, Shashank Tripathi, Ji- ahe Zhao, Michael J. Black, and Dima Damen. To- wards in-the-wild egocentric 3D hand-object pose estimation. InProc. Eur. Conf. Comput. Vis. (ECCV), 2026

  61. [69]

    HOI4D: A 4D egocentric dataset for category-level human-object interaction

    Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4D egocentric dataset for category-level human-object interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2022

  62. [70]

    Arctic: A dataset for dexterous bimanual hand-object manipulation

    Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J Black, and Otmar Hilliges. Arctic: A dataset for dexterous bimanual hand-object manipulation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023

  63. [71]

    Handoccnet: Occlusion-robust 3D hand mesh estimation network

    JoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hong- suk Choi, and Kyoung Mu Lee. Handoccnet: Occlusion-robust 3D hand mesh estimation network. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2022

  64. [72]

    Mobrecon: Mobile-friendly hand mesh reconstruction from monocular image

    Xingyu Chen, Yufeng Liu, Yajiao Dong, Xiong Zhang, Chongyang Ma, Yanmin Xiong, Yuan Zhang, and Xiaoyan Guo. Mobrecon: Mobile-friendly hand mesh reconstruction from monocular image. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  65. [73]

    A simple baseline for ef- ficient hand mesh reconstruction

    Zhishan Zhou, Shihao Zhou, Zhi Lv, Minqiang Zou, Yao Tang, and Jiajun Liang. A simple baseline for ef- ficient hand mesh reconstruction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  66. [74]

    Model-based 3D hand reconstruction via self- supervised learning

    Yujin Chen, Zhigang Tu, Di Kang, Linchao Bao, Ying Zhang, Xuefei Zhe, Ruizhi Chen, and Junsong Yuan. Model-based 3D hand reconstruction via self- supervised learning. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2021

  67. [75]

    Keypoint fusion for RGB-D based 3D hand pose estimation

    Xingyu Liu, Pengfei Ren, Yuanyuan Gao, Jingyu Wang, Haifeng Sun, Qi Qi, Zirui Zhuang, and Jianxin Liao. Keypoint fusion for RGB-D based 3D hand pose estimation. InProc. AAAI Conf. Artif. Intell., 2024

  68. [76]

    Hope-net: A graph-based model for hand-object pose estimation

    Bardia Doosti, Shujon Naha, Majid Mirbagheri, and David J Crandall. Hope-net: A graph-based model for hand-object pose estimation. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020

  69. [77]

    Keypoint transformer: Solv- ing joint identification in challenging hands and ob- ject interactions for accurate 3D pose estimation

    Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Keypoint transformer: Solv- ing joint identification in challenging hands and ob- ject interactions for accurate 3D pose estimation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  70. [78]

    Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision

    Ahmed Tawfik Aboukhadra, Jameel Malik, Ahmed El- hayek, Nadia Robertini, and Didier Stricker. Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision. InProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), 2023

  71. [79]

    HOISDF: Constraining 3D hand- object pose estimation with global signed distance fields

    Haozhe Qi, Chen Zhao, Mathieu Salzmann, and Alexander Mathis. HOISDF: Constraining 3D hand- object pose estimation with global signed distance fields. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  72. [80]

    Contactart: Learning 3D interaction priors for category-level ar- ticulated object and hand poses estimation

    Zehao Zhu, Jiashun Wang, Yuzhe Qin, Deqing Sun, Varun Jampani, and Xiaolong Wang. Contactart: Learning 3D interaction priors for category-level ar- ticulated object and hand poses estimation. InProc. Int. Conf. 3D Vis. (3DV), 2024

  73. [81]

    gSDF: Geometry-driven signed dis- tance functions for 3D hand-object reconstruction

    Zerui Chen, Shizhe Chen, Cordelia Schmid, and Ivan Laptev. gSDF: Geometry-driven signed dis- tance functions for 3D hand-object reconstruction. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023

  74. [82]

    Chord: Category-level hand-held object recon- struction via shape deformation

    Kailin Li, Lixin Yang, Haoyu Zhen, Zenan Lin, Xinyu Zhan, Licheng Zhong, Jian Xu, Kejian Wu, and Cewu Lu. Chord: Category-level hand-held object recon- struction via shape deformation. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023

  75. [83]

    Reconstructing hand-held objects from monoc- ular video

    Di Huang, Xiaopeng Ji, Xingyi He, Jiaming Sun, Tong He, Qing Shuai, Wanli Ouyang, and Xiaowei Zhou. Reconstructing hand-held objects from monoc- ular video. InProc. ACM SIGGRAPH Asia, 2022

  76. [84]

    Moho: Learning single-view hand-held object reconstruction with multi-view occlusion-aware su- pervision

    Chenyangguang Zhang, Guanlong Jiao, Yan Di, Gu Wang, Ziqin Huang, Ruida Zhang, Fabian Man- hardt, Bowen Fu, Federico Tombari, and Xiangyang Ji. Moho: Learning single-view hand-held object reconstruction with multi-view occlusion-aware su- pervision. InProc. IEEE/CVF Conf. Comp...

  77. [85]

    Texhoi: Reconstructing textures of 3D unknown ob- jects in monocular hand-object interaction scenes

    Alakh Aggarwal, Ningna Wang, and Xiaohu Guo. Texhoi: Reconstructing textures of 3D unknown ob- jects in monocular hand-object interaction scenes. IEEE Trans. Vis. Comput. Graph., 2025. 27 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation...

  78. [86]

    SeqHAND: RGB-sequence-based 3D hand pose and shape estimation

    John Yang, Hyung Jin Chang, Seungeui Lee, and Nojun Kwak. SeqHAND: RGB-sequence-based 3D hand pose and shape estimation. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020

  79. [87]

    Dyn-hamr: Recovering 4D interacting hand motion from a dynamic camera

    Zhengdi Yu, Stefanos Zafeiriou, and Tolga Birdal. Dyn-hamr: Recovering 4D interacting hand motion from a dynamic camera. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  80. [88]

    Interactionfusion: real-time reconstruction of hand poses and deformable objects in hand-object interac- tions.ACM Trans

    Hao Zhang, Zi-Hao Bo, Jun-Hai Yong, and Feng Xu. Interactionfusion: real-time reconstruction of hand poses and deformable objects in hand-object interac- tions.ACM Trans. Graph., 38(4), 2019

  81. [89]

    Interaction-aware 4D gaussian splat- ting for dynamic hand-object interaction reconstruc- tion.arXiv preprint arXiv:2511.14540, 2025

    Hao Tian, Chenyangguang Zhang, Rui Liu, Wen Shen, and Xiaolin Qin. Interaction-aware 4D gaussian splat- ting for dynamic hand-object interaction reconstruc- tion.arXiv preprint arXiv:2511.14540, 2025

  82. [90]

    Physics-aware hand-object interaction denoising

    Haowen Luo, Yunze Liu, and Li Yi. Physics-aware hand-object interaction denoising. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  83. [91]

    Grab: A dataset of whole-body hu- man grasping of objects

    Omid Taheri, Nima Ghorbani, Michael J Black, and Dimitrios Tzionas. Grab: A dataset of whole-body hu- man grasping of objects. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020

  84. [92]

    Hand-object contact consistency reason- ing for human grasps generation

    Hanwen Jiang, Shaowei Liu, Jiashun Wang, and Xiao- long Wang. Hand-object contact consistency reason- ing for human grasps generation. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021

  85. [93]

    Contactgen: Generative contact modeling for grasp generation

    Shaowei Liu, Yang Zhou, Jimei Yang, Saurabh Gupta, and Shenlong Wang. Contactgen: Generative contact modeling for grasp generation. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023

  86. [94]

    Contact2grasp: 3D grasp synthesis via hand-object contact constraint

    Haoming Li, Xinzhuo Lin, Yang Zhou, Xiang Li, Yuchi Huo, Jiming Chen, and Qi Ye. Contact2grasp: 3D grasp synthesis via hand-object contact constraint. InProc. Int. Joint Conf. Artif. Intell. (IJCAI), 2023. doi: 10.24963/IJCAI.2023/117

  87. [95]

    G-hop: Generative hand-object prior for interaction reconstruction and grasp synthesis

    Yufei Ye, Abhinav Gupta, Kris Kitani, and Shubham Tulsiani. G-hop: Generative hand-object prior for interaction reconstruction and grasp synthesis. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  88. [96]

    Clickdiff: click to induce semantic contact map for controllable grasp generation with diffusion models

    Peiming Li, Ziyi Wang, Mengyuan Liu, Hong Liu, and Chen Chen. Clickdiff: click to induce semantic contact map for controllable grasp generation with diffusion models. InProc. ACM Int. Conf. Multimedia (ACM MM), 2024

  89. [97]

    Fastgrasp: Efficient grasp synthesis with diffusion

    Xiaofei Wu, Tao Liu, Caoji Li, Yuexin Ma, Yujiao Shi, and Xuming He. Fastgrasp: Efficient grasp synthesis with diffusion. InProc. Int. Conf. 3D Vis. (3DV), 2025

  90. [98]

    Text2HOI: Text-guided 3D motion gener- ation for hand-object interaction

    Junuk Cha, Jihyeon Kim, Jae Shin Yoon, and Seun- gryul Baek. Text2HOI: Text-guided 3D motion gener- ation for hand-object interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  91. [99]

    Bimart: A unified approach for the synthesis of 3D bimanual in- teraction with articulated objects

    Wanyue Zhang, Rishabh Dabral, Vladislav Golyanik, Vasileios Choutas, Eduardo Alvarado, Thabo Beeler, Marc Habermann, and Christian Theobalt. Bimart: A unified approach for the synthesis of 3D bimanual in- teraction with articulated objects. InProc. IEEE/CVF Conf. Comput. Vis. ...

  92. [100]

    Manivideo: Generating hand-object ma- nipulation video with dexterous and generalizable grasping

    Youxin Pang, Ruizhi Shao, Jiajun Zhang, Hanzhang Tu, Yun Liu, Boyao Zhou, Hongwen Zhang, and Yebin Liu. Manivideo: Generating hand-object ma- nipulation video with dexterous and generalizable grasping. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2025

  93. [101]

    Hogsa: Bimanual hand-object interaction understanding with 3D gaussian splatting based data augmentation

    Wentian Qu, Jiahe Li, Jian Cheng, Jian Shi, Chenyu Meng, Cuixia Ma, Hongan Wang, Xiaoming Deng, and Yinda Zhang. Hogsa: Bimanual hand-object interaction understanding with 3D gaussian splatting based data augmentation. InProc. AAAI Conf. Artif. Intell., 2025

  94. [102]

    Gears: Local geometry-aware hand-object interaction synthesis

    Keyang Zhou, Bharat Lal Bhatnagar, Jan Eric Lenssen, and Gerard Pons-Moll. Gears: Local geometry-aware hand-object interaction synthesis. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  95. [103]

    SIGHT: synthesizing image-text conditioned and geometry-guided 3d hand-object trajectories.arXiv preprint arXiv:2503.22869, 2025

    Alexey Gavryushin, Alexandros Delitzas, Luc Van Gool, Marc Pollefeys, Kaichun Mo, and Xi Wang. SIGHT: synthesizing image-text conditioned and geometry-guided 3d hand-object trajectories.arXiv preprint arXiv:2503.22869, 2025

  96. [104]

    Gaze-guided hand-object interaction syn- thesis: Dataset and method.arXiv preprint arXiv:2403.16169, 2024

    Jie Tian, Ran Ji, Lingxiao Yang, Suting Ni, Yuexin Ma, Lan Xu, Jingyi Yu, Ye Shi, and Jingya Wang. Gaze-guided hand-object interaction syn- thesis: Dataset and method.arXiv preprint arXiv:2403.16169, 2024

  97. [105]

    How do i do that? synthesizing 3D hand motion and contacts for everyday interactions

    Aditya Prakash, Benjamin Lundell, Dmitry Andrey- chuk, David Forsyth, Saurabh Gupta, and Harpreet Sawhney. How do i do that? synthesizing 3D hand motion and contacts for everyday interactions. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025. 28 Hand-Object ...

  98. [106]

    Geohand: Unlocking prior geome- try knowledge for monocular 3D hand reconstruction

    Weiquan Lin, Yaoqing Hu, Liangchen Dai, Xu Tang, and Xingyu Chen. Geohand: Unlocking prior geome- try knowledge for monocular 3D hand reconstruction. arXiv preprint arXiv:2605.17354, 2026

  99. [107]

    MoGe-2: Accurate monoc- ular geometry with metric scale and sharp details

    Ruicheng Wang, Sicheng Xu, Yue Dong, Yu Deng, Jianfeng Xiang, Zelong Lv, Guangzhong Sun, Xin Tong, and Jiaolong Yang. MoGe-2: Accurate monoc- ular geometry with metric scale and sharp details. In Adv. Neural Inf. Process. Syst. (NeurIPS), volume 38, 2025

  100. [108]

    Lisa: Rea- soning segmentation via large language model

    Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. Lisa: Rea- soning segmentation via large language model. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  101. [109]

    Affordance diffusion: Synthe- sizing hand-object interactions

    Yufei Ye, Xueting Li, Abhinav Gupta, Shalini De Mello, Stan Birchfield, Jiaming Song, Shubham Tulsiani, and Sifei Liu. Affordance diffusion: Synthe- sizing hand-object interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023

  102. [110]

    Genie: Text-to-3D generation

    Luma AI. Genie: Text-to-3D generation. https:// lumalabs.ai/, 2024. Accessed: Dec. 15, 2024

  103. [111]

    ChatGPT, 2022

    OpenAI. ChatGPT, 2022. URL https://openai. com/blog/chatgpt/. Accessed: 2026-03-24

  104. [112]

    GPT-4V(ision) system card

    OpenAI. GPT-4V(ision) system card. Technical report, OpenAI, 2023. URL https://openai. com/index/gpt-4v-system-card/ . Ac- cessed: Dec. 15, 2024

  105. [113]

    Follow my hold: Hand-object in- teraction reconstruction through geometric guidance

    Ayce Idil Aytekin, Helge Rhodin, Rishabh Dabral, and Christian Theobalt. Follow my hold: Hand-object in- teraction reconstruction through geometric guidance. InProc. Int. Conf. 3D Vis. (3DV), 2026

  106. [114]

    Hunyuan3d 2.5: Towards high-fidelity 3D assets generation with ultimate details.arXiv preprint arXiv:2506.16504, 2025

    Zeqiang Lai, Yunfei Zhao, Haolin Liu, Zibo Zhao, Qingxiang Lin, Huiwen Shi, Xianghui Yang, Mingxin Yang, Shuhui Yang, Yifei Feng, et al. Hunyuan3d 2.5: Towards high-fidelity 3D assets generation with ultimate details.arXiv preprint arXiv:2506.16504, 2025

  107. [115]

    Host3r: Keypoint-free hand-object 3D reconstruction from RGB images

    Anilkumar Swamy, Vincent Leroy, Philippe Wein- zaepfel, Jean-S´ebastien Franco, and Gr´egory Rogez. Host3r: Keypoint-free hand-object 3D reconstruction from RGB images. InProc. IEEE/CVF Int. Conf. Comput. Vis. Workshops (ICCVW), 2025

  108. [116]

    Video depth anything: Consistent depth estimation for super- long videos

    Sili Chen, Hengkai Guo, Shengnan Zhu, Feihu Zhang, Zilong Huang, Jiashi Feng, and Bingyi Kang. Video depth anything: Consistent depth estimation for super- long videos. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  109. [117]

    Unidepthv2: Universal monocular metric depth estimation made simpler.IEEE Trans

    Luigi Piccinelli, Christos Sakaridis, Yung-Hsu Yang, Mattia Segu, Siyuan Li, Wim Abbeloos, and Luc Van Gool. Unidepthv2: Universal monocular metric depth estimation made simpler.IEEE Trans. Pattern Anal. Mach. Intell., 2025

  110. [118]

    SAM 2: Segment anything in im- ages and videos

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Rong- hang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R ¨adle, Chloe Rolland, Laura Gustafson, et al. SAM 2: Segment anything in im- ages and videos. InProc. Int. Conf. Learn. Represent. (ICLR), 2025

  111. [119]

    Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (...

  112. [120]

    Hawor: World-space hand motion reconstruction from egocentric videos

    Jinglei Zhang, Jiankang Deng, Chao Ma, and Rolan- dos Alexandros Potamias. Hawor: World-space hand motion reconstruction from egocentric videos. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  113. [121]

    Metric3d: Towards zero-shot metric 3D prediction from a single image

    Wei Yin, Chi Zhang, Hao Chen, Zhipeng Cai, Gang Yu, Kaixuan Wang, Xiaozhi Chen, and Chun Shen. Metric3d: Towards zero-shot metric 3D prediction from a single image. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023

  114. [122]

    SAM 3D: 3Dfy anything in images

    Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang, Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, et al. SAM 3D: 3Dfy anything in images. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026

  115. [123]

    MV-SAM3D: Adaptive multi-view fusion for layout-aware 3D gen- eration.arXiv preprint arXiv:2603.11633, 2026

    Baicheng Li, Dong Wu, Jun Li, Shunkai Zhou, Zecui Zeng, Lusong Li, and Hongbin Zha. MV-SAM3D: Adaptive multi-view fusion for layout-aware 3D gen- eration.arXiv preprint arXiv:2603.11633, 2026

  116. [124]

    SAM 3: Segment anything with concepts.arXiv preprint arXiv:2511.16719, 2025

    Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chai- tanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al. SAM 3: Segment anything with concepts.arXiv preprint arXiv:2511.16719, 2025

  117. [125]

    Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647, 2025

    Haotong Lin, Sili Chen, Junhao Liew, Donny Y Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647, 2025

  118. [126]

    Handos: 3D hand 29 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer reconstruction in one stage

    Xingyu Chen, Zhuheng Song, Xiaoke Jiang, Yaoqing Hu, Junzhi Yu, and Lei Zhang. Handos: 3D hand 29 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer reconstruction in one stage. InProc. IEEE/CVF Conf. Comput. Vis. P...

  119. [127]

    Choir: Contact-aware 4D hand- object interaction reconstruction.arXiv preprint arXiv:2605.20992, 2026

    Hao Xu, Yilin Liu, Yinqiao Wang, Chi-Wing Fu, and Niloy J Mitra. Choir: Contact-aware 4D hand- object interaction reconstruction.arXiv preprint arXiv:2605.20992, 2026

  120. [128]

    Using diffusion priors for video amodal segmenta- tion

    Kaihua Chen, Deva Ramanan, and Tarasha Khurana. Using diffusion priors for video amodal segmenta- tion. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  121. [129]

    Affordgrasp: Cross-modal diffusion for affordance-aware grasp synthesis

    Xiaofei Wu, Yi Zhang, Yumeng Liu, Yuexin Ma, Yu- jiao Shi, and Xuming He. Affordgrasp: Cross-modal diffusion for affordance-aware grasp synthesis. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026

  122. [130]

    Roberta: A robustly optimized bert pretraining approach.arXiv preprint arXiv:1907.11692, 2019

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach.arXiv preprint arXiv:1907.11692, 2019

  123. [131]

    Qwen2 technical re- port.arXiv preprint arXiv:2407.10671, 2024

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze...

  124. [132]

    G-dexgrasp: General- izable dexterous grasping synthesis via part-aware prior retrieval and prior-assisted generation

    Juntao Jian, Xiuping Liu, Zixuan Chen, Manyi Li, Jian Liu, and Ruizhen Hu. G-dexgrasp: General- izable dexterous grasping synthesis via part-aware prior retrieval and prior-assisted generation. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025

  125. [133]

    GPT-4o system card

    OpenAI. GPT-4o system card. Technical report, Ope- nAI, 2024. URL https://cdn.openai.com/ gpt-4o-system-card.pdf

  126. [134]

    Grounded language-image pre-training

    Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al. Grounded language-image pre-training. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  127. [135]

    NL2Contact: Natural language guided 3D hand- object contact modeling with diffusion model

    Zhongqun Zhang, Hengfei Wang, Ziwei Yu, Yi- hua Cheng, Angela Yao, and Hyung Jin Chang. NL2Contact: Natural language guided 3D hand- object contact modeling with diffusion model. In Proc. Eur. Conf. Comput. Vis. (ECCV), 2024

  128. [136]

    Bert: Pre-training of deep bidi- rectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidi- rectional transformers for language understanding. In Proc. Conf. North Amer. Chapter Assoc. Comput. Lin- guistics: Human Lang. Technol. (NAACL-HLT), 2019

  129. [137]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT qual- ity, March 2023. URL https://lmsys.or...

  130. [138]

    GPT-4 technical report.arXiv preprint arXiv:2303.08774, 2023

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. GPT-4 technical report.arXiv preprint arXiv:2303.08774, 2023

  131. [139]

    HOIGPT: Learning long-sequence hand-object inter- action with language models

    Mingzhen Huang, Fu-Jen Chu, Bugra Tekin, Kevin J Liang, Haoyu Ma, Weiyao Wang, Xingyu Chen, Pierre Gleize, Hongfei Xue, Siwei Lyu, et al. HOIGPT: Learning long-sequence hand-object inter- action with language models. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  132. [140]

    Llama: Open and efficient foundation language mod- els.arXiv preprint arXiv:2302.13971, 2023

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth ´ee Lacroix, Baptiste Rozi`ere, Naman Goyal, Eric Ham- bro, Faisal Azhar, Aur ´elien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation l...

  133. [141]

    OpenHOI: Open-world hand-object interaction synthesis with multimodal large language model

    Zhenhao Zhang, Ye Shi, Lingxiao Yang, Suting Ni, Qi Ye, and Jingya Wang. OpenHOI: Open-world hand-object interaction synthesis with multimodal large language model. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 38, 2025

  134. [142]

    ShapeLLM: Universal 3D object understanding for embodied interaction

    Zekun Qi, Runpei Dong, Shaochen Zhang, Haoran Geng, Chunrui Han, Zheng Ge, Li Yi, and Kaisheng Ma. ShapeLLM: Universal 3D object understanding for embodied interaction. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024. 30 Hand-Object Interaction in the Age of Large Foundation Mode...

  135. [143]

    Diffh2o: Diffusion-based syn- thesis of hand-object interactions from textual descrip- tions

    Sammy Christen, Shreyas Hampali, Fadime Sener, Edoardo Remelli, Tomas Hodan, Eric Sauser, Shugao Ma, and Bugra Tekin. Diffh2o: Diffusion-based syn- thesis of hand-object interactions from textual descrip- tions. InProc. ACM SIGGRAPH Asia, 2024

  136. [144]

    Forehoi: Feed-forward 3D object reconstruc- tion from daily hand-object interaction videos

    Yuantao Chen, Jiahao Chang, Chongjie Ye, Chaoran Zhang, Zhaojie Fang, Chenghong Li, and Xiaoguang Han. Forehoi: Feed-forward 3D object reconstruc- tion from daily hand-object interaction videos. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026

  137. [145]

    Zero- 1-to-3: Zero-shot one image to 3D object

    Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl V ondrick. Zero- 1-to-3: Zero-shot one image to 3D object. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023

  138. [146]

    Bigs: Bimanual category-agnostic interaction reconstruction from monocular videos via 3D gaussian splatting

    Jeongwan On, Kyeonghwan Gwak, Gunyoung Kang, Junuk Cha, Soohyun Hwang, Hyein Hwang, and Se- ungryul Baek. Bigs: Bimanual category-agnostic interaction reconstruction from monocular videos via 3D gaussian splatting. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2025

  139. [147]

    Human universal grasping.arXiv preprint arXiv:2606.17054, 2026

    Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, and Lerrel Pinto. Human universal grasping.arXiv preprint arXiv:2606.17054, 2026

  140. [148]

    Hvg-3D: Bridging real and simulation domains for 3D-conditional hand-object interaction video synthesis

    Mingjin Chen, Junhao Chen, Zhaoxin Fan, Yujian Lee, Zichen Dang, Lili Wang, Yawen Cui, Lap-Pui Chau, and Yi Wang. Hvg-3D: Bridging real and simulation domains for 3D-conditional hand-object interaction video synthesis. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026

  141. [149]

    Colin Raffel, Noam Shazeer, Adam Roberts, Kather- ine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of trans- fer learning with a unified text-to-text transformer.J. Mach. Learn. Res., 21(140), 2020

  142. [150]

    Re-hold: Video hand object interaction reenactment via adaptive layout- instructed diffusion model

    Yingying Fan, Quanwei Yang, Kaisiyuan Wang, Hang Zhou, Yingying Li, Haocheng Feng, Errui Ding, Yu Wu, and Jingdong Wang. Re-hold: Video hand object interaction reenactment via adaptive layout- instructed diffusion model. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (C...

  143. [151]

    GLIDE: To- wards photorealistic image generation and editing with text-guided diffusion models

    Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob Mc- Grew, Ilya Sutskever, and Mark Chen. GLIDE: To- wards photorealistic image generation and editing with text-guided diffusion models. InProc. Int. Conf. Mach. Learn. (ICML), 2022

  144. [152]

    Agile: Hand-object interaction re- construction from video via agentic generation.arXiv preprint arXiv:2602.04672, 2026

    Jin-Chuan Shi, Binhong Ye, Tao Liu, Junzhe He, Yangjinhui Xu, Xiaoyang Liu, Zeju Li, Hao Chen, and Chunhua Shen. Agile: Hand-object interaction re- construction from video via agentic generation.arXiv preprint arXiv:2602.04672, 2026

  145. [153]

    VGGT: Visual geometry grounded trans- former

    Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. VGGT: Visual geometry grounded trans- former. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2025

  146. [154]

    Efros, and Angjoo Kanazawa

    Qianqian Wang, Yifei Zhang, Aleksander Holynski, Alexei A. Efros, and Angjoo Kanazawa. Continu- ous 3D perception model with persistent state. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  147. [155]

    Hggt: Robust and flexi- ble 3D hand mesh reconstruction from uncalibrated images.arXiv preprint arXiv:2603.23997, 2026

    Yumeng Liu, Xiao-Xiao Long, Marc Habermann, Xu- anze Yang, Cheng Lin, Yuan Liu, Yuexin Ma, Wen- ping Wang, and Ligang Liu. Hggt: Robust and flexi- ble 3D hand mesh reconstruction from uncalibrated images.arXiv preprint arXiv:2603.23997, 2026

  148. [156]

    Egograsp: World-space hand-object interaction estimation from egocentric videos.arXiv preprint arXiv:2601.01050, 2026

    Hongming Fu, Wenjia Wang, Xiaozhen Qiao, Rolan- dos Alexandros Potamias, Taku Komura, Shuo Yang, Zheng Liu, and Bo Zhao. Egograsp: World-space hand-object interaction estimation from egocentric videos.arXiv preprint arXiv:2601.01050, 2026

  149. [157]

    Hand3r: Online 4D hand-scene reconstruction in the wild.arXiv preprint arXiv:2602.03200, 2026

    Wendi Hu, Haonan Zhou, Wenhao Hu, and Gaoang Wang. Hand3r: Online 4D hand-scene reconstruction in the wild.arXiv preprint arXiv:2602.03200, 2026

  150. [158]

    ScaleHP: Esti- mating hand pose in metric space.arXiv preprint arXiv:2606.25619, 2026

    Ruitao Jing, Xingyu Chen, Hongyang Li, Qing Jiang, Yukai Shi, and Lei Zhang. ScaleHP: Esti- mating hand pose in metric space.arXiv preprint arXiv:2606.25619, 2026

  151. [159]

    Qwen technical report.arXiv preprint arXiv:2309.16609, 2023

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report.arXiv preprint arXiv:2309.16609, 2023

  152. [160]

    Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learn- ers. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 33, 2020

  153. [161]

    Multi-graspllm: A multimodal LLM for multi-hand semantic guided grasp generation.arXiv preprint arXiv:2412.08468, 2024

    Haosheng Li, Weixin Mao, Weipeng Deng, Chenyu Meng, Haoqiang Fan, Tiancai Wang, Yoshie Os- amu, Ping Tan, Hongan Wang, and Xiaoming Deng. Multi-graspllm: A multimodal LLM for multi-hand semantic guided grasp generation.arXiv preprint arXiv:2412.08468, 2024. 31 Hand-Object Inte...

  154. [162]

    Afford- dexgrasp: Open-set language-guided dexterous grasp with generalizable-instructive affordance

    Yi-Lin Wei, Mu Lin, Yuhao Lin, Jian-Jian Jiang, Xiao- Ming Wu, Ling-An Zeng, and Wei-Shi Zheng. Afford- dexgrasp: Open-set language-guided dexterous grasp with generalizable-instructive affordance. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025

  155. [163]

    Jeni, and Junhyug Noh

    Mingyeong Song, Jungbin Cho, Jisoo Kim, Ananya Bal, Kartik Sharma, Youngjae Yu, Laszlo A. Jeni, and Junhyug Noh. Jointhoi: Jointly generating contact maps enhances hand object interaction generation. arXiv preprint arXiv:2607.01768, 2026

  156. [164]

    Structbihoi: Structured articulation modeling for long–horizon bimanual hand–object interaction generation.arXiv preprint arXiv:2603.08390, 2026

    Zhi Wang, Liu Liu, Ruonan Liu, Dan Guo, and Meng Wang. Structbihoi: Structured articulation modeling for long–horizon bimanual hand–object interaction generation.arXiv preprint arXiv:2603.08390, 2026

  157. [165]

    Synhlma: Synthesizing hand language manipulation for articulated object with discrete hu- man object interaction representation.arXiv preprint arXiv:2510.25268, 2025

    Zhi Wang, Yuyan Liu, Liu Liu, Li Zhang, Ruixuan Lu, and Dan Guo. Synhlma: Synthesizing hand language manipulation for articulated object with discrete hu- man object interaction representation.arXiv preprint arXiv:2510.25268, 2025

  158. [166]

    Touch: Text-guided controllable gen- eration of free-form hand-object interactions.arXiv preprint arXiv:2510.14874, 2025

    Guangyi Han, Wei Zhai, Yuhang Yang, Yang Cao, and Zheng-Jun Zha. Touch: Text-guided controllable gen- eration of free-form hand-object interactions.arXiv preprint arXiv:2510.14874, 2025

  159. [167]

    Megohand: Multimodal egocentric hand-object interaction motion generation

    Bohan Zhou, Yi Zhan, Zhongbin Zhang, and Zongqing Lu. Megohand: Multimodal egocentric hand-object interaction motion generation. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 38, 2025

  160. [168]

    Eagle 2: Building post-training data strategies from scratch for frontier vision-language models.arXiv preprint arXiv:2501.14818, 2025

    Zhiqi Li, Guo Chen, Shilong Liu, Shihao Wang, VS Vibashan, Yishen Ji, Shiyi Lan, Hao Zhang, Yilin Zhao, Subhashree Radhakrishnan, et al. Eagle 2: Building post-training data strategies from scratch for frontier vision-language models.arXiv preprint arXiv:2501.14818, 2025

  161. [169]

    Training- free dense hand contact estimation with multi- modal large language models.arXiv preprint arXiv:2605.05886, 2026

    Daniel Sungho Jung and Kyoung Mu Lee. Training- free dense hand contact estimation with multi- modal large language models.arXiv preprint arXiv:2605.05886, 2026

  162. [170]

    Affordance-guided diffusion prior for 3D hand re- construction.arXiv preprint arXiv:2510.00506, 2025

    Naru Suzuki, Takehiko Ohkawa, Tatsuro Banno, Jihyun Lee, Ryosuke Furuta, and Yoichi Sato. Affordance-guided diffusion prior for 3D hand re- construction.arXiv preprint arXiv:2510.00506, 2025

  163. [171]

    Hort: Monocular hand- held objects reconstruction with transformers

    Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, and Cordelia Schmid. Hort: Monocular hand- held objects reconstruction with transformers. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025

  164. [172]

    Grasp as you dream: Imitating functional grasping from generated human demonstrations.arXiv preprint arXiv:2604.07517, 2026

    Chao Tang, Jiacheng Xu, Haofei Lu, Bolin Zou, Wen- long Dong, Hong Zhang, and Danica Kragic. Grasp as you dream: Imitating functional grasping from generated human demonstrations.arXiv preprint arXiv:2604.07517, 2026

  165. [173]

    Web2grasp: Learning functional grasps from web images of hand-object interactions.arXiv preprint arXiv:2505.05517, 2025

    Hongyi Chen, Yunchao Yao, Yufei Ye, Zhixuan Xu, Homanga Bharadhwaj, Jiashun Wang, Shubham Tulsiani, Zackory Erickson, and Jeffrey Ichnowski. Web2grasp: Learning functional grasps from web images of hand-object interactions.arXiv preprint arXiv:2505.05517, 2025

  166. [174]

    Sdxl: Improving latent diffusion models for high-resolution image synthesis

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. InProc. Int. Conf. Learn. Represent. (ICLR), 2024

  167. [175]

    FLUX.1 kontext: Flow matching for in- context image generation and editing in latent space

    Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne, Tim Dockhorn, Jack English, Zion English, Patrick Esser, Sumith Kulal, Kyle Lacey, Yam Levi, Cheng Li, Dominik Lorenz, Jonas M ¨uller, Dustin Podell, Robin Rombach, Harry Sain...

  168. [176]

    Adding conditional control to text-to-image diffusion models

    Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023

  169. [177]

    Single-view image to novel- view generation for hand-object interactions

    Zhongqun Zhang, Yihua Cheng, Eduardo P ´erez- Pellitero, Yiren Zhou, Jiankang Deng, Hyung Jin Chang, and Jifei Song. Single-view image to novel- view generation for hand-object interactions. InProc. AAAI Conf. Artif. Intell., 2025

  170. [178]

    Objaverse-XL: A universe of 10m+ 3D objects

    Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram V oleti, Samir Yitzhak Gadre, et al. Objaverse-XL: A universe of 10m+ 3D objects. InAdv. Neural Inf. Process. Syst. (NeurIPS), Datasets and Benchmarks Tr...

  171. [179]

    Prompt-propose- verify: A reliable hand-object-interaction data gener- ation framework using foundational models.arXiv preprint arXiv:2312.15247, 2023

    Gurusha Juneja and Sukrit Kumar. Prompt-propose- verify: A reliable hand-object-interaction data gener- ation framework using foundational models.arXiv preprint arXiv:2312.15247, 2023

  172. [180]

    Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

    Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. InProc. 32 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generat...

  173. [181]

    Hand1000: Generating realistic hands from text with only 1,000 images

    Haozhuo Zhang, Bin Zhu, Yu Cao, and Yanbin Hao. Hand1000: Generating realistic hands from text with only 1,000 images. InProc. AAAI Conf. Artif. Intell., 2025

  174. [182]

    Rhands: Re- fining malformed hands for generated images with de- coupled structure and style guidance

    Chengrui Wang, Pengfei Liu, Min Zhou, Ming Zeng, Xubin Li, Tiezheng Ge, and Bo Zheng. Rhands: Re- fining malformed hands for generated images with de- coupled structure and style guidance. InProc. AAAI Conf. Artif. Intell., 2025

  175. [183]

    At- tentionhand: Text-driven controllable hand image gen- eration for 3D hand reconstruction in the wild

    Junho Park, Kyeongbo Kong, and Suk-Ju Kang. At- tentionhand: Text-driven controllable hand image gen- eration for 3D hand reconstruction in the wild. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024

  176. [184]

    Dynamicrafter: An- imating open-domain images with video diffusion priors

    Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang, Ying Shan, and Tien-Tsin Wong. Dynamicrafter: An- imating open-domain images with video diffusion priors. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024

  177. [185]

    Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025

    Wan Team, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025

  178. [186]

    Taste-rob: Advancing video generation of task-oriented hand- object interaction for generalizable robotic manipula- tion

    Hongxiang Zhao, Xingchen Liu, Mutian Xu, Yiming Hao, Weikai Chen, and Xiaoguang Han. Taste-rob: Advancing video generation of task-oriented hand- object interaction for generalizable robotic manipula- tion. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  179. [187]

    idit- HOI: Inpainting-based hand object interaction reenact- ment via video diffusion transformer.arXiv preprint arXiv:2506.12847, 2025

    Zhelun Shen, Chenming Wu, Junsheng Zhou, Chen Zhao, Kaisiyuan Wang, Hang Zhou, Yingying Li, Haocheng Feng, Wei He, and Jingdong Wang. idit- HOI: Inpainting-based hand object interaction reenact- ment via video diffusion transformer.arXiv preprint arXiv:2506.12847, 2025

  180. [188]

    Open-world hand- object interaction video generation based on structure and contact-aware representation

    Haodong Yan, Hang Yu, Zhide Zhong, Weilin Yuan, Xin Gong, Zehang Luo, Chengxi Heyu, Junfeng Li, Wenxuan Song, Shunbo Zhou, et al. Open-world hand- object interaction video generation based on structure and contact-aware representation. InProc. IEEE/CVF Conf. Comput. Vis. Patte...

  181. [189]

    PAM: A pose-appearance-motion engine for sim-to-real HOI video generation.arXiv preprint arXiv:2603.22193, 2026

    Mingze Gao, Kai Yang, Hongbo Gao, Bo Li, Aoxiang Ding, Wenxiang Li, Yujie Yu, Jian Liu, Shugong Xu, Yi Niu, Haoyu Chi, He Chen, Hui Tang, Li Yi, and Hao Zhao. PAM: A pose-appearance-motion engine for sim-to-real HOI video generation.arXiv preprint arXiv:2603.22193, 2026

  182. [190]

    Egocentric world model for photorealis- tic hand-object interaction synthesis.arXiv preprint arXiv:2603.13615, 2026

    Dayou Li, Lulin Liu, Bangya Liu, Shijie Zhou, Jiu Feng, Ziqi Lu, Minghui Zheng, Chenyu You, and Zhiwen Fan. Egocentric world model for photorealis- tic hand-object interaction synthesis.arXiv preprint arXiv:2603.13615, 2026

  183. [191]

    Hand2world: Autoregressive egocentric interaction generation via free-space hand gestures.arXiv preprint arXiv:2602.09600, 2026

    Yuxi Wang, Wenqi Ouyang, Tianyi Wei, Yi Dong, Zhiqi Shen, and Xingang Pan. Hand2world: Autoregressive egocentric interaction generation via free-space hand gestures.arXiv preprint arXiv:2602.09600, 2026

  184. [192]

    Generated real- ity: Human-centric world simulation using interactive video generation with hand and camera control

    Linxi Xie, Lisong C Sun, Ashley Neall, Tong Wu, Shengqu Cai, and Gordon Wetzstein. Generated real- ity: Human-centric world simulation using interactive video generation with hand and camera control. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026

  185. [193]

    Dexterous world models

    Byungjun Kim, Taeksoo Kim, Junyoung Lee, and Hanbyul Joo. Dexterous world models. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026

  186. [194]

    Handsonworld: Unconstrained egocentric video generation with camera-disentangled hand control.arXiv preprint arXiv:2607.02075, 2026

    Yushuo Chen, Xiaoyu Shi, Xiaoshi Wu, Xintao Wang, Pengfei Wan, and Yebin Liu. Handsonworld: Unconstrained egocentric video generation with camera-disentangled hand control.arXiv preprint arXiv:2607.02075, 2026

  187. [195]

    Wh0: Generative world models as scalable sources of ego- centric human hand manipulation data.arXiv preprint arXiv:2606.22136, 2026

    Yangtao Chen, Zixuan Chen, Peiyang Wang, Yong- Lu Li, Jing Huo, Jieqi Shi, and Yang Gao. Wh0: Generative world models as scalable sources of ego- centric human hand manipulation data.arXiv preprint arXiv:2606.22136, 2026

  188. [196]

    Moto: Latent motion token as the bridging language for learning robot manipulation from videos

    Yi Chen, Yuying Ge, Weiliang Tang, Yizhuo Li, Yix- iao Ge, Mingyu Ding, Ying Shan, and Xihui Liu. Moto: Latent motion token as the bridging language for learning robot manipulation from videos. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025

  189. [197]

    Igor: Image-goal representations are the atomic control units for foundation models in embodied ai.arXiv preprint arXiv:2411.00785, 2024

    Xiaoyu Chen, Junliang Guo, Tianyu He, Chuheng Zhang, Pushi Zhang, Derek Cathera Yang, Li Zhao, and Jiang Bian. Igor: Image-goal representations are the atomic control units for foundation models in embodied ai.arXiv preprint arXiv:2411.00785, 2024

  190. [198]

    In-n-on: Scaling egocentric manipulation with in-the-wild and on-task data.arXiv preprint arXiv:2511.15704, 2025

    Xiongyi Cai, Ri-Zhao Qiu, Geng Chen, Lai Wei, Is- abella Liu, Tianshu Huang, Xuxin Cheng, and Xiao- long Wang. In-n-on: Scaling egocentric manipulation with in-the-wild and on-task data.arXiv preprint arXiv:2511.15704, 2025. 33 Hand-Object Interaction in the Age of Large Found...

  191. [199]

    Emergence of human to robot trans- fer in vision-language-action models.arXiv preprint arXiv:2512.22414, 2025

    Simar Kareer, Karl Pertsch, James Darpinian, Judy Hoffman, Danfei Xu, Sergey Levine, Chelsea Finn, and Suraj Nair. Emergence of human to robot trans- fer in vision-language-action models.arXiv preprint arXiv:2512.22414, 2025

  192. [200]

    Clap: Contrastive latent action pretraining for learning vision-language-action models from human videos.arXiv preprint arXiv:2601.04061, 2026

    Chubin Zhang, Jianan Wang, Zifeng Gao, Yue Su, Tianru Dai, Cai Zhou, Jiwen Lu, and Yansong Tang. Clap: Contrastive latent action pretraining for learning vision-language-action models from human videos.arXiv preprint arXiv:2601.04061, 2026

  193. [201]

    Univla: Learning to act anywhere with task-centric latent actions.arXiv preprint arXiv:2505.06111, 2025

    Qingwen Bu, Yanting Yang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Yao, Ping Luo, and Hongyang Li. Univla: Learning to act anywhere with task-centric latent actions.arXiv preprint arXiv:2505.06111, 2025

  194. [202]

    Villa-x: enhancing latent action modeling in vision-language- action models.arXiv preprint arXiv:2507.23682, 2025

    Xiaoyu Chen, Hangxing Wei, Pushi Zhang, Chuheng Zhang, Kaixin Wang, Yanjiang Guo, Rushuai Yang, Yucen Wang, Xinquan Xiao, Li Zhao, et al. Villa-x: enhancing latent action modeling in vision-language- action models.arXiv preprint arXiv:2507.23682, 2025

  195. [203]

    Unleashing large-scale video gen- erative pre-training for visual robot manipulation

    Hongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen, Jiafeng Xu, Xinghang Li, Minghuan Liu, Hang Li, and Tao Kong. Unleashing large-scale video gen- erative pre-training for visual robot manipulation. In Proc. Int. Conf. Learn. Represent. (ICLR), 2024

  196. [204]

    Gr-2: A gen- erative video-language-action model with web-scale knowledge for robot manipulation.arXiv preprint arXiv:2410.06158, 2024

    Chi-Lam Cheang, Guangzeng Chen, Ya Jing, Tao Kong, Hang Li, Yifeng Li, Yuxiao Liu, Hongtao Wu, Jiafeng Xu, Yichu Yang, et al. Gr-2: A gen- erative video-language-action model with web-scale knowledge for robot manipulation.arXiv preprint arXiv:2410.06158, 2024

  197. [205]

    GR00T n1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734, 2025

    Johan Bjorck, Fernando Casta ˜neda, Nikita Cherni- adev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, et al. GR00T n1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734, 2025

  198. [206]

    Motus: A unified latent action world model

    Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, et al. Motus: A unified latent action world model. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026

  199. [207]

    Egoscale: Scaling dexterous manipulation with diverse egocentric human data.arXiv preprint arXiv:2602.16710, 2026

    Ruijie Zheng, Dantong Niu, Yuqi Xie, Jing Wang, Mengda Xu, Yunfan Jiang, Fernando Casta ˜neda, Fengyuan Hu, You Liang Tan, Letian Fu, et al. Egoscale: Scaling dexterous manipulation with diverse egocentric human data.arXiv preprint arXiv:2602.16710, 2026

  200. [208]

    Being-h0: vision- language-action pretraining from large-scale human videos.arXiv preprint arXiv:2507.15597, 2025

    Hao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng, Ye Wang, Haoqi Yuan, Jiazheng Liu, Chaoyi Xu, Qin Jin, and Zongqing Lu. Being-h0: vision- language-action pretraining from large-scale human videos.arXiv preprint arXiv:2507.15597, 2025

  201. [209]

    Scalable vision-language- action model pretraining for robotic manipulation with real-life human activity videos.arXiv preprint arXiv:2510.21571, 2025

    Qixiu Li, Yu Deng, Yaobo Liang, Lin Luo, Lei Zhou, Chengtang Yao, Lingqi Zeng, Zhiyuan Feng, Huizhi Liang, Sicheng Xu, et al. Scalable vision-language- action model pretraining for robotic manipulation with real-life human activity videos.arXiv preprint arXiv:2510.21571, 2025

  202. [210]

    Developing vision-language- action model from egocentric videos.arXiv preprint arXiv:2509.21986, 2025

    Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, and Shinsuke Mori. Developing vision-language- action model from egocentric videos.arXiv preprint arXiv:2509.21986, 2025

  203. [211]

    H-rdt: Human manipulation enhanced bimanual robotic ma- nipulation

    Hongzhe Bi, Lingxuan Wu, Tianwei Lin, Hengkai Tan, Zhizhong Su, Hang Su, and Jun Zhu. H-rdt: Human manipulation enhanced bimanual robotic ma- nipulation. InProc. AAAI Conf. Artif. Intell., 2026

  204. [212]

    Unihm: Unified dexterous hand manipula- tion with vision language model.arXiv preprint arXiv:2603.00732, 2026

    Zhenhao Zhang, Jiaxin Liu, Ye Shi, and Jingya Wang. Unihm: Unified dexterous hand manipula- tion with vision language model.arXiv preprint arXiv:2603.00732, 2026

  205. [213]

    Maniptrans: Efficient dexterous bi- manual manipulation transfer via residual learning

    Kailin Li, Puhao Li, Tengyu Liu, Yuyang Li, and Siyuan Huang. Maniptrans: Efficient dexterous bi- manual manipulation transfer via residual learning. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  206. [214]

    Dexmachina: Functional retargeting for bimanual dexterous manip- ulation.arXiv preprint arXiv:2505.24853, 2025

    Zhao Mandi, Yifan Hou, Dieter Fox, Yashraj Narang, Ajay Mandlekar, and Shuran Song. Dexmachina: Functional retargeting for bimanual dexterous manip- ulation.arXiv preprint arXiv:2505.24853, 2025

  207. [215]

    Learning dexterous manipulation using contact wrench guidance from hu- man demonstration.arXiv preprint arXiv:2607.00033, 2026

    Xinghao Zhu, Zixi Liu, Shalin Jain, Chenran Li, Mi- lad Noori, Huihua Zhao, John Welsh, Michael Andres Lin, Wei Liu, Tingwu Wang, et al. Learning dexterous manipulation using contact wrench guidance from hu- man demonstration.arXiv preprint arXiv:2607.00033, 2026

  208. [216]

    Egomimic: Scaling imitation learning via egocentric video

    Simar Kareer, Dhruv Patel, Ryan Punamiya, Pranay Mathur, Shuo Cheng, Chen Wang, Judy Hoffman, and Danfei Xu. Egomimic: Scaling imitation learning via egocentric video. InProc. IEEE Int. Conf. Robot. Autom. (ICRA), 2025

  209. [217]

    Dex- umi: Using human hand as the universal manipulation interface for dexterous manipulation.arXiv preprint arXiv:2505.21864, 2025

    Mengda Xu, Han Zhang, Yifan Hou, Zhenjia Xu, Linxi Fan, Manuela Veloso, and Shuran Song. Dex- umi: Using human hand as the universal manipulation interface for dexterous manipulation.arXiv preprint arXiv:2505.21864, 2025. 34 Hand-Object Interaction in the Age of Large Foundati...

  210. [218]

    Object-centric dexterous manipu- lation from human motion data.arXiv preprint arXiv:2411.04005, 2024

    Yuanpei Chen, Chen Wang, Yaodong Yang, and C Karen Liu. Object-centric dexterous manipu- lation from human motion data.arXiv preprint arXiv:2411.04005, 2024

  211. [219]

    Dexterous manipulation policies from RGB human videos via 3D hand-object trajectory recon- struction.arXiv preprint arXiv:2602.09013, 2026

    Hongyi Chen, Tony Dong, Tiancheng Wu, Liquan Wang, Yash Jangir, Yaru Niu, Yufei Ye, Homanga Bharadhwaj, Zackory Erickson, and Jeffrey Ich- nowski. Dexterous manipulation policies from RGB human videos via 3D hand-object trajectory recon- struction.arXiv preprint arXiv:2602.09013, 2026

  212. [220]

    Deximit: Learning bimanual dex- terous manipulation from monocular human videos

    Juncheng Mu, Sizhe Yang, Yiming Bao, Hojin Bae, Tianming Wei, Linning Xu, Boyi Li, Huazhe Xu, and Jiangmiao Pang. Deximit: Learning bimanual dex- terous manipulation from monocular human videos. arXiv preprint arXiv:2602.10105, 2026

  213. [221]

    Masquerade: Learning from in-the-wild human videos using data-editing.arXiv preprint arXiv:2508.09976, 2025

    Marion Lepert, Jiaying Fang, and Jeannette Bohg. Masquerade: Learning from in-the-wild human videos using data-editing.arXiv preprint arXiv:2508.09976, 2025

  214. [222]

    You only teach once: Learn one-shot bimanual robotic manipulation from video demonstrations.arXiv preprint arXiv:2501.14208, 2025

    Huayi Zhou, Ruixiang Wang, Yunxin Tai, Yueci Deng, Guiliang Liu, and Kui Jia. You only teach once: Learn one-shot bimanual robotic manipulation from video demonstrations.arXiv preprint arXiv:2501.14208, 2025

  215. [223]

    Gat-grasp: Gesture-driven affor- dance transfer for task-aware robotic grasping

    Ruixiang Wang, Huayi Zhou, Xinyue Yao, Guiliang Liu, and Kui Jia. Gat-grasp: Gesture-driven affor- dance transfer for task-aware robotic grasping. In Proc. IEEE/RSJ Int. Conf. Intell. Robots Syst. (IROS), 2025

  216. [224]

    Any- point trajectory modeling for policy learning

    Chuan Wen, Xingyu Lin, John Ian Reyes So, Kai Chen, Qi Dou, Yang Gao, and Pieter Abbeel. Any- point trajectory modeling for policy learning. InProc. Robot.: Sci. Syst. (RSS), 2024. doi: 10.15607/RSS. 2024.XX.092

  217. [225]

    Vidbot: Learning generalizable 3D actions from in-the-wild 2D hu- man videos for zero-shot robotic manipulation

    Hanzhi Chen, Boyang Sun, Anran Zhang, Marc Polle- feys, and Stefan Leutenegger. Vidbot: Learning generalizable 3D actions from in-the-wild 2D hu- man videos for zero-shot robotic manipulation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  218. [226]

    Flowhoi: Flow-based semantics-grounded generation of hand-object interactions for dexterous robot manip- ulation.arXiv preprint arXiv:2602.13444, 2026

    Huajian Zeng, Lingyun Chen, Jiaqi Yang, Yuantai Zhang, Fan Shi, Peidong Liu, and Xingxing Zuo. Flowhoi: Flow-based semantics-grounded generation of hand-object interactions for dexterous robot manip- ulation.arXiv preprint arXiv:2602.13444, 2026

  219. [227]

    Flow as the cross-domain manipulation interface

    Mengda Xu, Zhenjia Xu, Yinghao Xu, Cheng Chi, Gordon Wetzstein, Manuela Veloso, and Shuran Song. Flow as the cross-domain manipulation interface. arXiv preprint arXiv:2407.15208, 2024

  220. [228]

    3DFlowAction: Learning cross-embodiment manip- ulation from 3D flow world model.arXiv preprint arXiv:2506.06199, 2025

    Hongyan Zhi, Peihao Chen, Siyuan Zhou, Yubo Dong, Quanxi Wu, Lei Han, and Mingkui Tan. 3DFlowAction: Learning cross-embodiment manip- ulation from 3D flow world model.arXiv preprint arXiv:2506.06199, 2025

  221. [229]

    Novaflow: Zero-shot manipulation via actionable flow from gen- erated videos.arXiv preprint arXiv:2510.08568, 2025

    Hongyu Li, Lingfeng Sun, Yafei Hu, Duy Ta, Jennifer Barry, George Konidaris, and Jiahui Fu. Novaflow: Zero-shot manipulation via actionable flow from gen- erated videos.arXiv preprint arXiv:2510.08568, 2025

  222. [230]

    Dream2flow: Bridg- ing video generation and open-world manipulation with 3D object flow.arXiv preprint arXiv:2512.24766, 2025

    Karthik Dharmarajan, Wenlong Huang, Jiajun Wu, Li Fei-Fei, and Ruohan Zhang. Dream2flow: Bridg- ing video generation and open-world manipulation with 3D object flow.arXiv preprint arXiv:2512.24766, 2025

  223. [231]

    3PoinTr: 3D point tracks for learning ma- nipulation from unconstrained human videos.arXiv preprint arXiv:2603.08485, 2026

    Adam Hung, Bardienus Pieter Duisterhof, and Jeffrey Ichnowski. 3PoinTr: 3D point tracks for learning ma- nipulation from unconstrained human videos.arXiv preprint arXiv:2603.08485, 2026

  224. [232]

    Dex4D: Task-agnostic point track policy for sim-to-real dexterous manipulation

    Yuxuan Kuang, Sungjae Park, Katerina Fragkiadaki, and Shubham Tulsiani. Dex4D: Task-agnostic point track policy for sim-to-real dexterous manipulation. arXiv preprint arXiv:2602.15828, 2026

  225. [233]

    Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation

    Homanga Bharadhwaj, Debidatta Dwibedi, Abhinav Gupta, Shubham Tulsiani, Carl Doersch, Ted Xiao, Dhruv Shah, Fei Xia, Dorsa Sadigh, and Sean Kir- mani. Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation. arXiv preprint arXiv:2409.16283, 2024

  226. [234]

    Videopoet: A large language model for zero-shot video generation.arXiv preprint arXiv:2312.14125, 2023

    Dan Kondratyuk, Lijun Yu, Xiuye Gu, Jos´e Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, et al. Videopoet: A large language model for zero-shot video generation.arXiv preprint arXiv:2312.14125, 2023

  227. [235]

    Robowheel: A data engine from real-world human demonstra- tions for cross-embodiment robotic learning

    Yuhong Zhang, Zihan Gao, Shengpeng Li, Ling-Hao Chen, Kaisheng Liu, Runqing Cheng, Xiao Lin, Jun- jia Liu, Zhuoheng Li, Jingyi Feng, et al. Robowheel: A data engine from real-world human demonstra- tions for cross-embodiment robotic learning. In Proc. IEEE/CVF Conf. Comput. Vi...

  228. [236]

    Bhawna Paliwal, Haritheja Etukuru, William Liang, Pieter Abbeel, Nur Muhammad Mahi Shafiullah, and Jitendra Malik. Do as i do: Dexterous manipulation 35 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer data from e...

  229. [237]

    Egoinfinity: A web-scale 4D hand- object interaction data engine for any-view robot re- targeting and video-to-action robot learning.arXiv preprint arXiv:2606.17385, 2026

    Gaotian Wang, Kejia Ren, Andrew Morgan, Yiting Chen, Howard H Qian, Podshara Chanrungmaneekul, and Kaiyu Hang. Egoinfinity: A web-scale 4D hand- object interaction data engine for any-view robot re- targeting and video-to-action robot learning.arXiv preprint arXiv:2606.17385, 2026

  230. [238]

    Human2robot: Learning robot actions from paired human-robot videos

    Sicheng Xie, Haidong Cao, Zejia Weng, Zhen Xing, Haoran Chen, Shiwei Shen, Jiaqi Leng, Zuxuan Wu, and Yu-Gang Jiang. Human2robot: Learning robot actions from paired human-robot videos. InProc. AAAI Conf. Artif. Intell., 2026

  231. [239]

    Tracegen: World modeling in 3D trace-space enables learning from cross-embodiment videos.arXiv preprint arXiv:2511.21690, 2025

    Seungjae Lee, Yoonkyo Jung, Inkook Chun, Yao- Chih Lee, Zikui Cai, Hongjia Huang, Aayush Talreja, Tan Dat Dao, Yongyuan Liang, Jia-Bin Huang, and Furong Huang. Tracegen: World modeling in 3D trace-space enables learning from cross-embodiment videos.arXiv preprint arXiv:2511.21...

  232. [240]

    H2r-grounder: A paired-data-free paradigm for translating human interaction videos into physically grounded robot videos.arXiv preprint arXiv:2512.09406, 2025

    Hai Ci, Xiaokang Liu, Pei Yang, Yiren Song, and Mike Zheng Shou. H2r-grounder: A paired-data-free paradigm for translating human interaction videos into physically grounded robot videos.arXiv preprint arXiv:2512.09406, 2025

  233. [241]

    Qwen-robotmanip tech- nical report: Alignment unlocks scale for robotic manipulation foundation models.arXiv preprint arXiv:2606.17846, 2026

    Haoqi Yuan, Zhixuan Liang, Anzhe Chen, Ye Wang, Haoyang Li, Pei Lin, Yiyang Huang, Zixing Lei, Tong Zhang, Jiazhao Zhang, et al. Qwen-robotmanip tech- nical report: Alignment unlocks scale for robotic manipulation foundation models.arXiv preprint arXiv:2606.17846, 2026

  234. [242]

    Real-time joint tracking of a hand manip- ulating an object from RGB-D input

    Srinath Sridhar, Franziska Mueller, Michael Zoll- hoefer, Dan Casas, Antti Oulasvirta, and Christian Theobalt. Real-time joint tracking of a hand manip- ulating an object from RGB-D input. InProc. Eur. Conf. Comput. Vis. (ECCV), 2016

  235. [243]

    Real-time hand tracking under occlusion from an egocentric RGB-D sensor

    Franziska Mueller, Dushyant Mehta, Oleksandr Sot- nychenko, Srinath Sridhar, Dan Casas, and Christian Theobalt. Real-time hand tracking under occlusion from an egocentric RGB-D sensor. InProc. IEEE Int. Conf. Comput. Vis. (ICCV), 2017

  236. [244]

    First-person hand action benchmark with RGB-D videos and 3D hand pose annotations

    Guillermo Garcia-Hernando, Shanxin Yuan, Seun- gryul Baek, and Tae-Kyun Kim. First-person hand action benchmark with RGB-D videos and 3D hand pose annotations. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018

  237. [245]

    Frei- hand: A dataset for markerless capture of hand pose and shape from single RGB images

    Christian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan Russell, Max Argus, and Thomas Brox. Frei- hand: A dataset for markerless capture of hand pose and shape from single RGB images. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2019

  238. [246]

    Contactdb: Analyzing and pre- dicting grasp contact via thermal imaging

    Samarth Brahmbhatt, Cusuh Ham, Charles C Kemp, and James Hays. Contactdb: Analyzing and pre- dicting grasp contact via thermal imaging. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019

  239. [247]

    Ho-3D v3: Improving the accuracy of hand- object annotations of the HO-3D dataset.arXiv preprint arXiv:2107.00887, 2021

    Shreyas Hampali, Sayan Deb Sarkar, and Vincent Lepetit. Ho-3D v3: Improving the accuracy of hand- object annotations of the HO-3D dataset.arXiv preprint arXiv:2107.00887, 2021

  240. [248]

    Interhand2.6m: A dataset and baseline for 3D interacting hand pose es- timation from a single RGB image

    Gyeongsik Moon, Shoou-I Yu, He Wen, Takaaki Shiratori, and Kyoung Mu Lee. Interhand2.6m: A dataset and baseline for 3D interacting hand pose es- timation from a single RGB image. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020. doi: 10.1007/ 978-3-030-58565-5\ 33

  241. [249]

    Reconstructing hand-object interactions in the wild

    Zhe Cao, Ilija Radosavovic, Angjoo Kanazawa, and Ji- tendra Malik. Reconstructing hand-object interactions in the wild. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021

  242. [250]

    Dexycb: A benchmark for capturing hand grasping of objects

    Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, et al. Dexycb: A benchmark for capturing hand grasping of objects. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2021

  243. [251]

    H2o: Two hands manipu- lating objects for first person interaction recognition

    Taein Kwon, Bugra Tekin, Jan St ¨uhmer, Federica Bogo, and Marc Pollefeys. H2o: Two hands manipu- lating objects for first person interaction recognition. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021

  244. [252]

    Oakink: A large-scale knowledge repository for understanding hand-object interaction

    Lixin Yang, Kailin Li, Xinyu Zhan, Fei Wu, Anran Xu, Liu Liu, and Cewu Lu. Oakink: A large-scale knowledge repository for understanding hand-object interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  245. [253]

    Assembly- hands: Towards egocentric activity understanding via 3D hand pose estimation

    Takehiko Ohkawa, Kun He, Fadime Sener, Tomas Hodan, Luan Tran, and Cem Keskin. Assembly- hands: Towards egocentric activity understanding via 3D hand pose estimation. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023

  246. [254]

    Showme: Benchmarking object- agnostic hand-object 3D reconstruction

    Anilkumar Swamy, Vincent Leroy, Philippe Wein- zaepfel, Fabien Baradel, Salma Galaaoui, Romain Br´egier, Matthieu Armando, Jean-Sebastien Franco, and Gr´egory Rogez. Showme: Benchmarking object- agnostic hand-object 3D reconstruction. InProc. IEEE/CVF Int. Conf. Comput. Vis. (...

  247. [255]

    Dense hand- object (ho) graspnet with full grasping taxonomy and dynamics

    Woojin Cho, Jihyun Lee, Minjae Yi, Minje Kim, Taeyun Woo, Donghwan Kim, Taewook Ha, Hyokeun Lee, Je-Hwan Ryu, Woontack Woo, et al. Dense hand- object (ho) graspnet with full grasping taxonomy and dynamics. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024

  248. [256]

    Oakink2: A dataset of bimanual hands-object ma- nipulation in complex task completion

    Xinyu Zhan, Lixin Yang, Yifei Zhao, Kangrui Mao, Hanlin Xu, Zenan Lin, Kailin Li, and Cewu Lu. Oakink2: A dataset of bimanual hands-object ma- nipulation in complex task completion. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  249. [257]

    Taco: Bench- marking generalizable bimanual tool-action-object understanding

    Yun Liu, Haolin Yang, Xu Si, Ling Liu, Zipeng Li, Yuxiang Zhang, Yebin Liu, and Li Yi. Taco: Bench- marking generalizable bimanual tool-action-object understanding. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024

  250. [258]

    Gigahands: A massive annotated dataset of bimanual hand activi- ties

    Rao Fu, Dingxi Zhang, Alex Jiang, Wanjia Fu, Austin Funk, Daniel Ritchie, and Srinath Sridhar. Gigahands: A massive annotated dataset of bimanual hand activi- ties. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025

  251. [259]

    Ho-cap: A capture sys- tem and dataset for 3D reconstruction and pose track- ing of hand-object interaction

    Jikai Wang, Qifan Zhang, Yu-Wei Chao, Bowen Wen, Xiaohu Guo, and Yu Xiang. Ho-cap: A capture sys- tem and dataset for 3D reconstruction and pose track- ing of hand-object interaction. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 38, 2025

  252. [260]

    Anthony, Zhuorui Zhang, and Cewu Lu

    Zhenjun Yu, Wenqiang Xu, Pengfei Xie, Yutong Li, Brian W. Anthony, Zhuorui Zhang, and Cewu Lu. Dynamic reconstruction of hand-object interaction with distributed force-aware contact representation. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025

  253. [261]

    Hot3d: Hand and object track- ing in 3D from egocentric multi-view videos

    Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Shangchen Han, Fan Zhang, Lin- guang Zhang, Jade Fountain, Edward Miller, Se- len Basol, et al. Hot3d: Hand and object track- ing in 3D from egocentric multi-view videos. In Proc. IEEE/CVF Conf. Comput. Vis....

  254. [262]

    A VI-HT: Adaptive vision-IMU fusion for 3D hand tracking

    Ziyi Kou, Ankit Kumar, Mia Huang, Taylor Niehues, Vatsal Mehta, Ergys Ristani, and Li Guan. A VI-HT: Adaptive vision-IMU fusion for 3D hand tracking. arXiv preprint arXiv:2605.21714, 2026

  255. [263]

    Show3d: Capturing scenes of 3D hands and objects in the wild

    Patrick Rim, Kevin Harris, Braden Copple, Shangchen Han, Xu Xie, Ivan Shugurov, Sizhe An, He Wen, Alex Wong, Tomas Hodan, et al. Show3d: Capturing scenes of 3D hands and objects in the wild. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026

  256. [264]

    Handx: Scal- ing bimanual motion and interaction generation

    Zimu Zhang, Yucheng Zhang, Xiyan Xu, Ziyin Wang, Sirui Xu, Kai Zhou, Bing Zhou, Chuan Guo, Jian Wang, Yu-Xiong Wang, et al. Handx: Scal- ing bimanual motion and interaction generation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026

  257. [265]

    Activitynet: A large-scale video benchmark for human activity un- derstanding

    Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity un- derstanding. InProc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2015. doi: 10.1109/CVPR.2015. 7298698

  258. [266]

    Hollywood in homes: Crowdsourcing data collection for activ- ity understanding

    Gunnar A Sigurdsson, G¨ul Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. Hollywood in homes: Crowdsourcing data collection for activ- ity understanding. InProc. Eur. Conf. Comput. Vis. (ECCV), 2016

  259. [267]

    The kinetics human action video dataset.arXiv preprint arXiv:1705.06950, 2017

    Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset.arXiv preprint arXiv:1705.06950, 2017

  260. [268]

    The ”something something” video database for learn- ing and evaluating visual common sense

    Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fr¨und, Peter Yian- ilos, Moritz Mueller-Freitag, Florian Hoppe, Chris- tian Thurau, Ingo Bax, and Roland Memisevic. The ”something something” video d...

  261. [269]

    Ava: A video dataset of spatio-temporally localized atomic visual actions

    Chunhui Gu, Chen Sun, David A Ross, Carl V on- drick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al. Ava: A video dataset of spatio-temporally localized atomic visual actions. In Proc. IEEE/CVF Conf. Comput....

  262. [270]

    Hacs: Human action clips and segments dataset for recognition and temporal local- ization

    Hang Zhao, Antonio Torralba, Lorenzo Torresani, and Zhicheng Yan. Hacs: Human action clips and segments dataset for recognition and temporal local- ization. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2019

  263. [271]

    Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding 37 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer by watching hund...

  264. [272]

    Fin- egym: A hierarchical video dataset for fine-grained action understanding

    Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin. Fin- egym: A hierarchical video dataset for fine-grained action understanding. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2020

  265. [273]

    The epic-kitchens dataset: Collection, challenges and baselines.IEEE Trans

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. The epic-kitchens dataset: Collection, challenges and baselines.IEEE Trans. Pattern Anal. Mach. Intell., ...

  266. [274]

    Assembly101: A large-scale multi-view video dataset for understanding procedural activities

    Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao. Assembly101: A large-scale multi-view video dataset for understanding procedural activities. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  267. [275]

    Ego4d: Around the world in 3,000 hours of egocentric video

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022

  268. [276]

    Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world

    Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bugra Tekin, Felipe Vieira Frujeri, et al. Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world. InProc. IEEE/CVF Int. Conf....

  269. [277]

    Ego-exo4d: Understand- ing skilled human activity from first-and third-person perspectives

    Kristen Grauman, Andrew Westbury, Lorenzo Tor- resani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al. Ego-exo4d: Understand- ing skilled human activity from first-and third-person perspectives. InProc. IE...

  270. [278]

    Egodex: Learning dexterous manipulation from large-scale egocentric video.arXiv preprint arXiv:2505.11709, 2025

    Ryan Hoque, Peide Huang, David J Yoon, Mouli Siva- purapu, and Jian Zhang. Egodex: Learning dexterous manipulation from large-scale egocentric video.arXiv preprint arXiv:2505.11709, 2025

  271. [279]

    Openego: A large-scale multimodal egocentric dataset for dexterous manipu- lation.arXiv preprint arXiv:2509.05513, 2025

    Ahad Jawaid and Yu Xiang. Openego: A large-scale multimodal egocentric dataset for dexterous manipu- lation.arXiv preprint arXiv:2509.05513, 2025

  272. [280]

    Egolive: A large-scale egocentric dataset from real-world human tasks.arXiv preprint arXiv:2604.23570, 2026

    Yihang Li, Xuelong Wei, Jingzhou Luo, Yingjing Xiao, Yibo Bai, Guangyuan Zhou, Teng Zou, Chen- guang Gui, Jiajun Wen, He Zhang, et al. Egolive: A large-scale egocentric dataset from real-world human tasks.arXiv preprint arXiv:2604.23570, 2026

  273. [281]

    Egoverse: An egocentric human dataset for robot learning from around the world.arXiv preprint arXiv:2604.07607, 2026

    Ryan Punamiya, Simar Kareer, Zeyi Liu, Josh Cit- ron, Ri-Zhao Qiu, Xiongyi Cai, Alexey Gavryushin, Jiaqi Chen, Davide Liconti, Lawrence Y Zhu, et al. Egoverse: An egocentric human dataset for robot learning from around the world.arXiv preprint arXiv:2604.07607, 2026

  274. [282]

    Humannet: Scaling human-centric video learning to one million hours

    Yufan Deng and Daquan Zhou. Humannet: Scaling human-centric video learning to one million hours. arXiv preprint arXiv:2605.06747, 2026

  275. [283]

    FEEL (force-enhanced egocentric learning): A dataset for physical action understanding.arXiv preprint arXiv:2603.15847, 2026

    Eadom Dessalene, Botao He, Michael Maynord, Yonatan Tussa, Pavan Mantripragada, Yianni Kara- bati, Nirupam Roy, and Yiannis Aloimonos. FEEL (force-enhanced egocentric learning): A dataset for physical action understanding.arXiv preprint arXiv:2603.15847, 2026

  276. [284]

    Open- AoE: An open egocentric manipulation dataset and toolchain for embodied learning.arXiv preprint arXiv:2607.14183, 2026

    Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu, Hao Chen, Qingze Guan, Zhengxing Wu, Wanke Zhan, Yang Sun, Zhiyi Huang, et al. Open- AoE: An open egocentric manipulation dataset and toolchain for embodied learning.arXiv preprint arXiv:2607.14183, 2026

  277. [285]

    Human3.6m: Large scale datasets and predictive methods for 3D human sensing in nat- ural environments.IEEE Trans

    Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cris- tian Sminchisescu. Human3.6m: Large scale datasets and predictive methods for 3D human sensing in nat- ural environments.IEEE Trans. Pattern Anal. Mach. Intell., 36(7), 2014. doi: 10.1109/TPAMI.2013.248

  278. [286]

    Model based training, detection and pose estimation of texture-less 3D objects in heavily cluttered scenes

    Stefan Hinterstoisser, Vincent Lepetit, Slobodan Ilic, Stefan Holzer, Gary Bradski, Kurt Konolige, and Nassir Navab. Model based training, detection and pose estimation of texture-less 3D objects in heavily cluttered scenes. InProc. Asian Conf. Comput. Vis. (ACCV), 2012

  279. [287]

    PoseCNN: A convolutional neural network for 6D object pose estimation in cluttered scenes.arXiv preprint arXiv:1711.00199, 2017

    Yu Xiang, Tanner Schmidt, Venkatraman Narayanan, and Dieter Fox. PoseCNN: A convolutional neural network for 6D object pose estimation in cluttered scenes.arXiv preprint arXiv:1711.00199, 2017

  280. [288]

    A point set generation network for 3D object reconstruc- tion from a single image

    Haoqiang Fan, Hao Su, and Leonidas J Guibas. A point set generation network for 3D object reconstruc- tion from a single image. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2017

  281. [289]

    GANs trained by a two time-scale update rule con- verge to a local nash equilibrium

    Martin Heusel, Hubert Ramsauer, Thomas Un- terthiner, Bernhard Nessler, and Sepp Hochreiter. GANs trained by a two time-scale update rule con- verge to a local nash equilibrium. InAdv. Neural Inf. Process. Syst. (NeurIPS), 2017. 38 Hand-Object Interaction in the Age of Large F...

  282. [290]

    Towards accurate generative models of video: A new metric & challenges.arXiv preprint arXiv:1812.01717, 2018

    Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. Towards accurate generative models of video: A new metric & challenges.arXiv preprint arXiv:1812.01717, 2018

  283. [291]

    Image quality assessment: from error visibility to structural similarity.IEEE Trans

    Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity.IEEE Trans. Image Process., 13(4), 2004

  284. [292]

    The unreasonable ef- fectiveness of deep features as a perceptual metric

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable ef- fectiveness of deep features as a perceptual metric. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018. 39

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.