REVIEW 3 major objections 3 minor 292 references
Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
T0 review · 3 major / 3 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read This survey argues that the scattered field of foundation-model-assisted hand-object interaction is unified by a taxonomy of eight types of prior knowledge in three families, and that every method can be understood by which prior it injects
desk verdict A useful organizing taxonomy for HOI+foundation models, but the uncertainty-mitigation claims are inferred rather than evidenced. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central organizing device is the eight-sub-prior taxonomy combined with a pipeline abstraction: each method is characterized by its prior source, the representation it injects, and the injection operator it uses. This vocabulary—initialization, regularization, conditioning, token fusion, score-guided regularization, retargeting—describes how foundation-model knowledge enters the HOI backbone and task head. The taxonomy is linked to a five-uncertainty model (shape, spatial, physical, semantic, dynamic), so each prior family is mapped to the specific ambiguities it mitigates. For embodied transfer, the machinery is a five-route diagram tracing HOI evidence through a transferred signal and
What would settle it
A concrete test would be to re-run the survey's taxonomic tables under a broader inclusion rule that also counts task-specialized pretrained initializations as foundation priors; if the qualitative mapping between prior families and the five uncertainties still holds, the boundary is not load-bearing, but if the taxonomy's cells become crowded and the uncertainty mapping blurs, the paper's organizational claim is weakened. Alternatively, a single benchmark that jointly reports geometry, contact agreement, physical plausibility, and task success across the six tasks could settle whether current
Extended reading notes
Core claim
The paper's central claim is that the apparent fragmentation of HOI-plus-foundation-model work reflects a missing organizing dimension: what knowledge is introduced, where it enters, and which uncertainty it reduces. The authors define a deliberately narrow boundary—a method counts as foundation-model-prior only when an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge through predictions, representations, transferred parameters, adaptation, or distillation—and on that basis catalog eight sub-priors in three families. They argue this taxonomy covers six HOI tasks (pose estimation, hand-held object reconstruction, dynamic reconstruction, gr
Load-bearing premise
The load-bearing premise is the paper's stipulative boundary for what counts as a foundation-model prior: a method is included only when an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge through predictions, representations, transferred parameters, adaptation, or distillation, and task-specialized pretraining does not count.
Editorial extensions
If this is right
- HOI methods should be characterized and compared by the foundation-model priors they exploit, not only by architecture or training data.
- Geometric, semantic, and visual priors are complementary, so multi-prior systems are the natural next step for reducing shape, spatial, physical, semantic, and dynamic uncertainty together.
- Evaluation that reports only geometry or image fidelity is insufficient; contact agreement, physical plausibility, and functional task success must be reported alongside.
- Embodied transfer is best viewed as a downstream consumer of HOI reconstruction and generation outputs, with concrete transfer routes from human evidence to robot policy.
- The taxonomy provides a shared vocabulary that can make future HOI papers comparable and can guide the design of integrated, verifiable HOI systems.
Reading between the lines
- The taxonomy's gatekeeping boundary is an author choice rather than an empirical result; re-running the survey's tables with a broader definition of foundation priors (for example, including task-specialized pretrained initializations) would shift coverage and primary/auxiliary assignments.
- The eight sub-priors are not cleanly orthogonal in practice: the same vision-language model counts as a semantic grounding prior when used for localization and as a language reasoning prior when used for intent inference, suggesting the taxonomy is a lens for reading methods rather than a unique partition.
- A concrete extension the paper leaves implicit is a joint benchmark that measures geometry, contact, physical plausibility, and task success together; such a benchmark would directly test the survey's claim that current metrics have blind spots.
- The emerging line of action-conditioned video generation (world models) is identified as an inference pattern rather than an injection operator; editorially, this could mature into a fourth visual sub-prior or a new task category as interactive rollouts grow.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey organizes the hand-object interaction (HOI) literature around eight foundation-model sub-priors grouped into geometric, semantic, and visual families, and maps how these priors are represented, injected, and adapted across six HOI reconstruction/generation tasks and embodied transfer to robot learning. The paper claims to be the first systematic review of foundation-model priors for HOI, and supports the taxonomy with representative methods in Tables 3–5, a dataset/evaluation summary, and a live repository.
Significance. If the taxonomy and the uncertainty-mitigation mapping are accepted, the survey would provide a useful organizing principle for a rapidly growing but fragmented field. The manuscript is internally consistent, the gatekeeping definition of 'foundation-model prior' is explicit and applied transparently (e.g., excluding ViTPose initialization in HaMeR), and the paper repeatedly identifies evaluation blind spots. These are strengths. However, the central analytical output — the mapping from each prior to the HOI uncertainties it 'helps reduce' — is largely inferential and not backed by the surveyed papers' evidence, which weakens the strongest claim.
major comments (3)
- [Abstract, Table 3, Secs. 3.3/5.3/7.2.3] The abstract claims the survey reveals 'which HOI uncertainty it helps reduce,' and Table 3's Unc.↓ column operationalizes this. Yet the cited papers rarely ablate the foundation-model component against the listed uncertainty. The paper itself concedes: Sec. 3.3 states a reconstructed shape 'should not be interpreted as interaction evidence by itself'; Sec. 5.3 says 'image realism is not interaction correctness'; Sec. 7.2.3 notes physical metrics 'depend strongly on mesh quality, friction, contact modeling, and simulator settings.' The Unc.↓ assignments thus appear to be authors' inference, not literature evidence. Please either (a) reclassify the Unc.↓ column as a hypothesized mechanism with a clear caveat, or (b) for each Table 3 row, cite an ablation/experiment from the original paper that supports the assignment. This is load-bearing because the abstract frames the entire survey arou
- [Sec. 1] The manuscript claims 'the first systematic review' of foundation-model priors for HOI, but no search/selection methodology is provided: no databases, query terms, inclusion/exclusion criteria, screening process, or date cutoff are documented. Without this, the 'systematic' claim cannot be audited and the survey cannot be distinguished from an author-selected narrative review. Please add a methodology paragraph (or appendix) documenting the protocol, or soften the claim to 'first literature survey' / 'comprehensive review'.
- [Table 3 and Table 4] The representative method list and dataset tables rely on a large number of arXiv preprints and very recent 2026 venue entries (e.g., GeoHand arXiv:2605.17354, ScaleHP arXiv:2606.25619, several CVPR 2026 entries). Given the survey's cutoff is not stated, the 'first systematic' claim and the balanced coverage of the field are difficult to assess. State the literature cutoff date, and mark entries that are preprints or not yet peer-reviewed at that date. This is especially relevant because several Table 3 exemplars are from the authors' own group, and the selection criteria for 'representative' methods are not specified.
minor comments (3)
- [Fig. 7] The figure caption and labels refer to 'Sec. 4.2: Human-Data Pretraining', 'Sec. 4.3: Human-to-Robot Skill Transfer', and 'Sec. 4.4: HOI-to-Robot Data Engines', but the corresponding sections in the text are 6.2, 6.3, and 6.4. Update the figure numbering.
- [Sec. 6.2.1] The phrase 'frame-aligned action chunks' and '6DoF object trajectories' appear without prior definition in the HOI taxonomy; consider adding these to the interaction-representation list in Sec. 2.2.2 for terminological consistency.
- [Throughout] The text frequently uses 'systematically analyze' and 'systematic coverage' without a clear definition of systematicity. Align these terms with the (proposed) methodology, or use more neutral phrasing such as 'structured analysis.'
Circularity Check
No significant circularity: the taxonomy is stipulative and the uncertainty mapping is interpretive, but no result reduces to its inputs by construction.
full rationale
This is a survey, not a derivation. Its central deliverable is a taxonomy of foundation-model priors and a qualitative mapping from priors to five uncertainties. The boundary in Sec. 1 is explicitly stipulative: "we use a deliberately narrow boundary: a method is considered a foundation-model-prior method only when an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge to the HOI pipeline." Classification decisions are therefore authorial choices, not circular deductions. The Unc.↓ column in Table 3 is interpretive attribution rather than a fitted/predicted quantity; the paper itself qualifies the evidence base in several places: Sec. 3.3 says "the reconstructed shape should not be interpreted as interaction evidence by itself," Sec. 5.3 says "image realism is not interaction correctness," Sec. 7.2.3 says penetration and force-closure measures "depend strongly on mesh quality, friction, contact modeling, and simulator settings," and Sec. 8.4 lists prior reliability as an open problem. These caveats weaken the empirical support for the mapping, but they do not make it circular: no quantity is fitted and then presented as a prediction, and no equation or definition forces the survey's conclusions. The self-citations (GeoHand, HandOS, and ScaleHP in Table 3, with author overlap; MoGe-2 also has overlapping authors) are visible, but these are externally falsifiable method papers used as exemplars, the taxonomy does not depend on them, and per rule 4 self-citation alone is not circularity. The central claim—that the fragmented HOI literature can be organized by eight foundation-model sub-priors in three families—is a transparent analytic frame with independent content.
Assumptions & free parameters
assumptions (3)
- ad hoc to paper A method counts as foundation-model-prior only if an explicitly identified, large-scale general-purpose pretrained model contributes cross-domain knowledge through predictions, representations, transferred parameters, adaptation, or distillation.
- domain assumption HOI failure modes decompose into five uncertainties: shape, spatial, physical, semantic, dynamic.
- domain assumption Each method can be assigned a unique primary sub-prior and optional auxiliary sub-priors with a single injection operator.
Cite this review
Pith. "Pith review of Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer." pith.science (2026). https://pith.science/paper/FNALSTCS
@misc{pith2026260728394,
author = {Pith},
title = {Pith review of: Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer},
year = {2026},
howpublished = {\url{https://pith.science/paper/FNALSTCS}},
note = {Machine review of arXiv:2607.28394}
}
read the original abstract
Hand-object interaction (HOI) modeling remains challenging because it requires joint reasoning about hand articulation, object geometry, contact, semantics, and dynamics under severe visual uncertainty. Foundation models introduce transferable prior knowledge learned from large-scale cross-domain data, offering new ways to address these challenges beyond task-specific data and models. However, the rapidly growing literature remains fragmented, and existing studies typically describe these methods simply as ``using large models'' without systematically characterizing what knowledge is introduced, where it enters the HOI pipeline, or which HOI uncertainty it helps reduce. This survey presents the first systematic review of foundation-model priors for HOI. We organize the literature into six HOI tasks spanning reconstruction and generation. More importantly, we establish a taxonomy of eight foundation-model sub-priors grouped into geometric, semantic, and visual families. Geometric priors encompass shape retrieval, shape reconstruction, and spatial reconstruction; semantic priors include semantic grounding and language reasoning; and visual priors cover visual representation, image generation, and video generation. Based on this taxonomy, we systematically analyze how different priors are represented, injected, and adapted across HOI pipelines and tasks. Beyond how foundation models empower HOI, we further examine how HOI-derived knowledge is used in robot learning, including human-data pretraining, human-to-robot skill transfer, and HOI-to-robot data generation. Finally, we summarize datasets and evaluation protocols, and discuss limitations and future directions toward more generalizable HOI systems. To support long-term progress, we curate a live repository that continuously aggregates emerging methods and benchmarks.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
H+O: unified egocentric recognition of 3D hand- object poses and interactions
Bugra Tekin, Federica Bogo, and Marc Pollefeys. H+O: unified egocentric recognition of 3D hand- object poses and interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019. doi: 10.1109/CVPR.2019.00464
arXiv 2019
-
[2]
Learning joint reconstruction of hands and manipulated objects
Yana Hasson, Gul Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J Black, Ivan Laptev, and Cordelia Schmid. Learning joint reconstruction of hands and manipulated objects. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019
2019
-
[3]
CPF: Learning a contact potential field to model the hand-object interaction
Lixin Yang, Xinyu Zhan, Kailin Li, Wenqiang Xu, Jiefeng Li, and Cewu Lu. CPF: Learning a contact potential field to model the hand-object interaction. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021
2021
-
[4]
HOLD: Category-agnostic 3D re- construction of interacting hands and objects from video
Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen, Muhammed Kocabas, Michael J Black, and Otmar Hilliges. HOLD: Category-agnostic 3D re- construction of interacting hands and objects from video. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[5]
What’s in your hands? 3D reconstruction of generic objects in hands
Yufei Ye, Abhinav Gupta, and Shubham Tulsiani. What’s in your hands? 3D reconstruction of generic objects in hands. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[6]
Grasping field: Learning implicit representations for human grasps
Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael J Black, Krikamol Muandet, and Siyu Tang. Grasping field: Learning implicit representations for human grasps. InProc. Int. Conf. 3D Vis. (3DV), 2020
2020
-
[7]
DUSt3R: Geomet- ric 3D vision made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. DUSt3R: Geomet- ric 3D vision made easy. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[8]
Samarth Brahmbhatt, Chengcheng Tang, Christo- pher D Twigg, Charles C Kemp, and James Hays. 23 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer Contactpose: A dataset of grasps with object con- tact and hand pose. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020
2020
Show all 292 references
-
[9]
S2contact: Graph-based network for 3D hand-object contact estimation with semi-supervised learning
Tze Ho Elden Tse, Zhongqun Zhang, Kwang In Kim, Ales Leonardis, Feng Zheng, and Hyung Jin Chang. S2contact: Graph-based network for 3D hand-object contact estimation with semi-supervised learning. In Proc. Eur. Conf. Comput. Vis. (ECCV), 2022. doi: 10.1007/978-3-031-19769-7\ 33
2022 doi
-
[10]
Contactopt: Optimizing contact to improve grasps
Patrick Grady, Chengcheng Tang, Christopher D Twigg, Minh V o, Samarth Brahmbhatt, and Charles C Kemp. Contactopt: Optimizing contact to improve grasps. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2021
2021
-
[11]
Deep- SimHO: Stable pose estimation for hand-object in- teraction via physics simulation
Rong Wang, Wei Mao, and Hongdong Li. Deep- SimHO: Stable pose estimation for hand-object in- teraction via physics simulation. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023
2023
-
[12]
D- grasp: Physically plausible dynamic grasp synthesis for hand-object interactions
Sammy Christen, Muhammed Kocabas, Emre Aksan, Jemin Hwangbo, Jie Song, and Otmar Hilliges. D- grasp: Physically plausible dynamic grasp synthesis for hand-object interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[13]
Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jian- wei Yang, Hang Su, et al. Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024
2024
-
[14]
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023
2023
-
[15]
SemGrasp: Semantic grasp generation via language aligned discretization
Kailin Li, Jingbo Wang, Lixin Yang, Cewu Lu, and Bo Dai. SemGrasp: Semantic grasp generation via language aligned discretization. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024
2024
-
[16]
Text2grasp: Synthesis of grasps by text prompts for object grasping parts
Xiaoyun Chang and Yi Sun. Text2grasp: Synthesis of grasps by text prompts for object grasping parts. In Proc. Int. Symp. Neural Netw. (ISNN), 2025
2025
-
[17]
Towards unconstrained joint hand-object re- construction from RGB videos
Yana Hasson, G¨ul Varol, Cordelia Schmid, and Ivan Laptev. Towards unconstrained joint hand-object re- construction from RGB videos. InProc. Int. Conf. 3D Vis. (3DV), 2021
2021
-
[18]
Grasp in gaussians: Fast monocular recon- struction of dynamic hand-object interactions.arXiv preprint arXiv:2604.12929, 2026
Ayce Idil Aytekin, Xu Chen, Zhengyang Shen, Thabo Beeler, Helge Rhodin, Rishabh Dabral, and Christian Theobalt. Grasp in gaussians: Fast monocular recon- struction of dynamic hand-object interactions.arXiv preprint arXiv:2604.12929, 2026
2026 arXiv
-
[19]
EasyHOI: Unleashing the power of large models for reconstructing hand-object inter- actions in the wild
Yumeng Liu, Xiaoxiao Long, Zemin Yang, Yuan Liu, Marc Habermann, Christian Theobalt, Yuexin Ma, and Wenping Wang. EasyHOI: Unleashing the power of large models for reconstructing hand-object inter- actions in the wild. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[20]
MagicHOI: Leveraging 3D priors for accurate hand-object recon- struction from short monocular video clips
Shibo Wang, Haonan He, Maria Parelli, Christoph Gebhardt, Zicong Fan, and Jie Song. MagicHOI: Leveraging 3D priors for accurate hand-object recon- struction from short monocular video clips. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025
2025
-
[21]
Diffusion-guided reconstruction of everyday hand-object interaction clips
Yufei Ye, Poorvi Hebbar, Abhinav Gupta, and Shub- ham Tulsiani. Diffusion-guided reconstruction of everyday hand-object interaction clips. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023
2023
-
[22]
Hand-object interaction image gen- eration
Hezhen Hu, Weilun Wang, Wengang Zhou, and Houqiang Li. Hand-object interaction image gen- eration. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 35, 2022
2022
-
[23]
HOIDiffusion: Generating realistic 3D hand-object interaction data
Mengqi Zhang, Yang Fu, Zheng Ding, Sifei Liu, Zhuowen Tu, and Xiaolong Wang. HOIDiffusion: Generating realistic 3D hand-object interaction data. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2024
2024
-
[24]
Svimo: Synchro- nized diffusion for video and motion generation in hand-object interaction scenarios
Lingwei Dang, Ruizhi Shao, Hongwen Zhang, Wei Min, Yebin Liu, and Qingyao Wu. Svimo: Synchro- nized diffusion for video and motion generation in hand-object interaction scenarios. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 38, 2025
2025
-
[25]
Instantmesh: Efficient 3D mesh generation from a single image with sparse- view large reconstruction models.arXiv preprint arXiv:2404.07191, 2024
Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Efficient 3D mesh generation from a single image with sparse- view large reconstruction models.arXiv preprint arXiv:2404.07191, 2024
2024 arXiv
-
[26]
OpenShape: Scaling up 3D shape representation towards open-world understanding
Minghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu, Xuanlin Li, Shizhong Han, Hong Cai, Fatih Porikli, and Hao Su. OpenShape: Scaling up 3D shape representation towards open-world understanding. In Adv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023
2023
-
[27]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. InAdv. Neu- ral Inf. Process. Syst. (NeurIPS), volume 36, 2023. 24 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer
2023
-
[28]
Qwen-VL: A versatile vision-language model for understanding, localiza- tion, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A versatile vision-language model for understanding, localiza- tion, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023
2023 arXiv
-
[29]
DINOv2: Learning robust vi- sual features without supervision.arXiv preprint arXiv:2304.07193, 2023
Maxime Oquab, Timoth´ee Darcet, Th´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fer- nandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. DINOv2: Learning robust vi- sual features without supervision.arXiv preprint arXiv:2304.07193, 2023
2023 arXiv
-
[30]
High- resolution image synthesis with latent diffusion mod- els
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High- resolution image synthesis with latent diffusion mod- els. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[31]
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. InProc. Int. Conf. Learn. Repre- sent. (ICLR), 2025
2025
-
[32]
Objaverse: A universe of annotated 3D objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3D objects. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2023
2023
-
[33]
Learning trans- ferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sas- try, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning trans- ferable visual models from natural language supervi- sion. InProc. Int. C...
2021
-
[34]
Reconstructing hand-held objects in 3D from images and videos
Jane Wu, Georgios Pavlakos, Georgia Gkioxari, and Jitendra Malik. Reconstructing hand-held objects in 3D from images and videos. InProc. Int. Conf. 3D Vis. (3DV), 2026
2026
-
[35]
Ghost: Fast category- agnostic hand-object interaction reconstruction from RGB videos using gaussian splatting
Ahmed Tawfik Aboukhadra, Marcel Rogge, Nadia Robertini, Abdalla Arafa, Jameel Malik, Ahmed El- hayek, and Didier Stricker. Ghost: Fast category- agnostic hand-object interaction reconstruction from RGB videos using gaussian splatting. InProc. IEEE/CVF Conf. Comput. Vis. Patter...
2026
-
[36]
Hand-held object reconstruction from RGB video with dynamic interaction
Shijian Jiang, Qi Ye, Rengan Xie, Yuchi Huo, and Jim- ing Chen. Hand-held object reconstruction from RGB video with dynamic interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[37]
Reconstructing hands in 3D with transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstructing hands in 3D with transformers. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[38]
Wilor: End- to-end 3D hand localization and reconstruction in-the- wild
Rolandos Alexandros Potamias, Jinglei Zhang, Jiankang Deng, and Stefanos Zafeiriou. Wilor: End- to-end 3D hand localization and reconstruction in-the- wild. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[39]
ViTPose: Simple vision transformer baselines for human pose estimation
Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. ViTPose: Simple vision transformer baselines for human pose estimation. InAdv. Neural Inf. Process. Syst. (NeurIPS), 2022
2022
-
[40]
Latent action pretraining from videos
Seonghyeon Ye, Joel Jang, Byeongguk Jeon, Se June Joo, Jianwei Yang, Baolin Peng, Ajay Mandlekar, Reuben Tan, Yu-Wei Chao, Bill Yuchen Lin, et al. Latent action pretraining from videos. InProc. Int. Conf. Learn. Represent. (ICLR), 2025
2025
-
[41]
Egovla: Learning vision-language-action models from egocentric hu- man videos.arXiv preprint arXiv:2507.12440, 2025
Ruihan Yang, Qinxi Yu, Yecheng Wu, Rui Yan, Borui Li, An-Chieh Cheng, Xueyan Zou, Yunhao Fang, Xuxin Cheng, Ri-Zhao Qiu, et al. Egovla: Learning vision-language-action models from egocentric hu- man videos.arXiv preprint arXiv:2507.12440, 2025
2025 arXiv
-
[42]
Dexmv: Imitation learning for dexterous manipula- tion from human videos
Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang, Ruihan Yang, Yang Fu, and Xiaolong Wang. Dexmv: Imitation learning for dexterous manipula- tion from human videos. InProc. Eur. Conf. Comput. Vis. (ECCV), 2022
2022
-
[43]
Egoengine: From ego- centric human videos to high-fidelity dexterous robot demonstrations.arXiv preprint arXiv:2606.12604, 2026
Yangcen Liu, Shuo Cheng, Xinchen Yin, Woo Chul Shin, Alfred Cueva, Yiran Yang, Zhenyang Chen, Chuye Zhang, and Danfei Xu. Egoengine: From ego- centric human videos to high-fidelity dexterous robot demonstrations.arXiv preprint arXiv:2606.12604, 2026
2026 arXiv
-
[44]
Efficient annotation and learning for 3D hand pose estimation: A survey.Int
Takehiko Ohkawa, Ryosuke Furuta, and Yoichi Sato. Efficient annotation and learning for 3D hand pose estimation: A survey.Int. J. Comput. Vis., 131(12), 2023
2023
-
[45]
A survey of deep learning methods and datasets for hand pose estimation from hand-object interaction images.Comput
Taeyun Woo, Wonjung Park, Woohyun Jeong, and Jinah Park. A survey of deep learning methods and datasets for hand pose estimation from hand-object interaction images.Comput. Graph., 116, 2023
2023
-
[46]
Advances in vision-based deep learning methods for interacting hands reconstruction: A survey.Comput
Yu Miao and Yue Liu. Advances in vision-based deep learning methods for interacting hands reconstruction: A survey.Comput. Graph., 124, 2024. 25 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer
2024
-
[47]
An overview of learning-based dexterous grasping: recent advances and future directions.Artif
Xu Song, Yongyao Li, Yunfan Zhang, Yufei Liu, and Lei Jiang. An overview of learning-based dexterous grasping: recent advances and future directions.Artif. Intell. Rev., 58(10), 2025
2025
-
[48]
Arthoi: Taming foundation models for monocular 4D reconstruction of hand-articulated- object interactions
Zikai Wang, Zhilu Zhang, Yiqing Wang, Hui Li, and Wangmeng Zuo. Arthoi: Taming foundation models for monocular 4D reconstruction of hand-articulated- object interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026
2026
-
[49]
Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: modeling and capturing hands and bodies together.ACM Trans. Graph., 36 (6), 2017. doi: 10.1145/3130800.3130883
2017
-
[50]
Nimble: a non-rigid hand model with bones and muscles.ACM Trans
Yuwei Li, Longwen Zhang, Zesong Qiu, Yingwenqi Jiang, Nianyi Li, Yuexin Ma, Yuyao Zhang, Lan Xu, and Jingyi Yu. Nimble: a non-rigid hand model with bones and muscles.ACM Trans. Graph., 41(4), 2022
2022
-
[51]
HOnnotate: A method for 3D an- notation of hand and object poses
Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. HOnnotate: A method for 3D an- notation of hand and object poses. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020
2020
-
[52]
AlignSDF: Pose-aligned signed distance fields for hand-object reconstruction
Zerui Chen, Yana Hasson, Cordelia Schmid, and Ivan Laptev. AlignSDF: Pose-aligned signed distance fields for hand-object reconstruction. InProc. Eur. Conf. Comput. Vis. (ECCV), 2022
2022
-
[53]
HandNeRF: Learn- ing to reconstruct hand-object interaction scene from a single RGB image
Hongsuk Choi, Nikhil Chavan-Dafle, Jiacheng Yuan, V olkan Isler, and Hyunsoo Park. HandNeRF: Learn- ing to reconstruct hand-object interaction scene from a single RGB image. InProc. IEEE Int. Conf. Robot. Autom. (ICRA), 2024
2024
-
[54]
Simultaneous localization and mapping: part i.IEEE Robot
Hugh Durrant-Whyte and Tim Bailey. Simultaneous localization and mapping: part i.IEEE Robot. Autom. Mag., 13(2), 2006
2006
-
[55]
Seitz, and Richard Szeliski
Noah Snavely, Steven M. Seitz, and Richard Szeliski. Photo tourism: Exploring photo collections in 3D. In ACM SIGGRAPH 2006 Papers. ACM, 2006
2006
-
[56]
LatentHOI: On the generalizable hand object motion generation with latent hand diffusion
Muchen Li, Sammy Christen, Chengde Wan, Yu- jun Cai, Renjie Liao, Leonid Sigal, and Shugao Ma. LatentHOI: On the generalizable hand object motion generation with latent hand diffusion. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[57]
3D hand pose estimation in everyday egocentric images
Aditya Prakash, Ruisen Tu, Matthew Chang, and Saurabh Gupta. 3D hand pose estimation in everyday egocentric images. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024
2024
-
[58]
DDF-HO: Hand-held object recon- struction via conditional directed distance field
Chenyangguang Zhang, Yan Di, Ruida Zhang, Guangyao Zhai, Fabian Manhardt, Federico Tombari, and Xiangyang Ji. DDF-HO: Hand-held object recon- struction via conditional directed distance field. In Adv. Neural Inf. Process. Syst. (NeurIPS), volume 36, 2023
2023
-
[59]
Ncrf: neural contact radiance fields for free-viewpoint rendering of hand-object interaction
Zhongqun Zhang, Jifei Song, Eduardo P´erez-Pellitero, Yiren Zhou, Hyung Jin Chang, and Ale ˇs Leonardis. Ncrf: neural contact radiance fields for free-viewpoint rendering of hand-object interaction. InProc. Int. Conf. 3D Vis. (3DV), 2024
2024
-
[60]
Gan- hand: Predicting human grasp affordances in multi- object scenes
Enric Corona, Albert Pumarola, Guillem Alenya, Francesc Moreno-Noguer, and Gr´egory Rogez. Gan- hand: Predicting human grasp affordances in multi- object scenes. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020
2020
-
[61]
Artigrasp: Physically plausible synthesis of bi-manual dexterous grasping and articulation
Hui Zhang, Sammy Christen, Zicong Fan, Luocheng Zheng, Jemin Hwangbo, Jie Song, and Otmar Hilliges. Artigrasp: Physically plausible synthesis of bi-manual dexterous grasping and articulation. InProc. Int. Conf. 3D Vis. (3DV), 2024
2024
-
[62]
Deepsdf: Learning continuous signed distance functions for shape representation
Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2019
2019
-
[63]
A skeleton-driven neural occupancy representation for articulated hands
Korrawe Karunratanakul, Adrian Spurr, Zicong Fan, Otmar Hilliges, and Siyu Tang. A skeleton-driven neural occupancy representation for articulated hands. InProc. Int. Conf. 3D Vis. (3DV), 2021
2021
-
[64]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020. doi: 10.1007/978-3-030-58452-8 \ 24
2020 doi
-
[65]
3D gaussian splat- ting for real-time radiance field rendering.ACM Trans
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk¨uhler, and George Drettakis. 3D gaussian splat- ting for real-time radiance field rendering.ACM Trans. Graph., 42(4), 2023
2023
-
[66]
Affordpose: A large-scale dataset of hand- object interactions with affordance-driven hand pose
Juntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu, and Jian Liu. Affordpose: A large-scale dataset of hand- object interactions with affordance-driven hand pose. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023
2023
-
[67]
Th´eo Morales, Omid Taheri, and Gerard Lacey. A versatile and differentiable hand-object interaction 26 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer representation. InProc. IEEE/CVF Winter Conf. Appl. Comput. ...
2025
-
[68]
Black, and Dima Damen
Siddhant Bansal, Zhifan Zhu, Shashank Tripathi, Ji- ahe Zhao, Michael J. Black, and Dima Damen. To- wards in-the-wild egocentric 3D hand-object pose estimation. InProc. Eur. Conf. Comput. Vis. (ECCV), 2026
2026
-
[69]
HOI4D: A 4D egocentric dataset for category-level human-object interaction
Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4D egocentric dataset for category-level human-object interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2022
2022
-
[70]
Arctic: A dataset for dexterous bimanual hand-object manipulation
Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J Black, and Otmar Hilliges. Arctic: A dataset for dexterous bimanual hand-object manipulation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023
2023
-
[71]
Handoccnet: Occlusion-robust 3D hand mesh estimation network
JoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hong- suk Choi, and Kyoung Mu Lee. Handoccnet: Occlusion-robust 3D hand mesh estimation network. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR), 2022
2022
-
[72]
Mobrecon: Mobile-friendly hand mesh reconstruction from monocular image
Xingyu Chen, Yufeng Liu, Yajiao Dong, Xiong Zhang, Chongyang Ma, Yanmin Xiong, Yuan Zhang, and Xiaoyan Guo. Mobrecon: Mobile-friendly hand mesh reconstruction from monocular image. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[73]
A simple baseline for ef- ficient hand mesh reconstruction
Zhishan Zhou, Shihao Zhou, Zhi Lv, Minqiang Zou, Yao Tang, and Jiajun Liang. A simple baseline for ef- ficient hand mesh reconstruction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[74]
Model-based 3D hand reconstruction via self- supervised learning
Yujin Chen, Zhigang Tu, Di Kang, Linchao Bao, Ying Zhang, Xuefei Zhe, Ruizhi Chen, and Junsong Yuan. Model-based 3D hand reconstruction via self- supervised learning. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2021
2021
-
[75]
Keypoint fusion for RGB-D based 3D hand pose estimation
Xingyu Liu, Pengfei Ren, Yuanyuan Gao, Jingyu Wang, Haifeng Sun, Qi Qi, Zirui Zhuang, and Jianxin Liao. Keypoint fusion for RGB-D based 3D hand pose estimation. InProc. AAAI Conf. Artif. Intell., 2024
2024
-
[76]
Hope-net: A graph-based model for hand-object pose estimation
Bardia Doosti, Shujon Naha, Majid Mirbagheri, and David J Crandall. Hope-net: A graph-based model for hand-object pose estimation. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020
2020
-
[77]
Keypoint transformer: Solv- ing joint identification in challenging hands and ob- ject interactions for accurate 3D pose estimation
Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Keypoint transformer: Solv- ing joint identification in challenging hands and ob- ject interactions for accurate 3D pose estimation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[78]
Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision
Ahmed Tawfik Aboukhadra, Jameel Malik, Ahmed El- hayek, Nadia Robertini, and Didier Stricker. Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision. InProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), 2023
2023
-
[79]
HOISDF: Constraining 3D hand- object pose estimation with global signed distance fields
Haozhe Qi, Chen Zhao, Mathieu Salzmann, and Alexander Mathis. HOISDF: Constraining 3D hand- object pose estimation with global signed distance fields. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[80]
Contactart: Learning 3D interaction priors for category-level ar- ticulated object and hand poses estimation
Zehao Zhu, Jiashun Wang, Yuzhe Qin, Deqing Sun, Varun Jampani, and Xiaolong Wang. Contactart: Learning 3D interaction priors for category-level ar- ticulated object and hand poses estimation. InProc. Int. Conf. 3D Vis. (3DV), 2024
2024
-
[81]
gSDF: Geometry-driven signed dis- tance functions for 3D hand-object reconstruction
Zerui Chen, Shizhe Chen, Cordelia Schmid, and Ivan Laptev. gSDF: Geometry-driven signed dis- tance functions for 3D hand-object reconstruction. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023
2023
-
[82]
Chord: Category-level hand-held object recon- struction via shape deformation
Kailin Li, Lixin Yang, Haoyu Zhen, Zenan Lin, Xinyu Zhan, Licheng Zhong, Jian Xu, Kejian Wu, and Cewu Lu. Chord: Category-level hand-held object recon- struction via shape deformation. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023
2023
-
[83]
Reconstructing hand-held objects from monoc- ular video
Di Huang, Xiaopeng Ji, Xingyi He, Jiaming Sun, Tong He, Qing Shuai, Wanli Ouyang, and Xiaowei Zhou. Reconstructing hand-held objects from monoc- ular video. InProc. ACM SIGGRAPH Asia, 2022
2022
-
[84]
Moho: Learning single-view hand-held object reconstruction with multi-view occlusion-aware su- pervision
Chenyangguang Zhang, Guanlong Jiao, Yan Di, Gu Wang, Ziqin Huang, Ruida Zhang, Fabian Man- hardt, Bowen Fu, Federico Tombari, and Xiangyang Ji. Moho: Learning single-view hand-held object reconstruction with multi-view occlusion-aware su- pervision. InProc. IEEE/CVF Conf. Comp...
2024
-
[85]
Texhoi: Reconstructing textures of 3D unknown ob- jects in monocular hand-object interaction scenes
Alakh Aggarwal, Ningna Wang, and Xiaohu Guo. Texhoi: Reconstructing textures of 3D unknown ob- jects in monocular hand-object interaction scenes. IEEE Trans. Vis. Comput. Graph., 2025. 27 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation...
2025
-
[86]
SeqHAND: RGB-sequence-based 3D hand pose and shape estimation
John Yang, Hyung Jin Chang, Seungeui Lee, and Nojun Kwak. SeqHAND: RGB-sequence-based 3D hand pose and shape estimation. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020
2020
-
[87]
Dyn-hamr: Recovering 4D interacting hand motion from a dynamic camera
Zhengdi Yu, Stefanos Zafeiriou, and Tolga Birdal. Dyn-hamr: Recovering 4D interacting hand motion from a dynamic camera. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[88]
Interactionfusion: real-time reconstruction of hand poses and deformable objects in hand-object interac- tions.ACM Trans
Hao Zhang, Zi-Hao Bo, Jun-Hai Yong, and Feng Xu. Interactionfusion: real-time reconstruction of hand poses and deformable objects in hand-object interac- tions.ACM Trans. Graph., 38(4), 2019
2019
-
[89]
Interaction-aware 4D gaussian splat- ting for dynamic hand-object interaction reconstruc- tion.arXiv preprint arXiv:2511.14540, 2025
Hao Tian, Chenyangguang Zhang, Rui Liu, Wen Shen, and Xiaolin Qin. Interaction-aware 4D gaussian splat- ting for dynamic hand-object interaction reconstruc- tion.arXiv preprint arXiv:2511.14540, 2025
2025 arXiv
-
[90]
Physics-aware hand-object interaction denoising
Haowen Luo, Yunze Liu, and Li Yi. Physics-aware hand-object interaction denoising. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[91]
Grab: A dataset of whole-body hu- man grasping of objects
Omid Taheri, Nima Ghorbani, Michael J Black, and Dimitrios Tzionas. Grab: A dataset of whole-body hu- man grasping of objects. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020
2020
-
[92]
Hand-object contact consistency reason- ing for human grasps generation
Hanwen Jiang, Shaowei Liu, Jiashun Wang, and Xiao- long Wang. Hand-object contact consistency reason- ing for human grasps generation. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021
2021
-
[93]
Contactgen: Generative contact modeling for grasp generation
Shaowei Liu, Yang Zhou, Jimei Yang, Saurabh Gupta, and Shenlong Wang. Contactgen: Generative contact modeling for grasp generation. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023
2023
-
[94]
Contact2grasp: 3D grasp synthesis via hand-object contact constraint
Haoming Li, Xinzhuo Lin, Yang Zhou, Xiang Li, Yuchi Huo, Jiming Chen, and Qi Ye. Contact2grasp: 3D grasp synthesis via hand-object contact constraint. InProc. Int. Joint Conf. Artif. Intell. (IJCAI), 2023. doi: 10.24963/IJCAI.2023/117
2023 doi
-
[95]
G-hop: Generative hand-object prior for interaction reconstruction and grasp synthesis
Yufei Ye, Abhinav Gupta, Kris Kitani, and Shubham Tulsiani. G-hop: Generative hand-object prior for interaction reconstruction and grasp synthesis. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[96]
Clickdiff: click to induce semantic contact map for controllable grasp generation with diffusion models
Peiming Li, Ziyi Wang, Mengyuan Liu, Hong Liu, and Chen Chen. Clickdiff: click to induce semantic contact map for controllable grasp generation with diffusion models. InProc. ACM Int. Conf. Multimedia (ACM MM), 2024
2024
-
[97]
Fastgrasp: Efficient grasp synthesis with diffusion
Xiaofei Wu, Tao Liu, Caoji Li, Yuexin Ma, Yujiao Shi, and Xuming He. Fastgrasp: Efficient grasp synthesis with diffusion. InProc. Int. Conf. 3D Vis. (3DV), 2025
2025
-
[98]
Text2HOI: Text-guided 3D motion gener- ation for hand-object interaction
Junuk Cha, Jihyeon Kim, Jae Shin Yoon, and Seun- gryul Baek. Text2HOI: Text-guided 3D motion gener- ation for hand-object interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[99]
Bimart: A unified approach for the synthesis of 3D bimanual in- teraction with articulated objects
Wanyue Zhang, Rishabh Dabral, Vladislav Golyanik, Vasileios Choutas, Eduardo Alvarado, Thabo Beeler, Marc Habermann, and Christian Theobalt. Bimart: A unified approach for the synthesis of 3D bimanual in- teraction with articulated objects. InProc. IEEE/CVF Conf. Comput. Vis. ...
2025
-
[100]
Manivideo: Generating hand-object ma- nipulation video with dexterous and generalizable grasping
Youxin Pang, Ruizhi Shao, Jiajun Zhang, Hanzhang Tu, Yun Liu, Boyao Zhou, Hongwen Zhang, and Yebin Liu. Manivideo: Generating hand-object ma- nipulation video with dexterous and generalizable grasping. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2025
2025
-
[101]
Hogsa: Bimanual hand-object interaction understanding with 3D gaussian splatting based data augmentation
Wentian Qu, Jiahe Li, Jian Cheng, Jian Shi, Chenyu Meng, Cuixia Ma, Hongan Wang, Xiaoming Deng, and Yinda Zhang. Hogsa: Bimanual hand-object interaction understanding with 3D gaussian splatting based data augmentation. InProc. AAAI Conf. Artif. Intell., 2025
2025
-
[102]
Gears: Local geometry-aware hand-object interaction synthesis
Keyang Zhou, Bharat Lal Bhatnagar, Jan Eric Lenssen, and Gerard Pons-Moll. Gears: Local geometry-aware hand-object interaction synthesis. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[103]
SIGHT: synthesizing image-text conditioned and geometry-guided 3d hand-object trajectories.arXiv preprint arXiv:2503.22869, 2025
Alexey Gavryushin, Alexandros Delitzas, Luc Van Gool, Marc Pollefeys, Kaichun Mo, and Xi Wang. SIGHT: synthesizing image-text conditioned and geometry-guided 3d hand-object trajectories.arXiv preprint arXiv:2503.22869, 2025
2025 arXiv
-
[104]
Gaze-guided hand-object interaction syn- thesis: Dataset and method.arXiv preprint arXiv:2403.16169, 2024
Jie Tian, Ran Ji, Lingxiao Yang, Suting Ni, Yuexin Ma, Lan Xu, Jingyi Yu, Ye Shi, and Jingya Wang. Gaze-guided hand-object interaction syn- thesis: Dataset and method.arXiv preprint arXiv:2403.16169, 2024
2024
-
[105]
How do i do that? synthesizing 3D hand motion and contacts for everyday interactions
Aditya Prakash, Benjamin Lundell, Dmitry Andrey- chuk, David Forsyth, Saurabh Gupta, and Harpreet Sawhney. How do i do that? synthesizing 3D hand motion and contacts for everyday interactions. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025. 28 Hand-Object ...
2025
-
[106]
Geohand: Unlocking prior geome- try knowledge for monocular 3D hand reconstruction
Weiquan Lin, Yaoqing Hu, Liangchen Dai, Xu Tang, and Xingyu Chen. Geohand: Unlocking prior geome- try knowledge for monocular 3D hand reconstruction. arXiv preprint arXiv:2605.17354, 2026
2026 arXiv
-
[107]
MoGe-2: Accurate monoc- ular geometry with metric scale and sharp details
Ruicheng Wang, Sicheng Xu, Yue Dong, Yu Deng, Jianfeng Xiang, Zelong Lv, Guangzhong Sun, Xin Tong, and Jiaolong Yang. MoGe-2: Accurate monoc- ular geometry with metric scale and sharp details. In Adv. Neural Inf. Process. Syst. (NeurIPS), volume 38, 2025
2025
-
[108]
Lisa: Rea- soning segmentation via large language model
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. Lisa: Rea- soning segmentation via large language model. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[109]
Affordance diffusion: Synthe- sizing hand-object interactions
Yufei Ye, Xueting Li, Abhinav Gupta, Shalini De Mello, Stan Birchfield, Jiaming Song, Shubham Tulsiani, and Sifei Liu. Affordance diffusion: Synthe- sizing hand-object interactions. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023
2023
-
[110]
Genie: Text-to-3D generation
Luma AI. Genie: Text-to-3D generation. https:// lumalabs.ai/, 2024. Accessed: Dec. 15, 2024
2024
-
[111]
ChatGPT, 2022
OpenAI. ChatGPT, 2022. URL https://openai. com/blog/chatgpt/. Accessed: 2026-03-24
2022
-
[112]
GPT-4V(ision) system card
OpenAI. GPT-4V(ision) system card. Technical report, OpenAI, 2023. URL https://openai. com/index/gpt-4v-system-card/ . Ac- cessed: Dec. 15, 2024
2023
-
[113]
Follow my hold: Hand-object in- teraction reconstruction through geometric guidance
Ayce Idil Aytekin, Helge Rhodin, Rishabh Dabral, and Christian Theobalt. Follow my hold: Hand-object in- teraction reconstruction through geometric guidance. InProc. Int. Conf. 3D Vis. (3DV), 2026
2026
-
[114]
Hunyuan3d 2.5: Towards high-fidelity 3D assets generation with ultimate details.arXiv preprint arXiv:2506.16504, 2025
Zeqiang Lai, Yunfei Zhao, Haolin Liu, Zibo Zhao, Qingxiang Lin, Huiwen Shi, Xianghui Yang, Mingxin Yang, Shuhui Yang, Yifei Feng, et al. Hunyuan3d 2.5: Towards high-fidelity 3D assets generation with ultimate details.arXiv preprint arXiv:2506.16504, 2025
2025 arXiv
-
[115]
Host3r: Keypoint-free hand-object 3D reconstruction from RGB images
Anilkumar Swamy, Vincent Leroy, Philippe Wein- zaepfel, Jean-S´ebastien Franco, and Gr´egory Rogez. Host3r: Keypoint-free hand-object 3D reconstruction from RGB images. InProc. IEEE/CVF Int. Conf. Comput. Vis. Workshops (ICCVW), 2025
2025
-
[116]
Video depth anything: Consistent depth estimation for super- long videos
Sili Chen, Hengkai Guo, Shengnan Zhu, Feihu Zhang, Zilong Huang, Jiashi Feng, and Bingyi Kang. Video depth anything: Consistent depth estimation for super- long videos. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[117]
Unidepthv2: Universal monocular metric depth estimation made simpler.IEEE Trans
Luigi Piccinelli, Christos Sakaridis, Yung-Hsu Yang, Mattia Segu, Siyuan Li, Wim Abbeloos, and Luc Van Gool. Unidepthv2: Universal monocular metric depth estimation made simpler.IEEE Trans. Pattern Anal. Mach. Intell., 2025
2025
-
[118]
SAM 2: Segment anything in im- ages and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Rong- hang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R ¨adle, Chloe Rolland, Laura Gustafson, et al. SAM 2: Segment anything in im- ages and videos. InProc. Int. Conf. Learn. Represent. (ICLR), 2025
2025
-
[119]
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (...
2024
-
[120]
Hawor: World-space hand motion reconstruction from egocentric videos
Jinglei Zhang, Jiankang Deng, Chao Ma, and Rolan- dos Alexandros Potamias. Hawor: World-space hand motion reconstruction from egocentric videos. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[121]
Metric3d: Towards zero-shot metric 3D prediction from a single image
Wei Yin, Chi Zhang, Hao Chen, Zhipeng Cai, Gang Yu, Kaixuan Wang, Xiaozhi Chen, and Chun Shen. Metric3d: Towards zero-shot metric 3D prediction from a single image. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023
2023
-
[122]
SAM 3D: 3Dfy anything in images
Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang, Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, et al. SAM 3D: 3Dfy anything in images. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026
2026
-
[123]
MV-SAM3D: Adaptive multi-view fusion for layout-aware 3D gen- eration.arXiv preprint arXiv:2603.11633, 2026
Baicheng Li, Dong Wu, Jun Li, Shunkai Zhou, Zecui Zeng, Lusong Li, and Hongbin Zha. MV-SAM3D: Adaptive multi-view fusion for layout-aware 3D gen- eration.arXiv preprint arXiv:2603.11633, 2026
2026 arXiv
-
[124]
SAM 3: Segment anything with concepts.arXiv preprint arXiv:2511.16719, 2025
Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chai- tanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al. SAM 3: Segment anything with concepts.arXiv preprint arXiv:2511.16719, 2025
2025 arXiv
-
[125]
Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647, 2025
Haotong Lin, Sili Chen, Junhao Liew, Donny Y Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647, 2025
2025 arXiv
-
[126]
Handos: 3D hand 29 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer reconstruction in one stage
Xingyu Chen, Zhuheng Song, Xiaoke Jiang, Yaoqing Hu, Junzhi Yu, and Lei Zhang. Handos: 3D hand 29 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer reconstruction in one stage. InProc. IEEE/CVF Conf. Comput. Vis. P...
2025
-
[127]
Choir: Contact-aware 4D hand- object interaction reconstruction.arXiv preprint arXiv:2605.20992, 2026
Hao Xu, Yilin Liu, Yinqiao Wang, Chi-Wing Fu, and Niloy J Mitra. Choir: Contact-aware 4D hand- object interaction reconstruction.arXiv preprint arXiv:2605.20992, 2026
2026 arXiv
-
[128]
Using diffusion priors for video amodal segmenta- tion
Kaihua Chen, Deva Ramanan, and Tarasha Khurana. Using diffusion priors for video amodal segmenta- tion. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[129]
Affordgrasp: Cross-modal diffusion for affordance-aware grasp synthesis
Xiaofei Wu, Yi Zhang, Yumeng Liu, Yuexin Ma, Yu- jiao Shi, and Xuming He. Affordgrasp: Cross-modal diffusion for affordance-aware grasp synthesis. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026
2026
-
[130]
Roberta: A robustly optimized bert pretraining approach.arXiv preprint arXiv:1907.11692, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach.arXiv preprint arXiv:1907.11692, 2019
1907 arXiv
-
[131]
Qwen2 technical re- port.arXiv preprint arXiv:2407.10671, 2024
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze...
2024 arXiv
-
[132]
G-dexgrasp: General- izable dexterous grasping synthesis via part-aware prior retrieval and prior-assisted generation
Juntao Jian, Xiuping Liu, Zixuan Chen, Manyi Li, Jian Liu, and Ruizhen Hu. G-dexgrasp: General- izable dexterous grasping synthesis via part-aware prior retrieval and prior-assisted generation. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025
2025
-
[133]
GPT-4o system card
OpenAI. GPT-4o system card. Technical report, Ope- nAI, 2024. URL https://cdn.openai.com/ gpt-4o-system-card.pdf
2024
-
[134]
Grounded language-image pre-training
Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al. Grounded language-image pre-training. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[135]
NL2Contact: Natural language guided 3D hand- object contact modeling with diffusion model
Zhongqun Zhang, Hengfei Wang, Ziwei Yu, Yi- hua Cheng, Angela Yao, and Hyung Jin Chang. NL2Contact: Natural language guided 3D hand- object contact modeling with diffusion model. In Proc. Eur. Conf. Comput. Vis. (ECCV), 2024
2024
-
[136]
Bert: Pre-training of deep bidi- rectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidi- rectional transformers for language understanding. In Proc. Conf. North Amer. Chapter Assoc. Comput. Lin- guistics: Human Lang. Technol. (NAACL-HLT), 2019
2019
-
[137]
Gonzalez, Ion Stoica, and Eric P
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT qual- ity, March 2023. URL https://lmsys.or...
2023
-
[138]
GPT-4 technical report.arXiv preprint arXiv:2303.08774, 2023
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. GPT-4 technical report.arXiv preprint arXiv:2303.08774, 2023
2023 arXiv
-
[139]
HOIGPT: Learning long-sequence hand-object inter- action with language models
Mingzhen Huang, Fu-Jen Chu, Bugra Tekin, Kevin J Liang, Haoyu Ma, Weiyao Wang, Xingyu Chen, Pierre Gleize, Hongfei Xue, Siwei Lyu, et al. HOIGPT: Learning long-sequence hand-object inter- action with language models. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[140]
Llama: Open and efficient foundation language mod- els.arXiv preprint arXiv:2302.13971, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth ´ee Lacroix, Baptiste Rozi`ere, Naman Goyal, Eric Ham- bro, Faisal Azhar, Aur ´elien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation l...
2023 arXiv
-
[141]
OpenHOI: Open-world hand-object interaction synthesis with multimodal large language model
Zhenhao Zhang, Ye Shi, Lingxiao Yang, Suting Ni, Qi Ye, and Jingya Wang. OpenHOI: Open-world hand-object interaction synthesis with multimodal large language model. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 38, 2025
2025
-
[142]
ShapeLLM: Universal 3D object understanding for embodied interaction
Zekun Qi, Runpei Dong, Shaochen Zhang, Haoran Geng, Chunrui Han, Zheng Ge, Li Yi, and Kaisheng Ma. ShapeLLM: Universal 3D object understanding for embodied interaction. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024. 30 Hand-Object Interaction in the Age of Large Foundation Mode...
2024
-
[143]
Diffh2o: Diffusion-based syn- thesis of hand-object interactions from textual descrip- tions
Sammy Christen, Shreyas Hampali, Fadime Sener, Edoardo Remelli, Tomas Hodan, Eric Sauser, Shugao Ma, and Bugra Tekin. Diffh2o: Diffusion-based syn- thesis of hand-object interactions from textual descrip- tions. InProc. ACM SIGGRAPH Asia, 2024
2024
-
[144]
Forehoi: Feed-forward 3D object reconstruc- tion from daily hand-object interaction videos
Yuantao Chen, Jiahao Chang, Chongjie Ye, Chaoran Zhang, Zhaojie Fang, Chenghong Li, and Xiaoguang Han. Forehoi: Feed-forward 3D object reconstruc- tion from daily hand-object interaction videos. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026
2026
-
[145]
Zero- 1-to-3: Zero-shot one image to 3D object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl V ondrick. Zero- 1-to-3: Zero-shot one image to 3D object. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023
2023
-
[146]
Bigs: Bimanual category-agnostic interaction reconstruction from monocular videos via 3D gaussian splatting
Jeongwan On, Kyeonghwan Gwak, Gunyoung Kang, Junuk Cha, Soohyun Hwang, Hyein Hwang, and Se- ungryul Baek. Bigs: Bimanual category-agnostic interaction reconstruction from monocular videos via 3D gaussian splatting. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[147]
Human universal grasping.arXiv preprint arXiv:2606.17054, 2026
Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, and Lerrel Pinto. Human universal grasping.arXiv preprint arXiv:2606.17054, 2026
2026 arXiv
-
[148]
Hvg-3D: Bridging real and simulation domains for 3D-conditional hand-object interaction video synthesis
Mingjin Chen, Junhao Chen, Zhaoxin Fan, Yujian Lee, Zichen Dang, Lili Wang, Yawen Cui, Lap-Pui Chau, and Yi Wang. Hvg-3D: Bridging real and simulation domains for 3D-conditional hand-object interaction video synthesis. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026
2026
-
[149]
Colin Raffel, Noam Shazeer, Adam Roberts, Kather- ine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of trans- fer learning with a unified text-to-text transformer.J. Mach. Learn. Res., 21(140), 2020
2020
-
[150]
Re-hold: Video hand object interaction reenactment via adaptive layout- instructed diffusion model
Yingying Fan, Quanwei Yang, Kaisiyuan Wang, Hang Zhou, Yingying Li, Haocheng Feng, Errui Ding, Yu Wu, and Jingdong Wang. Re-hold: Video hand object interaction reenactment via adaptive layout- instructed diffusion model. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (C...
2025
-
[151]
GLIDE: To- wards photorealistic image generation and editing with text-guided diffusion models
Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob Mc- Grew, Ilya Sutskever, and Mark Chen. GLIDE: To- wards photorealistic image generation and editing with text-guided diffusion models. InProc. Int. Conf. Mach. Learn. (ICML), 2022
2022
-
[152]
Agile: Hand-object interaction re- construction from video via agentic generation.arXiv preprint arXiv:2602.04672, 2026
Jin-Chuan Shi, Binhong Ye, Tao Liu, Junzhe He, Yangjinhui Xu, Xiaoyang Liu, Zeju Li, Hao Chen, and Chunhua Shen. Agile: Hand-object interaction re- construction from video via agentic generation.arXiv preprint arXiv:2602.04672, 2026
2026 arXiv
-
[153]
VGGT: Visual geometry grounded trans- former
Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. VGGT: Visual geometry grounded trans- former. InProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR), 2025
2025
-
[154]
Efros, and Angjoo Kanazawa
Qianqian Wang, Yifei Zhang, Aleksander Holynski, Alexei A. Efros, and Angjoo Kanazawa. Continu- ous 3D perception model with persistent state. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[155]
Hggt: Robust and flexi- ble 3D hand mesh reconstruction from uncalibrated images.arXiv preprint arXiv:2603.23997, 2026
Yumeng Liu, Xiao-Xiao Long, Marc Habermann, Xu- anze Yang, Cheng Lin, Yuan Liu, Yuexin Ma, Wen- ping Wang, and Ligang Liu. Hggt: Robust and flexi- ble 3D hand mesh reconstruction from uncalibrated images.arXiv preprint arXiv:2603.23997, 2026
2026
-
[156]
Egograsp: World-space hand-object interaction estimation from egocentric videos.arXiv preprint arXiv:2601.01050, 2026
Hongming Fu, Wenjia Wang, Xiaozhen Qiao, Rolan- dos Alexandros Potamias, Taku Komura, Shuo Yang, Zheng Liu, and Bo Zhao. Egograsp: World-space hand-object interaction estimation from egocentric videos.arXiv preprint arXiv:2601.01050, 2026
2026
-
[157]
Hand3r: Online 4D hand-scene reconstruction in the wild.arXiv preprint arXiv:2602.03200, 2026
Wendi Hu, Haonan Zhou, Wenhao Hu, and Gaoang Wang. Hand3r: Online 4D hand-scene reconstruction in the wild.arXiv preprint arXiv:2602.03200, 2026
2026
-
[158]
ScaleHP: Esti- mating hand pose in metric space.arXiv preprint arXiv:2606.25619, 2026
Ruitao Jing, Xingyu Chen, Hongyang Li, Qing Jiang, Yukai Shi, and Lei Zhang. ScaleHP: Esti- mating hand pose in metric space.arXiv preprint arXiv:2606.25619, 2026
2026 arXiv
-
[159]
Qwen technical report.arXiv preprint arXiv:2309.16609, 2023
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report.arXiv preprint arXiv:2309.16609, 2023
2023 arXiv
-
[160]
Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learn- ers. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 33, 2020
2020
-
[161]
Multi-graspllm: A multimodal LLM for multi-hand semantic guided grasp generation.arXiv preprint arXiv:2412.08468, 2024
Haosheng Li, Weixin Mao, Weipeng Deng, Chenyu Meng, Haoqiang Fan, Tiancai Wang, Yoshie Os- amu, Ping Tan, Hongan Wang, and Xiaoming Deng. Multi-graspllm: A multimodal LLM for multi-hand semantic guided grasp generation.arXiv preprint arXiv:2412.08468, 2024. 31 Hand-Object Inte...
2024 arXiv
-
[162]
Afford- dexgrasp: Open-set language-guided dexterous grasp with generalizable-instructive affordance
Yi-Lin Wei, Mu Lin, Yuhao Lin, Jian-Jian Jiang, Xiao- Ming Wu, Ling-An Zeng, and Wei-Shi Zheng. Afford- dexgrasp: Open-set language-guided dexterous grasp with generalizable-instructive affordance. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025
2025
-
[163]
Jeni, and Junhyug Noh
Mingyeong Song, Jungbin Cho, Jisoo Kim, Ananya Bal, Kartik Sharma, Youngjae Yu, Laszlo A. Jeni, and Junhyug Noh. Jointhoi: Jointly generating contact maps enhances hand object interaction generation. arXiv preprint arXiv:2607.01768, 2026
2026 arXiv
-
[164]
Structbihoi: Structured articulation modeling for long–horizon bimanual hand–object interaction generation.arXiv preprint arXiv:2603.08390, 2026
Zhi Wang, Liu Liu, Ruonan Liu, Dan Guo, and Meng Wang. Structbihoi: Structured articulation modeling for long–horizon bimanual hand–object interaction generation.arXiv preprint arXiv:2603.08390, 2026
2026 arXiv
-
[165]
Synhlma: Synthesizing hand language manipulation for articulated object with discrete hu- man object interaction representation.arXiv preprint arXiv:2510.25268, 2025
Zhi Wang, Yuyan Liu, Liu Liu, Li Zhang, Ruixuan Lu, and Dan Guo. Synhlma: Synthesizing hand language manipulation for articulated object with discrete hu- man object interaction representation.arXiv preprint arXiv:2510.25268, 2025
2025
-
[166]
Touch: Text-guided controllable gen- eration of free-form hand-object interactions.arXiv preprint arXiv:2510.14874, 2025
Guangyi Han, Wei Zhai, Yuhang Yang, Yang Cao, and Zheng-Jun Zha. Touch: Text-guided controllable gen- eration of free-form hand-object interactions.arXiv preprint arXiv:2510.14874, 2025
2025
-
[167]
Megohand: Multimodal egocentric hand-object interaction motion generation
Bohan Zhou, Yi Zhan, Zhongbin Zhang, and Zongqing Lu. Megohand: Multimodal egocentric hand-object interaction motion generation. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 38, 2025
2025
-
[168]
Eagle 2: Building post-training data strategies from scratch for frontier vision-language models.arXiv preprint arXiv:2501.14818, 2025
Zhiqi Li, Guo Chen, Shilong Liu, Shihao Wang, VS Vibashan, Yishen Ji, Shiyi Lan, Hao Zhang, Yilin Zhao, Subhashree Radhakrishnan, et al. Eagle 2: Building post-training data strategies from scratch for frontier vision-language models.arXiv preprint arXiv:2501.14818, 2025
2025 arXiv
-
[169]
Training- free dense hand contact estimation with multi- modal large language models.arXiv preprint arXiv:2605.05886, 2026
Daniel Sungho Jung and Kyoung Mu Lee. Training- free dense hand contact estimation with multi- modal large language models.arXiv preprint arXiv:2605.05886, 2026
2026 arXiv
-
[170]
Affordance-guided diffusion prior for 3D hand re- construction.arXiv preprint arXiv:2510.00506, 2025
Naru Suzuki, Takehiko Ohkawa, Tatsuro Banno, Jihyun Lee, Ryosuke Furuta, and Yoichi Sato. Affordance-guided diffusion prior for 3D hand re- construction.arXiv preprint arXiv:2510.00506, 2025
2025 arXiv
-
[171]
Hort: Monocular hand- held objects reconstruction with transformers
Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, and Cordelia Schmid. Hort: Monocular hand- held objects reconstruction with transformers. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025
2025
-
[172]
Grasp as you dream: Imitating functional grasping from generated human demonstrations.arXiv preprint arXiv:2604.07517, 2026
Chao Tang, Jiacheng Xu, Haofei Lu, Bolin Zou, Wen- long Dong, Hong Zhang, and Danica Kragic. Grasp as you dream: Imitating functional grasping from generated human demonstrations.arXiv preprint arXiv:2604.07517, 2026
2026 arXiv
-
[173]
Web2grasp: Learning functional grasps from web images of hand-object interactions.arXiv preprint arXiv:2505.05517, 2025
Hongyi Chen, Yunchao Yao, Yufei Ye, Zhixuan Xu, Homanga Bharadhwaj, Jiashun Wang, Shubham Tulsiani, Zackory Erickson, and Jeffrey Ichnowski. Web2grasp: Learning functional grasps from web images of hand-object interactions.arXiv preprint arXiv:2505.05517, 2025
2025 arXiv
-
[174]
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. InProc. Int. Conf. Learn. Represent. (ICLR), 2024
2024
-
[175]
FLUX.1 kontext: Flow matching for in- context image generation and editing in latent space
Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne, Tim Dockhorn, Jack English, Zion English, Patrick Esser, Sumith Kulal, Kyle Lacey, Yam Levi, Cheng Li, Dominik Lorenz, Jonas M ¨uller, Dustin Podell, Robin Rombach, Harry Sain...
2025 arXiv
-
[176]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023
2023
-
[177]
Single-view image to novel- view generation for hand-object interactions
Zhongqun Zhang, Yihua Cheng, Eduardo P ´erez- Pellitero, Yiren Zhou, Jiankang Deng, Hyung Jin Chang, and Jifei Song. Single-view image to novel- view generation for hand-object interactions. InProc. AAAI Conf. Artif. Intell., 2025
2025
-
[178]
Objaverse-XL: A universe of 10m+ 3D objects
Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram V oleti, Samir Yitzhak Gadre, et al. Objaverse-XL: A universe of 10m+ 3D objects. InAdv. Neural Inf. Process. Syst. (NeurIPS), Datasets and Benchmarks Tr...
2023
-
[179]
Prompt-propose- verify: A reliable hand-object-interaction data gener- ation framework using foundational models.arXiv preprint arXiv:2312.15247, 2023
Gurusha Juneja and Sukrit Kumar. Prompt-propose- verify: A reliable hand-object-interaction data gener- ation framework using foundational models.arXiv preprint arXiv:2312.15247, 2023
2023 arXiv
-
[180]
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. InProc. 32 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generat...
2023
-
[181]
Hand1000: Generating realistic hands from text with only 1,000 images
Haozhuo Zhang, Bin Zhu, Yu Cao, and Yanbin Hao. Hand1000: Generating realistic hands from text with only 1,000 images. InProc. AAAI Conf. Artif. Intell., 2025
2025
-
[182]
Rhands: Re- fining malformed hands for generated images with de- coupled structure and style guidance
Chengrui Wang, Pengfei Liu, Min Zhou, Ming Zeng, Xubin Li, Tiezheng Ge, and Bo Zheng. Rhands: Re- fining malformed hands for generated images with de- coupled structure and style guidance. InProc. AAAI Conf. Artif. Intell., 2025
2025
-
[183]
At- tentionhand: Text-driven controllable hand image gen- eration for 3D hand reconstruction in the wild
Junho Park, Kyeongbo Kong, and Suk-Ju Kang. At- tentionhand: Text-driven controllable hand image gen- eration for 3D hand reconstruction in the wild. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024
2024
-
[184]
Dynamicrafter: An- imating open-domain images with video diffusion priors
Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang, Ying Shan, and Tien-Tsin Wong. Dynamicrafter: An- imating open-domain images with video diffusion priors. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024
2024
-
[185]
Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Wan Team, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
2025 arXiv
-
[186]
Taste-rob: Advancing video generation of task-oriented hand- object interaction for generalizable robotic manipula- tion
Hongxiang Zhao, Xingchen Liu, Mutian Xu, Yiming Hao, Weikai Chen, and Xiaoguang Han. Taste-rob: Advancing video generation of task-oriented hand- object interaction for generalizable robotic manipula- tion. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[187]
idit- HOI: Inpainting-based hand object interaction reenact- ment via video diffusion transformer.arXiv preprint arXiv:2506.12847, 2025
Zhelun Shen, Chenming Wu, Junsheng Zhou, Chen Zhao, Kaisiyuan Wang, Hang Zhou, Yingying Li, Haocheng Feng, Wei He, and Jingdong Wang. idit- HOI: Inpainting-based hand object interaction reenact- ment via video diffusion transformer.arXiv preprint arXiv:2506.12847, 2025
2025 arXiv
-
[188]
Open-world hand- object interaction video generation based on structure and contact-aware representation
Haodong Yan, Hang Yu, Zhide Zhong, Weilin Yuan, Xin Gong, Zehang Luo, Chengxi Heyu, Junfeng Li, Wenxuan Song, Shunbo Zhou, et al. Open-world hand- object interaction video generation based on structure and contact-aware representation. InProc. IEEE/CVF Conf. Comput. Vis. Patte...
2026
-
[189]
PAM: A pose-appearance-motion engine for sim-to-real HOI video generation.arXiv preprint arXiv:2603.22193, 2026
Mingze Gao, Kai Yang, Hongbo Gao, Bo Li, Aoxiang Ding, Wenxiang Li, Yujie Yu, Jian Liu, Shugong Xu, Yi Niu, Haoyu Chi, He Chen, Hui Tang, Li Yi, and Hao Zhao. PAM: A pose-appearance-motion engine for sim-to-real HOI video generation.arXiv preprint arXiv:2603.22193, 2026
2026
-
[190]
Egocentric world model for photorealis- tic hand-object interaction synthesis.arXiv preprint arXiv:2603.13615, 2026
Dayou Li, Lulin Liu, Bangya Liu, Shijie Zhou, Jiu Feng, Ziqi Lu, Minghui Zheng, Chenyu You, and Zhiwen Fan. Egocentric world model for photorealis- tic hand-object interaction synthesis.arXiv preprint arXiv:2603.13615, 2026
2026
-
[191]
Hand2world: Autoregressive egocentric interaction generation via free-space hand gestures.arXiv preprint arXiv:2602.09600, 2026
Yuxi Wang, Wenqi Ouyang, Tianyi Wei, Yi Dong, Zhiqi Shen, and Xingang Pan. Hand2world: Autoregressive egocentric interaction generation via free-space hand gestures.arXiv preprint arXiv:2602.09600, 2026
2026
-
[192]
Generated real- ity: Human-centric world simulation using interactive video generation with hand and camera control
Linxi Xie, Lisong C Sun, Ashley Neall, Tong Wu, Shengqu Cai, and Gordon Wetzstein. Generated real- ity: Human-centric world simulation using interactive video generation with hand and camera control. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026
2026
-
[193]
Dexterous world models
Byungjun Kim, Taeksoo Kim, Junyoung Lee, and Hanbyul Joo. Dexterous world models. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026
2026
-
[194]
Handsonworld: Unconstrained egocentric video generation with camera-disentangled hand control.arXiv preprint arXiv:2607.02075, 2026
Yushuo Chen, Xiaoyu Shi, Xiaoshi Wu, Xintao Wang, Pengfei Wan, and Yebin Liu. Handsonworld: Unconstrained egocentric video generation with camera-disentangled hand control.arXiv preprint arXiv:2607.02075, 2026
2026 arXiv
-
[195]
Wh0: Generative world models as scalable sources of ego- centric human hand manipulation data.arXiv preprint arXiv:2606.22136, 2026
Yangtao Chen, Zixuan Chen, Peiyang Wang, Yong- Lu Li, Jing Huo, Jieqi Shi, and Yang Gao. Wh0: Generative world models as scalable sources of ego- centric human hand manipulation data.arXiv preprint arXiv:2606.22136, 2026
2026 arXiv
-
[196]
Moto: Latent motion token as the bridging language for learning robot manipulation from videos
Yi Chen, Yuying Ge, Weiliang Tang, Yizhuo Li, Yix- iao Ge, Mingyu Ding, Ying Shan, and Xihui Liu. Moto: Latent motion token as the bridging language for learning robot manipulation from videos. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025
2025
-
[197]
Igor: Image-goal representations are the atomic control units for foundation models in embodied ai.arXiv preprint arXiv:2411.00785, 2024
Xiaoyu Chen, Junliang Guo, Tianyu He, Chuheng Zhang, Pushi Zhang, Derek Cathera Yang, Li Zhao, and Jiang Bian. Igor: Image-goal representations are the atomic control units for foundation models in embodied ai.arXiv preprint arXiv:2411.00785, 2024
2024 arXiv
-
[198]
In-n-on: Scaling egocentric manipulation with in-the-wild and on-task data.arXiv preprint arXiv:2511.15704, 2025
Xiongyi Cai, Ri-Zhao Qiu, Geng Chen, Lai Wei, Is- abella Liu, Tianshu Huang, Xuxin Cheng, and Xiao- long Wang. In-n-on: Scaling egocentric manipulation with in-the-wild and on-task data.arXiv preprint arXiv:2511.15704, 2025. 33 Hand-Object Interaction in the Age of Large Found...
2025
-
[199]
Emergence of human to robot trans- fer in vision-language-action models.arXiv preprint arXiv:2512.22414, 2025
Simar Kareer, Karl Pertsch, James Darpinian, Judy Hoffman, Danfei Xu, Sergey Levine, Chelsea Finn, and Suraj Nair. Emergence of human to robot trans- fer in vision-language-action models.arXiv preprint arXiv:2512.22414, 2025
2025
-
[200]
Clap: Contrastive latent action pretraining for learning vision-language-action models from human videos.arXiv preprint arXiv:2601.04061, 2026
Chubin Zhang, Jianan Wang, Zifeng Gao, Yue Su, Tianru Dai, Cai Zhou, Jiwen Lu, and Yansong Tang. Clap: Contrastive latent action pretraining for learning vision-language-action models from human videos.arXiv preprint arXiv:2601.04061, 2026
2026
-
[201]
Univla: Learning to act anywhere with task-centric latent actions.arXiv preprint arXiv:2505.06111, 2025
Qingwen Bu, Yanting Yang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Yao, Ping Luo, and Hongyang Li. Univla: Learning to act anywhere with task-centric latent actions.arXiv preprint arXiv:2505.06111, 2025
2025 arXiv
-
[202]
Villa-x: enhancing latent action modeling in vision-language- action models.arXiv preprint arXiv:2507.23682, 2025
Xiaoyu Chen, Hangxing Wei, Pushi Zhang, Chuheng Zhang, Kaixin Wang, Yanjiang Guo, Rushuai Yang, Yucen Wang, Xinquan Xiao, Li Zhao, et al. Villa-x: enhancing latent action modeling in vision-language- action models.arXiv preprint arXiv:2507.23682, 2025
2025 arXiv
-
[203]
Unleashing large-scale video gen- erative pre-training for visual robot manipulation
Hongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen, Jiafeng Xu, Xinghang Li, Minghuan Liu, Hang Li, and Tao Kong. Unleashing large-scale video gen- erative pre-training for visual robot manipulation. In Proc. Int. Conf. Learn. Represent. (ICLR), 2024
2024
-
[204]
Gr-2: A gen- erative video-language-action model with web-scale knowledge for robot manipulation.arXiv preprint arXiv:2410.06158, 2024
Chi-Lam Cheang, Guangzeng Chen, Ya Jing, Tao Kong, Hang Li, Yifeng Li, Yuxiao Liu, Hongtao Wu, Jiafeng Xu, Yichu Yang, et al. Gr-2: A gen- erative video-language-action model with web-scale knowledge for robot manipulation.arXiv preprint arXiv:2410.06158, 2024
2024 arXiv
-
[205]
GR00T n1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734, 2025
Johan Bjorck, Fernando Casta ˜neda, Nikita Cherni- adev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, et al. GR00T n1: An open foundation model for generalist humanoid robots.arXiv preprint arXiv:2503.14734, 2025
2025 arXiv
-
[206]
Motus: A unified latent action world model
Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, et al. Motus: A unified latent action world model. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026
2026
-
[207]
Egoscale: Scaling dexterous manipulation with diverse egocentric human data.arXiv preprint arXiv:2602.16710, 2026
Ruijie Zheng, Dantong Niu, Yuqi Xie, Jing Wang, Mengda Xu, Yunfan Jiang, Fernando Casta ˜neda, Fengyuan Hu, You Liang Tan, Letian Fu, et al. Egoscale: Scaling dexterous manipulation with diverse egocentric human data.arXiv preprint arXiv:2602.16710, 2026
2026
-
[208]
Being-h0: vision- language-action pretraining from large-scale human videos.arXiv preprint arXiv:2507.15597, 2025
Hao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng, Ye Wang, Haoqi Yuan, Jiazheng Liu, Chaoyi Xu, Qin Jin, and Zongqing Lu. Being-h0: vision- language-action pretraining from large-scale human videos.arXiv preprint arXiv:2507.15597, 2025
2025 arXiv
-
[209]
Scalable vision-language- action model pretraining for robotic manipulation with real-life human activity videos.arXiv preprint arXiv:2510.21571, 2025
Qixiu Li, Yu Deng, Yaobo Liang, Lin Luo, Lei Zhou, Chengtang Yao, Lingqi Zeng, Zhiyuan Feng, Huizhi Liang, Sicheng Xu, et al. Scalable vision-language- action model pretraining for robotic manipulation with real-life human activity videos.arXiv preprint arXiv:2510.21571, 2025
2025
-
[210]
Developing vision-language- action model from egocentric videos.arXiv preprint arXiv:2509.21986, 2025
Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, and Shinsuke Mori. Developing vision-language- action model from egocentric videos.arXiv preprint arXiv:2509.21986, 2025
2025
-
[211]
H-rdt: Human manipulation enhanced bimanual robotic ma- nipulation
Hongzhe Bi, Lingxuan Wu, Tianwei Lin, Hengkai Tan, Zhizhong Su, Hang Su, and Jun Zhu. H-rdt: Human manipulation enhanced bimanual robotic ma- nipulation. InProc. AAAI Conf. Artif. Intell., 2026
2026
-
[212]
Unihm: Unified dexterous hand manipula- tion with vision language model.arXiv preprint arXiv:2603.00732, 2026
Zhenhao Zhang, Jiaxin Liu, Ye Shi, and Jingya Wang. Unihm: Unified dexterous hand manipula- tion with vision language model.arXiv preprint arXiv:2603.00732, 2026
2026
-
[213]
Maniptrans: Efficient dexterous bi- manual manipulation transfer via residual learning
Kailin Li, Puhao Li, Tengyu Liu, Yuyang Li, and Siyuan Huang. Maniptrans: Efficient dexterous bi- manual manipulation transfer via residual learning. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[214]
Dexmachina: Functional retargeting for bimanual dexterous manip- ulation.arXiv preprint arXiv:2505.24853, 2025
Zhao Mandi, Yifan Hou, Dieter Fox, Yashraj Narang, Ajay Mandlekar, and Shuran Song. Dexmachina: Functional retargeting for bimanual dexterous manip- ulation.arXiv preprint arXiv:2505.24853, 2025
2025 arXiv
-
[215]
Learning dexterous manipulation using contact wrench guidance from hu- man demonstration.arXiv preprint arXiv:2607.00033, 2026
Xinghao Zhu, Zixi Liu, Shalin Jain, Chenran Li, Mi- lad Noori, Huihua Zhao, John Welsh, Michael Andres Lin, Wei Liu, Tingwu Wang, et al. Learning dexterous manipulation using contact wrench guidance from hu- man demonstration.arXiv preprint arXiv:2607.00033, 2026
2026 arXiv
-
[216]
Egomimic: Scaling imitation learning via egocentric video
Simar Kareer, Dhruv Patel, Ryan Punamiya, Pranay Mathur, Shuo Cheng, Chen Wang, Judy Hoffman, and Danfei Xu. Egomimic: Scaling imitation learning via egocentric video. InProc. IEEE Int. Conf. Robot. Autom. (ICRA), 2025
2025
-
[217]
Dex- umi: Using human hand as the universal manipulation interface for dexterous manipulation.arXiv preprint arXiv:2505.21864, 2025
Mengda Xu, Han Zhang, Yifan Hou, Zhenjia Xu, Linxi Fan, Manuela Veloso, and Shuran Song. Dex- umi: Using human hand as the universal manipulation interface for dexterous manipulation.arXiv preprint arXiv:2505.21864, 2025. 34 Hand-Object Interaction in the Age of Large Foundati...
2025
-
[218]
Object-centric dexterous manipu- lation from human motion data.arXiv preprint arXiv:2411.04005, 2024
Yuanpei Chen, Chen Wang, Yaodong Yang, and C Karen Liu. Object-centric dexterous manipu- lation from human motion data.arXiv preprint arXiv:2411.04005, 2024
2024 arXiv
-
[219]
Dexterous manipulation policies from RGB human videos via 3D hand-object trajectory recon- struction.arXiv preprint arXiv:2602.09013, 2026
Hongyi Chen, Tony Dong, Tiancheng Wu, Liquan Wang, Yash Jangir, Yaru Niu, Yufei Ye, Homanga Bharadhwaj, Zackory Erickson, and Jeffrey Ich- nowski. Dexterous manipulation policies from RGB human videos via 3D hand-object trajectory recon- struction.arXiv preprint arXiv:2602.09013, 2026
2026
-
[220]
Deximit: Learning bimanual dex- terous manipulation from monocular human videos
Juncheng Mu, Sizhe Yang, Yiming Bao, Hojin Bae, Tianming Wei, Linning Xu, Boyi Li, Huazhe Xu, and Jiangmiao Pang. Deximit: Learning bimanual dex- terous manipulation from monocular human videos. arXiv preprint arXiv:2602.10105, 2026
2026
-
[221]
Masquerade: Learning from in-the-wild human videos using data-editing.arXiv preprint arXiv:2508.09976, 2025
Marion Lepert, Jiaying Fang, and Jeannette Bohg. Masquerade: Learning from in-the-wild human videos using data-editing.arXiv preprint arXiv:2508.09976, 2025
2025 arXiv
-
[222]
You only teach once: Learn one-shot bimanual robotic manipulation from video demonstrations.arXiv preprint arXiv:2501.14208, 2025
Huayi Zhou, Ruixiang Wang, Yunxin Tai, Yueci Deng, Guiliang Liu, and Kui Jia. You only teach once: Learn one-shot bimanual robotic manipulation from video demonstrations.arXiv preprint arXiv:2501.14208, 2025
2025 arXiv
-
[223]
Gat-grasp: Gesture-driven affor- dance transfer for task-aware robotic grasping
Ruixiang Wang, Huayi Zhou, Xinyue Yao, Guiliang Liu, and Kui Jia. Gat-grasp: Gesture-driven affor- dance transfer for task-aware robotic grasping. In Proc. IEEE/RSJ Int. Conf. Intell. Robots Syst. (IROS), 2025
2025
-
[224]
Any- point trajectory modeling for policy learning
Chuan Wen, Xingyu Lin, John Ian Reyes So, Kai Chen, Qi Dou, Yang Gao, and Pieter Abbeel. Any- point trajectory modeling for policy learning. InProc. Robot.: Sci. Syst. (RSS), 2024. doi: 10.15607/RSS. 2024.XX.092
2024 doi
-
[225]
Vidbot: Learning generalizable 3D actions from in-the-wild 2D hu- man videos for zero-shot robotic manipulation
Hanzhi Chen, Boyang Sun, Anran Zhang, Marc Polle- feys, and Stefan Leutenegger. Vidbot: Learning generalizable 3D actions from in-the-wild 2D hu- man videos for zero-shot robotic manipulation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[226]
Flowhoi: Flow-based semantics-grounded generation of hand-object interactions for dexterous robot manip- ulation.arXiv preprint arXiv:2602.13444, 2026
Huajian Zeng, Lingyun Chen, Jiaqi Yang, Yuantai Zhang, Fan Shi, Peidong Liu, and Xingxing Zuo. Flowhoi: Flow-based semantics-grounded generation of hand-object interactions for dexterous robot manip- ulation.arXiv preprint arXiv:2602.13444, 2026
2026
-
[227]
Flow as the cross-domain manipulation interface
Mengda Xu, Zhenjia Xu, Yinghao Xu, Cheng Chi, Gordon Wetzstein, Manuela Veloso, and Shuran Song. Flow as the cross-domain manipulation interface. arXiv preprint arXiv:2407.15208, 2024
2024 arXiv
-
[228]
3DFlowAction: Learning cross-embodiment manip- ulation from 3D flow world model.arXiv preprint arXiv:2506.06199, 2025
Hongyan Zhi, Peihao Chen, Siyuan Zhou, Yubo Dong, Quanxi Wu, Lei Han, and Mingkui Tan. 3DFlowAction: Learning cross-embodiment manip- ulation from 3D flow world model.arXiv preprint arXiv:2506.06199, 2025
2025 arXiv
-
[229]
Novaflow: Zero-shot manipulation via actionable flow from gen- erated videos.arXiv preprint arXiv:2510.08568, 2025
Hongyu Li, Lingfeng Sun, Yafei Hu, Duy Ta, Jennifer Barry, George Konidaris, and Jiahui Fu. Novaflow: Zero-shot manipulation via actionable flow from gen- erated videos.arXiv preprint arXiv:2510.08568, 2025
2025
-
[230]
Dream2flow: Bridg- ing video generation and open-world manipulation with 3D object flow.arXiv preprint arXiv:2512.24766, 2025
Karthik Dharmarajan, Wenlong Huang, Jiajun Wu, Li Fei-Fei, and Ruohan Zhang. Dream2flow: Bridg- ing video generation and open-world manipulation with 3D object flow.arXiv preprint arXiv:2512.24766, 2025
2025
-
[231]
3PoinTr: 3D point tracks for learning ma- nipulation from unconstrained human videos.arXiv preprint arXiv:2603.08485, 2026
Adam Hung, Bardienus Pieter Duisterhof, and Jeffrey Ichnowski. 3PoinTr: 3D point tracks for learning ma- nipulation from unconstrained human videos.arXiv preprint arXiv:2603.08485, 2026
2026 arXiv
-
[232]
Dex4D: Task-agnostic point track policy for sim-to-real dexterous manipulation
Yuxuan Kuang, Sungjae Park, Katerina Fragkiadaki, and Shubham Tulsiani. Dex4D: Task-agnostic point track policy for sim-to-real dexterous manipulation. arXiv preprint arXiv:2602.15828, 2026
2026
-
[233]
Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation
Homanga Bharadhwaj, Debidatta Dwibedi, Abhinav Gupta, Shubham Tulsiani, Carl Doersch, Ted Xiao, Dhruv Shah, Fei Xia, Dorsa Sadigh, and Sean Kir- mani. Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation. arXiv preprint arXiv:2409.16283, 2024
2024 arXiv
-
[234]
Videopoet: A large language model for zero-shot video generation.arXiv preprint arXiv:2312.14125, 2023
Dan Kondratyuk, Lijun Yu, Xiuye Gu, Jos´e Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, et al. Videopoet: A large language model for zero-shot video generation.arXiv preprint arXiv:2312.14125, 2023
2023 arXiv
-
[235]
Robowheel: A data engine from real-world human demonstra- tions for cross-embodiment robotic learning
Yuhong Zhang, Zihan Gao, Shengpeng Li, Ling-Hao Chen, Kaisheng Liu, Runqing Cheng, Xiao Lin, Jun- jia Liu, Zhuoheng Li, Jingyi Feng, et al. Robowheel: A data engine from real-world human demonstra- tions for cross-embodiment robotic learning. In Proc. IEEE/CVF Conf. Comput. Vi...
2026
-
[236]
Bhawna Paliwal, Haritheja Etukuru, William Liang, Pieter Abbeel, Nur Muhammad Mahi Shafiullah, and Jitendra Malik. Do as i do: Dexterous manipulation 35 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer data from e...
2026 arXiv
-
[237]
Egoinfinity: A web-scale 4D hand- object interaction data engine for any-view robot re- targeting and video-to-action robot learning.arXiv preprint arXiv:2606.17385, 2026
Gaotian Wang, Kejia Ren, Andrew Morgan, Yiting Chen, Howard H Qian, Podshara Chanrungmaneekul, and Kaiyu Hang. Egoinfinity: A web-scale 4D hand- object interaction data engine for any-view robot re- targeting and video-to-action robot learning.arXiv preprint arXiv:2606.17385, 2026
2026 arXiv
-
[238]
Human2robot: Learning robot actions from paired human-robot videos
Sicheng Xie, Haidong Cao, Zejia Weng, Zhen Xing, Haoran Chen, Shiwei Shen, Jiaqi Leng, Zuxuan Wu, and Yu-Gang Jiang. Human2robot: Learning robot actions from paired human-robot videos. InProc. AAAI Conf. Artif. Intell., 2026
2026
-
[239]
Tracegen: World modeling in 3D trace-space enables learning from cross-embodiment videos.arXiv preprint arXiv:2511.21690, 2025
Seungjae Lee, Yoonkyo Jung, Inkook Chun, Yao- Chih Lee, Zikui Cai, Hongjia Huang, Aayush Talreja, Tan Dat Dao, Yongyuan Liang, Jia-Bin Huang, and Furong Huang. Tracegen: World modeling in 3D trace-space enables learning from cross-embodiment videos.arXiv preprint arXiv:2511.21...
2025
-
[240]
H2r-grounder: A paired-data-free paradigm for translating human interaction videos into physically grounded robot videos.arXiv preprint arXiv:2512.09406, 2025
Hai Ci, Xiaokang Liu, Pei Yang, Yiren Song, and Mike Zheng Shou. H2r-grounder: A paired-data-free paradigm for translating human interaction videos into physically grounded robot videos.arXiv preprint arXiv:2512.09406, 2025
2025
-
[241]
Qwen-robotmanip tech- nical report: Alignment unlocks scale for robotic manipulation foundation models.arXiv preprint arXiv:2606.17846, 2026
Haoqi Yuan, Zhixuan Liang, Anzhe Chen, Ye Wang, Haoyang Li, Pei Lin, Yiyang Huang, Zixing Lei, Tong Zhang, Jiazhao Zhang, et al. Qwen-robotmanip tech- nical report: Alignment unlocks scale for robotic manipulation foundation models.arXiv preprint arXiv:2606.17846, 2026
2026 arXiv
-
[242]
Real-time joint tracking of a hand manip- ulating an object from RGB-D input
Srinath Sridhar, Franziska Mueller, Michael Zoll- hoefer, Dan Casas, Antti Oulasvirta, and Christian Theobalt. Real-time joint tracking of a hand manip- ulating an object from RGB-D input. InProc. Eur. Conf. Comput. Vis. (ECCV), 2016
2016
-
[243]
Real-time hand tracking under occlusion from an egocentric RGB-D sensor
Franziska Mueller, Dushyant Mehta, Oleksandr Sot- nychenko, Srinath Sridhar, Dan Casas, and Christian Theobalt. Real-time hand tracking under occlusion from an egocentric RGB-D sensor. InProc. IEEE Int. Conf. Comput. Vis. (ICCV), 2017
2017
-
[244]
First-person hand action benchmark with RGB-D videos and 3D hand pose annotations
Guillermo Garcia-Hernando, Shanxin Yuan, Seun- gryul Baek, and Tae-Kyun Kim. First-person hand action benchmark with RGB-D videos and 3D hand pose annotations. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018
2018
-
[245]
Frei- hand: A dataset for markerless capture of hand pose and shape from single RGB images
Christian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan Russell, Max Argus, and Thomas Brox. Frei- hand: A dataset for markerless capture of hand pose and shape from single RGB images. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2019
2019
-
[246]
Contactdb: Analyzing and pre- dicting grasp contact via thermal imaging
Samarth Brahmbhatt, Cusuh Ham, Charles C Kemp, and James Hays. Contactdb: Analyzing and pre- dicting grasp contact via thermal imaging. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019
2019
-
[247]
Ho-3D v3: Improving the accuracy of hand- object annotations of the HO-3D dataset.arXiv preprint arXiv:2107.00887, 2021
Shreyas Hampali, Sayan Deb Sarkar, and Vincent Lepetit. Ho-3D v3: Improving the accuracy of hand- object annotations of the HO-3D dataset.arXiv preprint arXiv:2107.00887, 2021
2021 arXiv
-
[248]
Interhand2.6m: A dataset and baseline for 3D interacting hand pose es- timation from a single RGB image
Gyeongsik Moon, Shoou-I Yu, He Wen, Takaaki Shiratori, and Kyoung Mu Lee. Interhand2.6m: A dataset and baseline for 3D interacting hand pose es- timation from a single RGB image. InProc. Eur. Conf. Comput. Vis. (ECCV), 2020. doi: 10.1007/ 978-3-030-58565-5\ 33
2020
-
[249]
Reconstructing hand-object interactions in the wild
Zhe Cao, Ilija Radosavovic, Angjoo Kanazawa, and Ji- tendra Malik. Reconstructing hand-object interactions in the wild. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021
2021
-
[250]
Dexycb: A benchmark for capturing hand grasping of objects
Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, et al. Dexycb: A benchmark for capturing hand grasping of objects. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2021
2021
-
[251]
H2o: Two hands manipu- lating objects for first person interaction recognition
Taein Kwon, Bugra Tekin, Jan St ¨uhmer, Federica Bogo, and Marc Pollefeys. H2o: Two hands manipu- lating objects for first person interaction recognition. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021
2021
-
[252]
Oakink: A large-scale knowledge repository for understanding hand-object interaction
Lixin Yang, Kailin Li, Xinyu Zhan, Fei Wu, Anran Xu, Liu Liu, and Cewu Lu. Oakink: A large-scale knowledge repository for understanding hand-object interaction. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[253]
Assembly- hands: Towards egocentric activity understanding via 3D hand pose estimation
Takehiko Ohkawa, Kun He, Fadime Sener, Tomas Hodan, Luan Tran, and Cem Keskin. Assembly- hands: Towards egocentric activity understanding via 3D hand pose estimation. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023
2023
-
[254]
Showme: Benchmarking object- agnostic hand-object 3D reconstruction
Anilkumar Swamy, Vincent Leroy, Philippe Wein- zaepfel, Fabien Baradel, Salma Galaaoui, Romain Br´egier, Matthieu Armando, Jean-Sebastien Franco, and Gr´egory Rogez. Showme: Benchmarking object- agnostic hand-object 3D reconstruction. InProc. IEEE/CVF Int. Conf. Comput. Vis. (...
2023
-
[255]
Dense hand- object (ho) graspnet with full grasping taxonomy and dynamics
Woojin Cho, Jihyun Lee, Minjae Yi, Minje Kim, Taeyun Woo, Donghwan Kim, Taewook Ha, Hyokeun Lee, Je-Hwan Ryu, Woontack Woo, et al. Dense hand- object (ho) graspnet with full grasping taxonomy and dynamics. InProc. Eur. Conf. Comput. Vis. (ECCV), 2024
2024
-
[256]
Oakink2: A dataset of bimanual hands-object ma- nipulation in complex task completion
Xinyu Zhan, Lixin Yang, Yifei Zhao, Kangrui Mao, Hanlin Xu, Zenan Lin, Kailin Li, and Cewu Lu. Oakink2: A dataset of bimanual hands-object ma- nipulation in complex task completion. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[257]
Taco: Bench- marking generalizable bimanual tool-action-object understanding
Yun Liu, Haolin Yang, Xu Si, Ling Liu, Zipeng Li, Yuxiang Zhang, Yebin Liu, and Li Yi. Taco: Bench- marking generalizable bimanual tool-action-object understanding. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024
2024
-
[258]
Gigahands: A massive annotated dataset of bimanual hand activi- ties
Rao Fu, Dingxi Zhang, Alex Jiang, Wanjia Fu, Austin Funk, Daniel Ritchie, and Srinath Sridhar. Gigahands: A massive annotated dataset of bimanual hand activi- ties. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2025
2025
-
[259]
Ho-cap: A capture sys- tem and dataset for 3D reconstruction and pose track- ing of hand-object interaction
Jikai Wang, Qifan Zhang, Yu-Wei Chao, Bowen Wen, Xiaohu Guo, and Yu Xiang. Ho-cap: A capture sys- tem and dataset for 3D reconstruction and pose track- ing of hand-object interaction. InAdv. Neural Inf. Process. Syst. (NeurIPS), volume 38, 2025
2025
-
[260]
Anthony, Zhuorui Zhang, and Cewu Lu
Zhenjun Yu, Wenqiang Xu, Pengfei Xie, Yutong Li, Brian W. Anthony, Zhuorui Zhang, and Cewu Lu. Dynamic reconstruction of hand-object interaction with distributed force-aware contact representation. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2025
2025
-
[261]
Hot3d: Hand and object track- ing in 3D from egocentric multi-view videos
Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Shangchen Han, Fan Zhang, Lin- guang Zhang, Jade Fountain, Edward Miller, Se- len Basol, et al. Hot3d: Hand and object track- ing in 3D from egocentric multi-view videos. In Proc. IEEE/CVF Conf. Comput. Vis....
2025
-
[262]
A VI-HT: Adaptive vision-IMU fusion for 3D hand tracking
Ziyi Kou, Ankit Kumar, Mia Huang, Taylor Niehues, Vatsal Mehta, Ergys Ristani, and Li Guan. A VI-HT: Adaptive vision-IMU fusion for 3D hand tracking. arXiv preprint arXiv:2605.21714, 2026
2026 arXiv
-
[263]
Show3d: Capturing scenes of 3D hands and objects in the wild
Patrick Rim, Kevin Harris, Braden Copple, Shangchen Han, Xu Xie, Ivan Shugurov, Sizhe An, He Wen, Alex Wong, Tomas Hodan, et al. Show3d: Capturing scenes of 3D hands and objects in the wild. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026
2026
-
[264]
Handx: Scal- ing bimanual motion and interaction generation
Zimu Zhang, Yucheng Zhang, Xiyan Xu, Ziyin Wang, Sirui Xu, Kai Zhou, Bing Zhou, Chuan Guo, Jian Wang, Yu-Xiong Wang, et al. Handx: Scal- ing bimanual motion and interaction generation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026
2026
-
[265]
Activitynet: A large-scale video benchmark for human activity un- derstanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity un- derstanding. InProc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2015. doi: 10.1109/CVPR.2015. 7298698
2015 doi
-
[266]
Hollywood in homes: Crowdsourcing data collection for activ- ity understanding
Gunnar A Sigurdsson, G¨ul Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. Hollywood in homes: Crowdsourcing data collection for activ- ity understanding. InProc. Eur. Conf. Comput. Vis. (ECCV), 2016
2016
-
[267]
The kinetics human action video dataset.arXiv preprint arXiv:1705.06950, 2017
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset.arXiv preprint arXiv:1705.06950, 2017
2017 arXiv
-
[268]
The ”something something” video database for learn- ing and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fr¨und, Peter Yian- ilos, Moritz Mueller-Freitag, Florian Hoppe, Chris- tian Thurau, Ingo Bax, and Roland Memisevic. The ”something something” video d...
2017 doi
-
[269]
Ava: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, David A Ross, Carl V on- drick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al. Ava: A video dataset of spatio-temporally localized atomic visual actions. In Proc. IEEE/CVF Conf. Comput....
2018
-
[270]
Hacs: Human action clips and segments dataset for recognition and temporal local- ization
Hang Zhao, Antonio Torralba, Lorenzo Torresani, and Zhicheng Yan. Hacs: Human action clips and segments dataset for recognition and temporal local- ization. InProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2019
2019
-
[271]
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding 37 Hand-Object Interaction in the Age of Large Foundation Models: Reconstruction, Generation, and Embodied Transfer by watching hund...
2019
-
[272]
Fin- egym: A hierarchical video dataset for fine-grained action understanding
Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin. Fin- egym: A hierarchical video dataset for fine-grained action understanding. InProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), 2020
2020
-
[273]
The epic-kitchens dataset: Collection, challenges and baselines.IEEE Trans
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. The epic-kitchens dataset: Collection, challenges and baselines.IEEE Trans. Pattern Anal. Mach. Intell., ...
2021
-
[274]
Assembly101: A large-scale multi-view video dataset for understanding procedural activities
Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao. Assembly101: A large-scale multi-view video dataset for understanding procedural activities. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[275]
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022
2022
-
[276]
Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world
Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bugra Tekin, Felipe Vieira Frujeri, et al. Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world. InProc. IEEE/CVF Int. Conf....
2023
-
[277]
Ego-exo4d: Understand- ing skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Tor- resani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al. Ego-exo4d: Understand- ing skilled human activity from first-and third-person perspectives. InProc. IE...
2024
-
[278]
Egodex: Learning dexterous manipulation from large-scale egocentric video.arXiv preprint arXiv:2505.11709, 2025
Ryan Hoque, Peide Huang, David J Yoon, Mouli Siva- purapu, and Jian Zhang. Egodex: Learning dexterous manipulation from large-scale egocentric video.arXiv preprint arXiv:2505.11709, 2025
2025 arXiv
-
[279]
Openego: A large-scale multimodal egocentric dataset for dexterous manipu- lation.arXiv preprint arXiv:2509.05513, 2025
Ahad Jawaid and Yu Xiang. Openego: A large-scale multimodal egocentric dataset for dexterous manipu- lation.arXiv preprint arXiv:2509.05513, 2025
2025 arXiv
-
[280]
Egolive: A large-scale egocentric dataset from real-world human tasks.arXiv preprint arXiv:2604.23570, 2026
Yihang Li, Xuelong Wei, Jingzhou Luo, Yingjing Xiao, Yibo Bai, Guangyuan Zhou, Teng Zou, Chen- guang Gui, Jiajun Wen, He Zhang, et al. Egolive: A large-scale egocentric dataset from real-world human tasks.arXiv preprint arXiv:2604.23570, 2026
2026 arXiv
-
[281]
Egoverse: An egocentric human dataset for robot learning from around the world.arXiv preprint arXiv:2604.07607, 2026
Ryan Punamiya, Simar Kareer, Zeyi Liu, Josh Cit- ron, Ri-Zhao Qiu, Xiongyi Cai, Alexey Gavryushin, Jiaqi Chen, Davide Liconti, Lawrence Y Zhu, et al. Egoverse: An egocentric human dataset for robot learning from around the world.arXiv preprint arXiv:2604.07607, 2026
2026 arXiv
-
[282]
Humannet: Scaling human-centric video learning to one million hours
Yufan Deng and Daquan Zhou. Humannet: Scaling human-centric video learning to one million hours. arXiv preprint arXiv:2605.06747, 2026
2026 arXiv
-
[283]
FEEL (force-enhanced egocentric learning): A dataset for physical action understanding.arXiv preprint arXiv:2603.15847, 2026
Eadom Dessalene, Botao He, Michael Maynord, Yonatan Tussa, Pavan Mantripragada, Yianni Kara- bati, Nirupam Roy, and Yiannis Aloimonos. FEEL (force-enhanced egocentric learning): A dataset for physical action understanding.arXiv preprint arXiv:2603.15847, 2026
2026
-
[284]
Open- AoE: An open egocentric manipulation dataset and toolchain for embodied learning.arXiv preprint arXiv:2607.14183, 2026
Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu, Hao Chen, Qingze Guan, Zhengxing Wu, Wanke Zhan, Yang Sun, Zhiyi Huang, et al. Open- AoE: An open egocentric manipulation dataset and toolchain for embodied learning.arXiv preprint arXiv:2607.14183, 2026
2026 arXiv
-
[285]
Human3.6m: Large scale datasets and predictive methods for 3D human sensing in nat- ural environments.IEEE Trans
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cris- tian Sminchisescu. Human3.6m: Large scale datasets and predictive methods for 3D human sensing in nat- ural environments.IEEE Trans. Pattern Anal. Mach. Intell., 36(7), 2014. doi: 10.1109/TPAMI.2013.248
2014 doi
-
[286]
Model based training, detection and pose estimation of texture-less 3D objects in heavily cluttered scenes
Stefan Hinterstoisser, Vincent Lepetit, Slobodan Ilic, Stefan Holzer, Gary Bradski, Kurt Konolige, and Nassir Navab. Model based training, detection and pose estimation of texture-less 3D objects in heavily cluttered scenes. InProc. Asian Conf. Comput. Vis. (ACCV), 2012
2012
-
[287]
PoseCNN: A convolutional neural network for 6D object pose estimation in cluttered scenes.arXiv preprint arXiv:1711.00199, 2017
Yu Xiang, Tanner Schmidt, Venkatraman Narayanan, and Dieter Fox. PoseCNN: A convolutional neural network for 6D object pose estimation in cluttered scenes.arXiv preprint arXiv:1711.00199, 2017
2017 arXiv
-
[288]
A point set generation network for 3D object reconstruc- tion from a single image
Haoqiang Fan, Hao Su, and Leonidas J Guibas. A point set generation network for 3D object reconstruc- tion from a single image. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2017
2017
-
[289]
GANs trained by a two time-scale update rule con- verge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Un- terthiner, Bernhard Nessler, and Sepp Hochreiter. GANs trained by a two time-scale update rule con- verge to a local nash equilibrium. InAdv. Neural Inf. Process. Syst. (NeurIPS), 2017. 38 Hand-Object Interaction in the Age of Large F...
2017
-
[290]
Towards accurate generative models of video: A new metric & challenges.arXiv preprint arXiv:1812.01717, 2018
Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. Towards accurate generative models of video: A new metric & challenges.arXiv preprint arXiv:1812.01717, 2018
2018 arXiv
-
[291]
Image quality assessment: from error visibility to structural similarity.IEEE Trans
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity.IEEE Trans. Image Process., 13(4), 2004
2004
-
[292]
The unreasonable ef- fectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable ef- fectiveness of deep features as a perceptual metric. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018. 39
2018
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.