REVIEW 3 major objections 6 minor 116 references
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
T0 review · 3 major / 6 minor · reviewed 2026-07-31 · grok-4.5
Pith's one-line read Everyday household skill can be recorded as one time-aligned stream of vision, body, hands, objects, sound, and touch in real homes.
desk verdict Solid dual-scale home capture with true multi-modal GT and honest SOTA gaps; limitations are real but already owned, and the paper deserves referees. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Ambient Capture Engine (ACE): two complementary home configurations (table-scale dense close-range capture; room-scale full-apartment coverage) that share optical-clock temporal sync to the motion-capture timeline and marker-bridged spatial calibration into one world frame, so every modality describes the same physical moment.
What would settle it
Train imitation or VLA policies on ACE-Data-0 and test zero-shot or few-shot transfer in an uninstrumented third home with novel layouts, lighting, and unmarked everyday objects; if performance collapses relative to in-distribution homes, the representativeness claim fails.
Extended reading notes
Core claim
The full perception–action loop of everyday home interaction can be captured holistically: dual-scale ambient instrumentation in real furnished homes yields a single metrically registered, millisecond-aligned multisensory stream (ego and multi-exo video, body and hand motion, object geometry and 6-DoF, audio, tactile), and on the resulting long-horizon ACE-Data-0 benchmark current methods show large, systematic gaps under contact, occlusion, egomotion, and extended time.
Load-bearing premise
Two instrumented homes, a few dozen pre-scanned markered objects, and people wearing suits, gloves, headsets, and markers still produce behavior and visuals representative enough to train general embodied agents.
Editorial extensions
If this is right
- Imitation and policy learning can supervise contact, kinematics, and multi-view vision from the same physical instant rather than stitching mismatched sources.
- World models and VLA systems gain long-horizon household chains with natural subtask order, hesitation, and recovery under goal-level instructions.
- Benchmarks can diagnose failure mode (missed contact vs bad object state vs trajectory drift) instead of only final task success.
- Egocentric and exocentric hand/body estimators can be compared and fused under identical measured ground truth.
- Tactile-from-vision becomes a learnable mapping because every frame is paired with glove pressure on a shared clock.
Reading between the lines
- If suit and marker appearance dominate learned features, domain randomization or appearance stripping may be required before robot transfer works.
- Extending ground truth beyond rigid marked objects to fluids, cloth, and articulated appliances is the natural next bottleneck the dataset itself surfaces.
- Providing measured headset motion as an oracle input would cleanly separate hand-reconstruction error from egomotion error in future ego baselines.
- Scaling sites and unmarked objects may matter more than scaling hours inside the same two layouts for open-world generalization.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces the Ambient Capture Engine (ACE), a dual-scale (table-scale and room-scale) capture system that turns furnished homes into calibrated, time-synchronized multi-sensor studios, and releases ACE-Data-0: ~150 hours / 17M frames / 75,000 episodes across 200 task categories from 50 participants in 2 sites. Streams include egocentric and multi-view exocentric video, full-body and hand motion, object meshes and 6-DoF trajectories, audio, and tactile pressure, registered to a common OptiTrack-referenced timeline and world frame. Tasks are goal-level (atomic HOI, long-horizon HOI chains, HSI) rather than scripted atomic actions. A hierarchical benchmark evaluates tactile-from-vision, human motion recovery, and hand motion from ego/exo views; zero-shot tests of 30+ SOTA methods report large gaps under contact, occlusion, egomotion, and long horizons. The authors position the corpus as supervision for imitation learning, world models, and VLA systems.
Significance. If the capture fidelity, scale, and released annotations hold as described, this is a substantial systems and resource contribution for embodied AI. Prior HOI and egocentric datasets typically fragment viewpoint, modality, or spatial scale; ACE’s explicit dual-scale design plus measured (not purely estimated) body/hand/object/tactile alignment in real homes is a clear advance over lab-only mocap HOI and uninstrumented egocentric video. Strengths include a concrete synchronization and calibration pipeline (§4.1: OptiTrack reference, QR optical clock, hand-eye ego calibration, reported ms residuals and <3 px median reprojection), goal-level long-horizon collection (§4.2), rich measured annotations (§4.3), and unusually broad zero-shot tables (Tables 3–6). For a dataset/benchmark paper, that package is significant even without new learning algorithms.
major comments (3)
- [§5; Abstract; §1; §6] The abstract, introduction, and conclusion repeatedly claim ACE-Data-0 as a “scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI,” yet §5 only reports zero-shot evaluation of external pretrained checkpoints. There is no fine-tuning, imitation, world-model, or VLA experiment that uses ACE-Data-0 as training supervision. For a resource paper this is not fatal, but the training-utility claim is load-bearing in the framing: either add a minimal transfer/finetune study (e.g., hand or body estimator fine-tuned on a train split, or a small behavior-cloning baseline) or substantially tone down claims that currently outrun the evidence in §5.
- [§3.2.2; §5.3.1; Tables 5–6] Hand ground truth is heterogeneous across scales: Manus gloves at room scale versus RANSAC triangulation of 2D keypoints from eight exo cameras plus manual refinement at table scale (§3.2.2, §5.3.1), with joints masked when under-observed. Tables 5–6 and the cross-view analysis (§5.3.3) treat these as a single comparable GT. The paper should quantify residual error of the table-scale pipeline (e.g., against a glove subset or multi-view consistency), state which benchmark splits use which GT source, and confirm that masked joints do not selectively remove the hardest contact frames—otherwise ego/exo and method rankings may partly reflect GT quality rather than viewpoint difficulty.
- [Abstract; §1; Table 1; Limitations (§6)] Domain breadth is a real limit on the “real home / general embodied” claim: only two sites, ~50 pre-scanned markered rigid objects, and participants in mocap suits, gloves, and headsets (Limitations; §3.2; §4.2). The Limitations section already states this, but the main text and Table 1 still present ACE-Data-0 as closing the home-scene gap relative to in-the-wild ego corpora. Please make the scope explicit earlier (abstract/intro/Table 1 caption): instrumented rigid HOI/HSI in two instrumented homes, not open-world domestic diversity—and avoid implying coverage of fluids, articulations, or deformables that are explicitly unannotated.
minor comments (6)
- [Table 1] Table 1 is very dense; a short caption note defining Sync/LH and the checkmark variants (measured vs estimated vs partial) would help readers compare rows without hunting the legend paragraph.
- [§4.2.3; Fig. 8] Dataset statistics: abstract and §4.2.3 say 17M frames / 150 hours / 75,000 episodes; Fig. 8 uses “15M+” in one place. Please reconcile all headline numbers.
- [§5.1.1; Table 3] §5.1.1 metrics (Temp Acc., C-IoU, V-IoU, CoP) need precise definitions or a short appendix (thresholds, min-max aggregation, palm/fingertip mask). Reproducibility of Table 3 depends on them.
- [§4.3.1–4.3.2] Textual annotations rely on Gemini-3.1-pro-preview with human correction (§4.3.2). State approximate correction rate and whether language labels are used in any benchmark track (they appear unused in §5).
- [Project page / §6] Release plan: project and HF links are given; please state what will be public at acceptance (raw vs processed streams, meshes, sync tables, train/test splits, licenses) so the “foundation” claim is checkable.
- [Table 4; References] Minor polish: “W A-MPJPE” spacing in Table 4; a few repeated figure callouts; ensure all baseline citations in Tables 3–6 match the reference list versions used.
Circularity Check
No significant circularity: dataset/capture paper with measured GT and external pretrained baselines.
full rationale
ACE-Data-0 is a systems, capture, and benchmark paper, not a first-principles derivation that recovers fitted constants as predictions. Load-bearing claims are engineering and empirical: dual-scale ambient capture in real homes, hardware/software sync and calibration to a common spatio-temporal frame, goal-level long-horizon collection, and a hierarchical benchmark on which released pretrained SOTA methods fail under contact, occlusion, egomotion, and long horizons. Ground truth is measured (OptiTrack markers, Manus gloves, tactile pressure, scanned/2DGS object meshes), not inferred from the models under test; annotations such as pose reprojections and boxes follow by projection from tracked states. Benchmark tables compare external checkpoints to held-out mocap/tactile GT. Text captions use an external VLM plus human correction—ordinary labeling, not a circular proof of multi-modal alignment. Author self-citations appear in related work and future-application context and are not load-bearing uniqueness or ansatz imports for the central dataset claim. No self-definitional loop, fitted-input-as-prediction, or renaming of a known result as a forced derivation was found.
Assumptions & free parameters
free parameters (3)
- Sync residual / timing model (offset + slow drift line fit to QR clock readings) =
ms-level residuals; ego mutual misalignment <2 ms; calibration re-estimated offsets <4 ms
- Held-out test split size =
10 hours
- Contact/pressure metric thresholds (C-IoU, V-IoU aggregation, CoP region definition)
assumptions (6)
- domain assumption Optical mocap marker trajectories plus glove sensors constitute metric ground truth for body, hand, object pose, and contact pressure in evaluation.
- domain assumption Two furnished homes with dense truss-mounted sensors remain ecologically valid ‘real home’ interaction distributions despite instrumentation.
- domain assumption Goal-level verbal instructions (not step scripts) induce natural long-horizon household behavior suitable for imitation/VLA supervision.
- domain assumption Zero-shot evaluation of officially released pretrained checkpoints fairly exposes method gaps on home long-horizon data.
- standard math Standard multi-view geometry, hand-eye calibration, bundle adjustment, and SMPL-X/MANO representations are adequate to register all streams into one spatio-temporal frame.
- ad hoc to paper Gemini-generated then human-corrected language descriptions are acceptable semantic annotations aligned to measured activity.
invented entities (3)
-
Ambient Capture Engine (ACE) dual-scale studio paradigm
independent evidence
-
ACE-Data-0 corpus and hierarchical signal→component→interaction benchmark
independent evidence
-
ACE-Ego-Head-V02 Lite / ACE-Sense-Glove Lite
Cite this review
Pith. "Pith review of ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine." pith.science (2026). https://pith.science/paper/UJK7YLOX
@misc{pith2026260728625,
author = {Pith},
title = {Pith review of: ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine},
year = {2026},
howpublished = {\url{https://pith.science/paper/UJK7YLOX}},
note = {Machine review of arXiv:2607.28625}
}
read the original abstract
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.
Reference graph
Works this paper leans on
-
[1]
Psychology press, 2014
James J Gibson.The ecological approach to visual perception: classic edition. Psychology press, 2014
2014
-
[2]
MIT press, 2006
Rolf Pfeifer and Josh Bongard.How the body shapes the way we think: a new view of intelligence. MIT press, 2006
2006
-
[3]
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. InCVPR, 2009
2009
-
[4]
LAION-5B: an open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. LAION-5B: an open large-scale dataset for training next generation image-text models. InNeurIPS, 2022
2022
-
[5]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InICML, 2021
2021
-
[6]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhari- wal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Lit...
2020
-
[7]
Ego-Exo4D: understanding skilled human activity from first- and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyl- los Afouras, et al. Ego-Exo4D: understanding skilled human activity from first- and third-person perspectives. InCVPR, 2024
2024
-
[8]
Rescal- ing egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100.IJCV, 130 (1):33–55, 2022
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Rescal- ing egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100.IJCV, 130 (1):33–55, 2022
2022
Show all 116 references
-
[9]
Xperience-10M: A large-scale egocentric multimodal dataset with structured 3D/4D annotations
Ropedia. Xperience-10M: A large-scale egocentric multimodal dataset with structured 3D/4D annotations. Hugging Face dataset, 2026. URLhttps://huggingface.co/datasets/ ropedia-ai/xperience-10m
2026
-
[10]
Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll
Bharat Lal Bhatnagar, Xianghui Xie, Ilya A. Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. BEHA VE: dataset and method for tracking human object interactions. In CVPR, 2022
2022
-
[11]
Black, and Dimitrios Tzionas
Omid Taheri, Nima Ghorbani, Michael J. Black, and Dimitrios Tzionas. GRAB: A dataset of whole-body human grasping of objects. InECCV, 2020
2020
-
[12]
Black, and Otmar Hilliges
Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, and Otmar Hilliges. ARCTIC: A dataset for dexterous bimanual hand-object manipulation. InCVPR, 2023
2023
-
[13]
OAKINK2: A dataset of bimanual hands-object manipulation in complex task completion
Xinyu Zhan, Lixin Yang, Yifei Zhao, Kangrui Mao, Hanlin Xu, Zenan Lin, Kailin Li, and Cewu Lu. OAKINK2: A dataset of bimanual hands-object manipulation in complex task completion. InCVPR, 2024. 29
2024
-
[14]
HOT3D: hand and object tracking in 3D from egocentric multi-view videos
Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Shangchen Han, Fan Zhang, Linguang Zhang, Jade Fountain, Edward Miller, Selen Basol, Richard Newcombe, Robert Wang, Jakob Julian Engel, and Tomas Hodan. HOT3D: hand and object tracking in 3D from egocentric ...
2025
-
[15]
Twigg, Charles C
Samarth Brahmbhatt, Chengcheng Tang, Christopher D. Twigg, Charles C. Kemp, and James Hays. ContactPose: A dataset of grasps with object contact and hand pose. InECCV, 2020
2020
-
[16]
Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Dieter Fox
Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S. Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Dieter Fox. DexYCB: A benchmark for capturing hand grasping of objects. InCVPR, 2021
2021
-
[17]
GigaHands: A massive annotated dataset of bimanual hand activities
Rao Fu, Dingxi Zhang, Alex Jiang, Wanjia Fu, Austin Funk, Daniel Ritchie, and Srinath Sridhar. GigaHands: A massive annotated dataset of bimanual hand activities. InCVPR, 2025
2025
-
[18]
H2O: two hands manipulating objects for first person interaction recognition
Taein Kwon, Bugra Tekin, Jan Stuhmer, Federica Bogo, and Marc Pollefeys. H2O: two hands manipulating objects for first person interaction recognition. InICCV, 2021
2021
-
[19]
World action models are zero-shot policies
Seonghyeon Ye, Yunhao Ge, Kaiyuan Zheng, Shenyuan Gao, Sihyun Yu, George Kurian, Suneel Indupuru, You Liang Tan, Chuning Zhu, Jiannan Xiang, Ayaan Malik, Kyungmin Lee, William Liang, Nadun Ranawaka, Jiasheng Gu, Yinzhen Xu, Guanzhi Wang, Fengyuan Hu, Avnish Narayan, Johan Bjor...
2026 arXiv
-
[20]
EgoVLA: learning vision-language-action models from egocentric human videos
Ruihan Yang, Qinxi Yu, Yecheng Wu, Rui Yan, Borui Li, An-Chieh Cheng, Xueyan Zou, Yunhao Fang, Xuxin Cheng, Ri-Zhao Qiu, Hongxu Yin, Sifei Liu, Song Han, Yao Lu, and Xiaolong Wang. EgoVLA: learning vision-language-action models from egocentric human videos. InCoRL, 2025
2025
-
[21]
Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Robert Equi, Chelsea Finn, Niccolo Fusai, Manuel Y . Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, brian ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey L...
2025
-
[22]
World Guidance: world modeling in condi- tion space for action generation.arXiv 2602.22010, 2026
Yue Su, Sijin Chen, Haixin Shi, Mingyu Liu, Zhengshen Zhang, Ningyuan Huang, Weiheng Zhong, Zhengbang Zhu, Yuxiao Liu, and Xihui Liu. World Guidance: world modeling in condi- tion space for action generation.arXiv 2602.22010, 2026
2026
-
[23]
Ego4D: around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, et al. Ego4D: around the world in 3,000 hours of egocentric video. InCVPR, 2022
2022
-
[24]
HoloAssist: an egocentric human interaction dataset for interactive AI assistants in the real world
Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bugra Tekin, Felipe Vieira Frujeri, Neel Joshi, and Marc Pollefeys. HoloAssist: an egocentric human interaction dataset for interactive AI assistants in the real world. In...
2023
-
[25]
EgoLife: towards egocentric life assistant
Jingkang Yang, Shuai Liu, Hongming Guo, Yuhao Dong, Xiamengwei Zhang, Sicheng Zhang, Pengyun Wang, Zitang Zhou, Binzhu Xie, Ziyue Wang, Bei Ouyang, Zhengyu Lin, Marco Cominelli, Zhongang Cai, Bo Li, Yuanhan Zhang, Peiyuan Zhang, Fangzhou Hong, Joerg Wid- mer, Francesco Gringol...
2025
-
[26]
HD-EPIC: A highly-detailed egocentric video dataset
Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Parida, Kait- ing Liu, Prajwal Gatti, Siddhant Bansal, Kevin Flanagan, Jacob Chalk, Zhifan Zhu, Rhodri Guer- rier, Fahd Abdelazim, Bin Zhu, Davide Moltisanti, Michael Wray, Hazel Doughty, and Dima ...
2025
-
[27]
HOnnotate: A method for 3D annotation of hand and object poses
Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. HOnnotate: A method for 3D annotation of hand and object poses. InCVPR, 2020
2020
-
[28]
OakInk: A large-scale knowledge repository for understanding hand-object interaction
Lixin Yang, Kailin Li, Xinyu Zhan, Fei Wu, Anran Xu, Liu Liu, and Cewu Lu. OakInk: A large-scale knowledge repository for understanding hand-object interaction. InCVPR, 2022
2022
-
[29]
HOI4D: A 4D egocentric dataset for category-level human-object interaction
Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4D egocentric dataset for category-level human-object interaction. InCVPR, 2022
2022
-
[30]
TACO: benchmarking generalizable bimanual Tool-ACtion-Object understanding
Yun Liu, Haolin Yang, Xu Si, Ling Liu, Zipeng Li, Yuxiang Zhang, Yebin Liu, and Li Yi. TACO: benchmarking generalizable bimanual Tool-ACtion-Object understanding. InCVPR, 2024
2024
-
[31]
Black, and Dimitrios Tzionas
Yinghao Huang, Omid Taheri, Michael J. Black, and Dimitrios Tzionas. InterCap: joint marker- less 3D tracking of humans and objects in interaction.IJCV, 132(7):2551–2566, 2024
2024
-
[32]
Full-body articulated human-object interaction
Nan Jiang, Tengyu Liu, Zhexuan Cao, Jieming Cui, Zhiyuan Zhang, Yixin Chen, He Wang, Yixin Zhu, and Siyuan Huang. Full-body articulated human-object interaction. InICCV, 2023
2023
-
[33]
EgoBody: human body shape and motion of interacting people from head- mounted devices
Siwei Zhang, Qianli Ma, Yan Zhang, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, and Siyu Tang. EgoBody: human body shape and motion of interacting people from head- mounted devices. InECCV, 2022
2022
-
[34]
Parkhi, Richard A
Xiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters, Thomas Whelan, Chen Kong, Omkar M. Parkhi, Richard A. Newcombe, and Carl Yuheng Ren. Aria Digital Twin: A new benchmark dataset for egocentric 3D machine perception. InICCV, 2023
2023
-
[35]
Karen Liu
Jiaman Li, Jiajun Wu, and C. Karen Liu. Object motion guided human motion synthesis.ACM TOG, 42(6):197:1–197:11, 2023
2023
-
[36]
HIMO: A new benchmark for full-body human interacting with multiple objects
Xintao Lv, Liang Xu, Yichao Yan, Xin Jin, Congsheng Xu, Shuwen Wu, Yifan Liu, Lincheng Li, Mengxiao Bi, Wenjun Zeng, and Xiaokang Yang. HIMO: A new benchmark for full-body human interacting with multiple objects. InECCV, 2024
2024
-
[37]
Scaling up dynamic human-scene interaction modeling
Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, and Siyuan Huang. Scaling up dynamic human-scene interaction modeling. InCVPR, 2024
2024
-
[38]
ParaHome: parameterizing ev- eryday home activities towards 3D generative modeling of human-object interactions
Jeonghwan Kim, Jisoo Kim, Jeonghyeon Na, and Hanbyul Joo. ParaHome: parameterizing ev- eryday home activities towards 3D generative modeling of human-object interactions. InCVPR, 2024
2024
-
[39]
Karen Liu, Ziwei Liu, Jakob J
Lingni Ma, Yuting Ye, Fangzhou Hong, Vladimir Guzov, Yifeng Jiang, Rowan Postyeni, Luis Pesqueira, Alexander Gamino, Vijay Baiyya, Hyo Jin Kim, Kevin Bailey, David Soriano Fosas, C. Karen Liu, Ziwei Liu, Jakob J. Engel, Renzo De Nardi, and Richard A. Newcombe. Nymeria: A massi...
2024
-
[40]
HUMOTO: A 4D dataset of mocap human object interactions
Jiaxin Lu, Chun-Hao Paul Huang, Uttaran Bhattacharya, Qixing Huang, and Yi Zhou. HUMOTO: A 4D dataset of mocap human object interactions. InICCV, 2025
2025
-
[41]
Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen- Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, Abraham Lee, Kuan Fang, Chelsea Finn, and Sergey Levine
Homer Rich Walke, Kevin Black, Tony Z. Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen- Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, Abraham Lee, Kuan Fang, Chelsea Finn, and Sergey Levine. BridgeData V2: A dataset for robot learning at scale. InCoRL, 2023. 31
2023
-
[42]
RH20T: A comprehensive robotic dataset for learning diverse skills in one-shot
Haoshu Fang, Hongjie Fang, Zhenyu Tang, Jirong Liu, Chenxi Wang, Junbo Wang, Haoyi Zhu, and Cewu Lu. RH20T: A comprehensive robotic dataset for learning diverse skills in one-shot. InICRA, 2024
2024
-
[43]
Karen Liu
Chen Wang, Haochen Shi, Weizhuo Wang, Ruohan Zhang, Li Fei-Fei, and C. Karen Liu. DexCap: scalable and portable mocap data collection system for dexterous manipulation. InRSS, 2024
2024
-
[44]
DROID: A large-scale in-the-wild robot manipulation dataset
Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, et al. DROID: A large-scale in-the-wild robot manipulation dataset. InRSS, 2024
2024
-
[45]
Ag- iBot World Colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems
AgiBot-World-Contributors, Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, et al. Ag- iBot World Colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems. InRSS, 2025
2025
-
[46]
Yoon, Mouli Sivapurapu, and Jian Zhang
Ryan Hoque, Peide Huang, David J. Yoon, Mouli Sivapurapu, and Jian Zhang. EgoDex: learning dexterous manipulation from large-scale egocentric video. InICLR, 2026
2026
-
[47]
Galaxea open-world dataset and G0 dual-system VLA model.arXiv 2509.00576, 2025
Tao Jiang, Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Jianning Cui, Xiao Liu, Shuiqi Cheng, Jiyang Gao, Huazhe Xu, and Hang Zhao. Galaxea open-world dataset and G0 dual-system VLA model.arXiv 2509.00576, 2025
2025 arXiv
-
[48]
EgoScale: scaling dexterous manipulation with diverse egocentric human data.arXiv 2602.16710, 2026
Ruijie Zheng, Dantong Niu, Yuqi Xie, Jing Wang, Mengda Xu, Yunfan Jiang, Fernando Casta˜neda, Fengyuan Hu, You Liang Tan, Letian Fu, Trevor Darrell, Furong Huang, Yuke Zhu, Danfei Xu, and Linxi Fan. EgoScale: scaling dexterous manipulation with diverse egocentric human data.ar...
2026
-
[49]
Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and cap- turing hands and bodies together.ACM TOG, 36(6):245:1–245:17, 2017
2017
-
[50]
Keypoint Transformer: solving joint identification in challenging hands and object interactions for accurate 3D pose estimation
Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Keypoint Transformer: solving joint identification in challenging hands and object interactions for accurate 3D pose estimation. InCVPR, 2022
2022
-
[51]
Reconstructing hands in 3D with transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Ji- tendra Malik. Reconstructing hands in 3D with transformers. InCVPR, 2024
2024
-
[52]
HaWoR: world- space hand motion reconstruction from egocentric videos
Jinglei Zhang, Jiankang Deng, Chao Ma, and Rolandos Alexandros Potamias. HaWoR: world- space hand motion reconstruction from egocentric videos. InCVPR, 2025
2025
-
[53]
Yufei Ye, Yao Feng, Omid Taheri, Haiwen Feng, Shubham Tulsiani, and Michael J. Black. Pre- dicting 4D hand trajectory from monocular videos. In3DV, 2026
2026
-
[54]
HORT: monoc- ular hand-held objects reconstruction with transformers
Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, and Cordelia Schmid. HORT: monoc- ular hand-held objects reconstruction with transformers. InICCV, 2025
2025
-
[55]
PhySIC: physically plausible 3D human-scene interaction and contact from a single image
Pradyumna Yalandur Muralidhar, Yuxuan Xue, Xianghui Xie, Margaret Kostyrko, and Gerard Pons-Moll. PhySIC: physically plausible 3D human-scene interaction and contact from a single image. InSIGGRAPH Asia, 2025
2025
-
[56]
Reconstructing 4D spatial intelligence: A survey.arXiv 2507.21045, 2025
Yukang Cao, Jiahao Lu, Zhisheng Huang, Zhuowen Shen, Chengfeng Zhao, Fangzhou Hong, Zhaoxi Chen, Xin Li, Wenping Wang, Yuan Liu, and Ziwei Liu. Reconstructing 4D spatial intelligence: A survey.arXiv 2507.21045, 2025
2025 arXiv
-
[57]
Humanoid policy∼human policy
Ri-Zhao Qiu, Shiqi Yang, Xuxin Cheng, Chaitanya Chawla, Jialong Li, Tairan He, Ge Yan, Lars Paulsen, Ge Yang, Sha Yi, Guanya Shi, and Xiaolong Wang. Humanoid policy∼human policy. InCoRL, 2025. 32
2025
-
[58]
ActionNet: A dataset for dexterous bimanual manipulation
Fourier ActionNet Team and Yao Mu. ActionNet: A dataset for dexterous bimanual manipulation. Dataset website, 2025. URLhttps://action-net.org/
2025
-
[59]
DexMV: imitation learning for dexterous manipulation from human videos
Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang, Ruihan Yang, Yang Fu, and Xiaolong Wang. DexMV: imitation learning for dexterous manipulation from human videos. InECCV, 2022
2022
-
[60]
EgoMimic: scaling imitation learning via egocentric video
Simar Kareer, Dhruv Patel, Ryan Punamiya, Pranay Mathur, Shuo Cheng, Chen Wang, Judy Hoffman, and Danfei Xu. EgoMimic: scaling imitation learning via egocentric video. InICRA, 2025
2025
-
[61]
Zhu, Simar Kareer, Judy Hoffman, and Danfei Xu
Ryan Punamiya, Dhruv Patel, Patcharapong Aphiwetsa, Pranav Kuppili, Lawrence Y . Zhu, Simar Kareer, Judy Hoffman, and Danfei Xu. EgoBridge: domain adaptation for generalizable imitation from egocentric human data. InNeurIPS, 2025
2025
-
[62]
UniVLA: learning to act anywhere with task-centric latent actions
Qingwen Bu, Yanting Yang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Yao, Ping Luo, and Hongyang Li. UniVLA: learning to act anywhere with task-centric latent actions. InRSS, 2025
2025
-
[63]
In-N-On: scaling egocentric manipulation with in-the-wild and on-task data
Xiongyi Cai, Ri-Zhao Qiu, Geng Chen, Lai Wei, Isabella Liu, Tianshu Huang, Xuxin Cheng, and Xiaolong Wang. In-N-On: scaling egocentric manipulation with in-the-wild and on-task data. arXiv 2511.15704, 2025
2025
-
[64]
Emergence of human to robot transfer in vision-language-action models
Simar Kareer, Karl Pertsch, James Darpinian, Judy Hoffman, Danfei Xu, Sergey Levine, Chelsea Finn, and Suraj Nair. Emergence of human to robot transfer in vision-language-action models. arXiv 2512.22414, 2025
2025
-
[65]
Do as I can, not as I say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, et al. Do as I can, not as I say: Grounding language in robotic affordances. InCoRL, 2022
2022
-
[66]
Ren, Haohuan Wang, Jiaming Tang, Kyle Stachowicz, Karan Dhabalia, Michael Equi, Quan Vuong, Jost Tobias Springenberg, Sergey Levine, Chelsea Finn, and Danny Driess
Marcel Torne, Karl Pertsch, Homer Walke, Kyle Vedder, Suraj Nair, Brian Ichter, Allen Z. Ren, Haohuan Wang, Jiaming Tang, Kyle Stachowicz, Karan Dhabalia, Michael Equi, Quan Vuong, Jost Tobias Springenberg, Sergey Levine, Chelsea Finn, and Danny Driess. MEM: multi-scale embodi...
2026
-
[67]
Bridg- ing semantic and kinematic conditions with diffusion-based discrete motion tokenizer.arXiv 2603.19227, 2026
Chenyang Gu, Mingyuan Zhang, Haozhe Xie, Zhongang Cai, Lei Yang, and Ziwei Liu. Bridg- ing semantic and kinematic conditions with diffusion-based discrete motion tokenizer.arXiv 2603.19227, 2026
2026
-
[68]
InfiniteDance: scalable 3D dance generation towards in-the-wild general- ization
Ronghui Li, Zhongyuan Hu, Li Siyao, Youliang Zhang, Haozhe Xie, Mingyuan Zhang, Jie Guo, Xiu Li, and Ziwei Liu. InfiniteDance: scalable 3D dance generation towards in-the-wild general- ization. InECCV, 2026
2026
-
[69]
Wong, and Ziwei Liu
Yukang Cao, Liang Pan, Kai Han, Kwan-Yee K. Wong, and Ziwei Liu. AvatarGO: zero-shot 4D human-object interaction generation and animation. InICLR, 2025
2025
-
[70]
Zhao, and Chelsea Finn
Zipeng Fu, Tony Z. Zhao, and Chelsea Finn. Mobile ALOHA: Learning bimanual mobile manip- ulation with low-cost whole-body teleoperation. InCoRL, 2024
2024
-
[71]
RoboCOIN: an open-sourced bimanual robotic data COllection for INtegrated manipulation.arXiv 2511.17441, 2025
Shihan Wu, Xuecheng Liu, Shaoxuan Xie, Pengwei Wang, Xinghang Li, et al. RoboCOIN: an open-sourced bimanual robotic data COllection for INtegrated manipulation.arXiv 2511.17441, 2025
2025 arXiv
-
[72]
RoboMIND 2.0: A mul- timodal, bimanual mobile manipulation dataset for generalizable embodied intelligence.arXiv 2512.24653, 2025
Chengkai Hou, Kun Wu, Jiaming Liu, Zhengping Che, Di Wu, et al. RoboMIND 2.0: A mul- timodal, bimanual mobile manipulation dataset for generalizable embodied intelligence.arXiv 2512.24653, 2025. 33
2025
-
[73]
HSImul3R: physics-in-the-loop reconstruction of simulation-ready human-scene interactions
Yukang Cao, Haozhe Xie, Fangzhou Hong, Long Zhuo, Zhaoxi Chen, Liang Pan, and Ziwei Liu. HSImul3R: physics-in-the-loop reconstruction of simulation-ready human-scene interactions. In ECCV, 2026
2026
-
[74]
WholeBodyVLA: towards unified latent VLA for whole-body loco-manipulation control
Haoran Jiang, Jin Chen, Qingwen Bu, Li Chen, Modi Shi, Yanjie Zhang, Delong Li, Chuanzhe Suo, Chuang Wang, Zhihui Peng, and Hongyang Li. WholeBodyVLA: towards unified latent VLA for whole-body loco-manipulation control. InICLR, 2026
2026
-
[75]
DynamicVLA: A vision-language-action model for dynamic object manipulation.arXiv 2601.22153, 2026
Haozhe Xie, Beichen Wen, Jiarui Zheng, Zhaoxi Chen, Fangzhou Hong, Haiwen Diao, and Ziwei Liu. DynamicVLA: A vision-language-action model for dynamic object manipulation.arXiv 2601.22153, 2026
2026
-
[76]
Zhu, Pranav Kuppili, Ryan Punamiya, Patcharapong Aphiwetsa, Dhruv Patel, Simar Kareer, Sehoon Ha, and Danfei Xu
Lawrence Y . Zhu, Pranav Kuppili, Ryan Punamiya, Patcharapong Aphiwetsa, Dhruv Patel, Simar Kareer, Sehoon Ha, and Danfei Xu. EMMA: scaling mobile manipulation via egocentric human data.IEEE RA-L, 11(3):3087–3094, 2025
2025
-
[77]
HoMMI: learning whole-body mobile manipulation from human demonstra- tions.arXiv 2603.03243, 2026
Xiaomeng Xu, Jisang Park, Han Zhang, Eric Cousineau, Aditya Bhat, Jose Barreiros, Dian Wang, and Shuran Song. HoMMI: learning whole-body mobile manipulation from human demonstra- tions.arXiv 2603.03243, 2026
2026 arXiv
-
[78]
2D Gaussian splat- ting for geometrically accurate radiance fields
Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2D Gaussian splat- ting for geometrically accurate radiance fields. InSIGGRAPH, 2024
2024
-
[79]
Garrido-Jurado, R
S. Garrido-Jurado, R. Mu ˜noz-Salinas, F.J. Madrid-Cuevas, and M.J. Mar´ın-Jim´enez. Automatic generation and detection of highly reliable fiducial markers under occlusion.PR, 47(6):2280– 2292, 2014
2014
-
[80]
Tsai and Reimar Lenz
Roger Y . Tsai and Reimar Lenz. A new technique for fully autonomous and efficient 3D robotics hand/eye calibration.IEEE T-RA, 5(3):345–358, 1989
1989
-
[81]
McLauchlan, Richard I
Bill Triggs, Philip F. McLauchlan, Richard I. Hartley, and Andrew W. Fitzgibbon. Bundle adjust- ment — A modern synthesis. InVision Algorithms: Theory and Practice, 2000
2000
-
[82]
Unified temporal and spatial calibration for multi-sensor systems
Paul Furgale, Joern Rehder, and Roland Siegwart. Unified temporal and spatial calibration for multi-sensor systems. InIROS, 2013
2013
-
[83]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. InCVPR, 2019
2019
-
[84]
Gemini: A family of highly capable multimodal models.arXiv 2312.11805, 2025
Gemini Team. Gemini: A family of highly capable multimodal models.arXiv 2312.11805, 2025
2025 arXiv
-
[85]
Twigg, Chengde Wan, James Hays, and Charles C
Patrick Grady, Chengcheng Tang, Samarth Brahmbhatt, Christopher D. Twigg, Chengde Wan, James Hays, and Charles C. Kemp. PressureVision: estimating hand pressure from a single RGB image. InECCV, 2022
2022
-
[86]
EgoTactile: learning grasp pressure for everyday objects from egocentric video
Yuan Zeng, Yujia Shi, Tiao Tan, Xingting Li, Yaqi Qin, Zongqing Lu, Wenming Yang, Jing-Hao Xue, and Qingmin Liao. EgoTactile: learning grasp pressure for everyday objects from egocentric video. InICML, 2026
2026
-
[87]
TouchAnything: A dataset and framework for bimanual tactile estimation from ego- centric video.arXiv 2605.13083, 2026
Jianyi Zhou, Ziteng Gao, Feiyang Hong, Zirui Liu, Guannan Zhang, Weisheng Dai, Ruichen Zhen, Chuqiao Lyu, Haotian Wu, Yinian Mao, Xushi Wang, Yuxiang Jiang, Wenbo Ding, and Shuo Yang. TouchAnything: A dataset and framework for bimanual tactile estimation from ego- centric vide...
2026 arXiv
-
[88]
FoundationPose: unified 6D pose esti- mation and tracking of novel objects
Bowen Wen, Wei Yang, Jan Kautz, and Stan Birchfield. FoundationPose: unified 6D pose esti- mation and tracking of novel objects. InCVPR, 2024. 34
2024
-
[89]
Multi-HMR: multi-person whole-body human mesh recovery in a single shot
Fabien Baradel, Matthieu Armando, Salma Galaaoui, Romain Br ´egier, Philippe Weinzaepfel, Gr´egory Rogez, and Thomas Lucas. Multi-HMR: multi-person whole-body human mesh recovery in a single shot. InECCV, 2024
2024
-
[90]
Multi-HMR 2: multi- person camera-centric human detection, mesh recovery and tracking.arXiv 2606.14841, 2026
Gu ´enol´e Fiche, Philippe Weinzaepfel, Romain Br´egier, and Fabien Baradel. Multi-HMR 2: multi- person camera-centric human detection, mesh recovery and tracking.arXiv 2606.14841, 2026
2026
-
[91]
SAM 3D Body: robust full-body human mesh recovery.arXiv 2602.15989, 2026
Xitong Yang, Devansh Kukreja, Don Pinkus, Anushka Sagar, Taosha Fan, Jinhyung Park, Soyong Shin, Jinkun Cao, Jiawei Liu, Nicolas Ugrinovic, Matt Feiszli, Jitendra Malik, Piotr Doll ´ar, and Kris Kitani. SAM 3D Body: robust full-body human mesh recovery.arXiv 2602.15989, 2026
2026
-
[92]
PyMAF-X: towards well-aligned full-body model regression from monocular images.IEEE TPAMI, 45(10):12287–12303, 2023
Hongwen Zhang, Yating Tian, Yuxiang Zhang, Mengcheng Li, Liang An, Zhenan Sun, and Yebin Liu. PyMAF-X: towards well-aligned full-body model regression from monocular images.IEEE TPAMI, 45(10):12287–12303, 2023
2023
-
[93]
Huang, Otmar Hilliges, and Michael J
Muhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, and Michael J. Black. PARE: part attention regressor for 3D human body estimation. InICCV, 2021
2021
-
[94]
Priyanka Patel and Michael J. Black. CameraHMR: aligning people with perspective. In3DV, 2025
2025
-
[95]
One-stage 3D whole-body mesh recovery with component aware transformer
Jing Lin, Ailing Zeng, Haoqian Wang, Lei Zhang, and Yu Li. One-stage 3D whole-body mesh recovery with component aware transformer. InCVPR, 2023
2023
-
[96]
SMPLer-X: scaling up expressive human pose and shape estimation
Zhongang Cai, Wanqi Yin, Ailing Zeng, Chen Wei, Qingping Sun, Wang Yanjun, Hui En Pang, Haiyi Mei, Mingyuan Zhang, Lei Zhang, Chen Change Loy, Lei Yang, and Ziwei Liu. SMPLer-X: scaling up expressive human pose and shape estimation. InNeurIPS, 2023
2023
-
[97]
SMPLest-X: ultimate scaling for expressive human pose and shape estimation.IEEE TPAMI, 48(2):1778–1794, 2026
Wanqi Yin, Zhongang Cai, Ruisi Wang, Ailing Zeng, Chen Wei, Qingping Sun, Haiyi Mei, Yan- jun Wang, Hui En Pang, Mingyuan Zhang, Lei Zhang, Chen Change Loy, Atsushi Yamashita, Lei Yang, and Ziwei Liu. SMPLest-X: ultimate scaling for expressive human pose and shape estimation.I...
2026
-
[98]
Humans in 4D: Reconstructing and tracking humans with transformers
Shubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa, and Jitendra Ma- lik. Humans in 4D: Reconstructing and tracking humans with transformers. InICCV, 2023
2023
-
[99]
World-grounded human motion recovery via gravity-view coordinates
Zehong Shen, Huaijin Pi, Yan Xia, Zhi Cen, Sida Peng, Zechen Hu, Hujun Bao, Ruizhen Hu, and Xiaowei Zhou. World-grounded human motion recovery via gravity-view coordinates. In SIGGRAPH Asia, 2024
2024
-
[100]
Soyong Shin, Juyong Kim, Eni Halilaj, and Michael J. Black. WHAM: reconstructing world- grounded humans with accurate 3D motion. InCVPR, 2024
2024
-
[101]
GitHub, 2021
EasyMoCap: make human motion capture easier. GitHub, 2021. URLhttps://github. com/zju3dv/EasyMocap
2021
-
[102]
Novel view synthesis of human interactions from sparse multi-view videos
Qing Shuai, Chen Geng, Qi Fang, Sida Peng, Wenhao Shen, Xiaowei Zhou, and Hujun Bao. Novel view synthesis of human interactions from sparse multi-view videos. InSIGGRAPH, 2022
2022
-
[103]
UniSH: unifying scene and human reconstruction in a feed-forward pass.arXiv 2601.01222, 2026
Mengfei Li, Peng Li, Zheng Zhang, Jiahao Lu, Chengfeng Zhao, Wei Xue, Qifeng Liu, Sida Peng, Wenxiao Zhang, Wenhan Luo, Yuan Liu, and Yike Guo. UniSH: unifying scene and human reconstruction in a feed-forward pass.arXiv 2601.01222, 2026
2026
-
[104]
Joint optimization for 4D human-scene reconstruction in the wild.arXiv 2501.02158, 2025
Zhizheng Liu, Joe Lin, Wayne Wu, and Bolei Zhou. Joint optimization for 4D human-scene reconstruction in the wild.arXiv 2501.02158, 2025
2025 arXiv
-
[105]
Hu- man3R: everyone everywhere all at once
Yue Chen, Xingyu Chen, Yuxuan Xue, Anpei Chen, Yuliang Xiu, and Gerard Pons-Moll. Hu- man3R: everyone everywhere all at once. InICLR, 2026. 35
2026
-
[106]
Hanz Cuevas-Velasquez, Anastasios Yiannakidis, Soyong Shin, Giorgio Becherini, Markus H¨oschle, Joachim Tesch, Taylor Obersat, Tsvetelina Alexiadis, and Michael J. Black. MAMMA: markerless and automatic multi-person motion action capture. InCVPR, 2026
2026
-
[107]
Human mesh recovery from arbitrary multi-view images.arXiv 2403.12434, 2024
Xiaoben Li, Mancheng Meng, Ziyan Wu, Terrence Chen, Fan Yang, and Dinggang Shen. Human mesh recovery from arbitrary multi-view images.arXiv 2403.12434, 2024
2024 arXiv
-
[108]
Reconstructing people, places, and cameras
Lea M ¨uller, Hongsuk Choi, Anthony Zhang, Brent Yi, Jitendra Malik, and Angjoo Kanazawa. Reconstructing people, places, and cameras. InCVPR, 2025
2025
-
[109]
Karen Liu, and Jiajun Wu
Jiaman Li, C. Karen Liu, and Jiajun Wu. Ego-body pose estimation via ego-head pose estimation. InCVPR, 2023
2023
-
[110]
Estimating body and hand motion in an ego-sensed world
Brent Yi, Vickie Ye, Maya Zheng, Yunqi Li, Lea M ¨uller, Georgios Pavlakos, Yi Ma, Jitendra Malik, and Angjoo Kanazawa. Estimating body and hand motion in an ego-sensed world. In CVPR, 2025
2025
-
[111]
Human3.6M: large scale datasets and predictive methods for 3D human sensing in natural environments.IEEE TPAMI, 36 (7):1325–1339, 2014
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6M: large scale datasets and predictive methods for 3D human sensing in natural environments.IEEE TPAMI, 36 (7):1325–1339, 2014
2014
-
[112]
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. SMPL: a skinned multi-person linear model.ACM TOG, 34(6):248:1–248:16, 2015
2015
-
[113]
3D hand pose estimation in everyday egocentric images
Aditya Prakash, Ruisen Tu, Matthew Chang, and Saurabh Gupta. 3D hand pose estimation in everyday egocentric images. InECCV, 2024
2024
-
[114]
Dyn-HaMR: recovering 4D interacting hand motion from a dynamic camera
Zhengdi Yu, Stefanos Zafeiriou, and Tolga Birdal. Dyn-HaMR: recovering 4D interacting hand motion from a dynamic camera. InCVPR, 2025
2025
-
[115]
WiLoR: end-to-end 3D hand localization and reconstruction in-the-wild
Rolandos Alexandros Potamias, Jinglei Zhang, Jiankang Deng, and Stefanos Zafeiriou. WiLoR: end-to-end 3D hand localization and reconstruction in-the-wild. InCVPR, 2025
2025
-
[116]
OmniHands: robust motion capture of interactive hands via a versatile trans- former.ACM TOG, 42(6):197:1–197:11, 2026
Dixuan Lin, Yuxiang Zhang, Mengcheng Li, Yebin Liu, Wei Jing, Qi Yan, Qianying Wang, and Hongwen Zhang. OmniHands: robust motion capture of interactive hands via a versatile trans- former.ACM TOG, 42(6):197:1–197:11, 2026. 36
2026
Reviewed July 31, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.