REVIEW 3 major objections 6 minor 116 references
Everyday household skill can be recorded as one time-aligned stream of vision, body, hands, objects, sound, and touch in real homes.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
ACE-Data-0 is a 150-hour home HOI dataset with millisecond-synced ego/exo video, mocap body/hands, object 6-DoF, audio, and tactile signals, plus a three-level benchmark exposing large SOTA gaps.
T0 review reviewed 2026-07-31 challenge →
load-bearing objection Solid dual-scale home capture with true multi-modal GT and honest SOTA gaps; limitations are real but already owned, and the paper deserves referees. the 3 major comments →
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The full perception–action loop of everyday home interaction can be captured holistically: dual-scale ambient instrumentation in real furnished homes yields a single metrically registered, millisecond-aligned multisensory stream (ego and multi-exo video, body and hand motion, object geometry and 6-DoF, audio, tactile), and on the resulting long-horizon ACE-Data-0 benchmark current methods show large, systematic gaps under contact, occlusion, egomotion, and extended time.
What carries the argument
Ambient Capture Engine (ACE): two complementary home configurations (table-scale dense close-range capture; room-scale full-apartment coverage) that share optical-clock temporal sync to the motion-capture timeline and marker-bridged spatial calibration into one world frame, so every modality describes the same physical moment.
Load-bearing premise
Two instrumented homes, a few dozen pre-scanned markered objects, and people wearing suits, gloves, headsets, and markers still produce behavior and visuals representative enough to train general embodied agents.
What would settle it
Train imitation or VLA policies on ACE-Data-0 and test zero-shot or few-shot transfer in an uninstrumented third home with novel layouts, lighting, and unmarked everyday objects; if performance collapses relative to in-distribution homes, the representativeness claim fails.
If this is right
- Imitation and policy learning can supervise contact, kinematics, and multi-view vision from the same physical instant rather than stitching mismatched sources.
- World models and VLA systems gain long-horizon household chains with natural subtask order, hesitation, and recovery under goal-level instructions.
- Benchmarks can diagnose failure mode (missed contact vs bad object state vs trajectory drift) instead of only final task success.
- Egocentric and exocentric hand/body estimators can be compared and fused under identical measured ground truth.
- Tactile-from-vision becomes a learnable mapping because every frame is paired with glove pressure on a shared clock.
Where Pith is reading between the lines
- If suit and marker appearance dominate learned features, domain randomization or appearance stripping may be required before robot transfer works.
- Extending ground truth beyond rigid marked objects to fluids, cloth, and articulated appliances is the natural next bottleneck the dataset itself surfaces.
- Providing measured headset motion as an oracle input would cleanly separate hand-reconstruction error from egomotion error in future ego baselines.
- Scaling sites and unmarked objects may matter more than scaling hours inside the same two layouts for open-world generalization.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces the Ambient Capture Engine (ACE), a dual-scale (table-scale and room-scale) capture system that turns furnished homes into calibrated, time-synchronized multi-sensor studios, and releases ACE-Data-0: ~150 hours / 17M frames / 75,000 episodes across 200 task categories from 50 participants in 2 sites. Streams include egocentric and multi-view exocentric video, full-body and hand motion, object meshes and 6-DoF trajectories, audio, and tactile pressure, registered to a common OptiTrack-referenced timeline and world frame. Tasks are goal-level (atomic HOI, long-horizon HOI chains, HSI) rather than scripted atomic actions. A hierarchical benchmark evaluates tactile-from-vision, human motion recovery, and hand motion from ego/exo views; zero-shot tests of 30+ SOTA methods report large gaps under contact, occlusion, egomotion, and long horizons. The authors position the corpus as supervision for imitation learning, world models, and VLA systems.
Significance. If the capture fidelity, scale, and released annotations hold as described, this is a substantial systems and resource contribution for embodied AI. Prior HOI and egocentric datasets typically fragment viewpoint, modality, or spatial scale; ACE’s explicit dual-scale design plus measured (not purely estimated) body/hand/object/tactile alignment in real homes is a clear advance over lab-only mocap HOI and uninstrumented egocentric video. Strengths include a concrete synchronization and calibration pipeline (§4.1: OptiTrack reference, QR optical clock, hand-eye ego calibration, reported ms residuals and <3 px median reprojection), goal-level long-horizon collection (§4.2), rich measured annotations (§4.3), and unusually broad zero-shot tables (Tables 3–6). For a dataset/benchmark paper, that package is significant even without new learning algorithms.
major comments (3)
- [§5; Abstract; §1; §6] The abstract, introduction, and conclusion repeatedly claim ACE-Data-0 as a “scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI,” yet §5 only reports zero-shot evaluation of external pretrained checkpoints. There is no fine-tuning, imitation, world-model, or VLA experiment that uses ACE-Data-0 as training supervision. For a resource paper this is not fatal, but the training-utility claim is load-bearing in the framing: either add a minimal transfer/finetune study (e.g., hand or body estimator fine-tuned on a train split, or a small behavior-cloning baseline) or substantially tone down claims that currently outrun the evidence in §5.
- [§3.2.2; §5.3.1; Tables 5–6] Hand ground truth is heterogeneous across scales: Manus gloves at room scale versus RANSAC triangulation of 2D keypoints from eight exo cameras plus manual refinement at table scale (§3.2.2, §5.3.1), with joints masked when under-observed. Tables 5–6 and the cross-view analysis (§5.3.3) treat these as a single comparable GT. The paper should quantify residual error of the table-scale pipeline (e.g., against a glove subset or multi-view consistency), state which benchmark splits use which GT source, and confirm that masked joints do not selectively remove the hardest contact frames—otherwise ego/exo and method rankings may partly reflect GT quality rather than viewpoint difficulty.
- [Abstract; §1; Table 1; Limitations (§6)] Domain breadth is a real limit on the “real home / general embodied” claim: only two sites, ~50 pre-scanned markered rigid objects, and participants in mocap suits, gloves, and headsets (Limitations; §3.2; §4.2). The Limitations section already states this, but the main text and Table 1 still present ACE-Data-0 as closing the home-scene gap relative to in-the-wild ego corpora. Please make the scope explicit earlier (abstract/intro/Table 1 caption): instrumented rigid HOI/HSI in two instrumented homes, not open-world domestic diversity—and avoid implying coverage of fluids, articulations, or deformables that are explicitly unannotated.
minor comments (6)
- [Table 1] Table 1 is very dense; a short caption note defining Sync/LH and the checkmark variants (measured vs estimated vs partial) would help readers compare rows without hunting the legend paragraph.
- [§4.2.3; Fig. 8] Dataset statistics: abstract and §4.2.3 say 17M frames / 150 hours / 75,000 episodes; Fig. 8 uses “15M+” in one place. Please reconcile all headline numbers.
- [§5.1.1; Table 3] §5.1.1 metrics (Temp Acc., C-IoU, V-IoU, CoP) need precise definitions or a short appendix (thresholds, min-max aggregation, palm/fingertip mask). Reproducibility of Table 3 depends on them.
- [§4.3.1–4.3.2] Textual annotations rely on Gemini-3.1-pro-preview with human correction (§4.3.2). State approximate correction rate and whether language labels are used in any benchmark track (they appear unused in §5).
- [Project page / §6] Release plan: project and HF links are given; please state what will be public at acceptance (raw vs processed streams, meshes, sync tables, train/test splits, licenses) so the “foundation” claim is checkable.
- [Table 4; References] Minor polish: “W A-MPJPE” spacing in Table 4; a few repeated figure callouts; ensure all baseline citations in Tables 3–6 match the reference list versions used.
Circularity Check
No significant circularity: dataset/capture paper with measured GT and external pretrained baselines.
full rationale
ACE-Data-0 is a systems, capture, and benchmark paper, not a first-principles derivation that recovers fitted constants as predictions. Load-bearing claims are engineering and empirical: dual-scale ambient capture in real homes, hardware/software sync and calibration to a common spatio-temporal frame, goal-level long-horizon collection, and a hierarchical benchmark on which released pretrained SOTA methods fail under contact, occlusion, egomotion, and long horizons. Ground truth is measured (OptiTrack markers, Manus gloves, tactile pressure, scanned/2DGS object meshes), not inferred from the models under test; annotations such as pose reprojections and boxes follow by projection from tracked states. Benchmark tables compare external checkpoints to held-out mocap/tactile GT. Text captions use an external VLM plus human correction—ordinary labeling, not a circular proof of multi-modal alignment. Author self-citations appear in related work and future-application context and are not load-bearing uniqueness or ansatz imports for the central dataset claim. No self-definitional loop, fitted-input-as-prediction, or renaming of a known result as a forced derivation was found.
Axiom & Free-Parameter Ledger
free parameters (3)
- Sync residual / timing model (offset + slow drift line fit to QR clock readings) =
ms-level residuals; ego mutual misalignment <2 ms; calibration re-estimated offsets <4 ms
- Held-out test split size =
10 hours
- Contact/pressure metric thresholds (C-IoU, V-IoU aggregation, CoP region definition)
axioms (6)
- domain assumption Optical mocap marker trajectories plus glove sensors constitute metric ground truth for body, hand, object pose, and contact pressure in evaluation.
- domain assumption Two furnished homes with dense truss-mounted sensors remain ecologically valid ‘real home’ interaction distributions despite instrumentation.
- domain assumption Goal-level verbal instructions (not step scripts) induce natural long-horizon household behavior suitable for imitation/VLA supervision.
- domain assumption Zero-shot evaluation of officially released pretrained checkpoints fairly exposes method gaps on home long-horizon data.
- standard math Standard multi-view geometry, hand-eye calibration, bundle adjustment, and SMPL-X/MANO representations are adequate to register all streams into one spatio-temporal frame.
- ad hoc to paper Gemini-generated then human-corrected language descriptions are acceptable semantic annotations aligned to measured activity.
invented entities (3)
-
Ambient Capture Engine (ACE) dual-scale studio paradigm
independent evidence
-
ACE-Data-0 corpus and hierarchical signal→component→interaction benchmark
independent evidence
-
ACE-Ego-Head-V02 Lite / ACE-Sense-Glove Lite
no independent evidence
Cite this review
Pith. "Pith review of ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine." pith.science (2026). https://pith.science/paper/UJK7YLOX
@misc{pith2026260728625,
author = {Pith},
title = {Pith review of: ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine},
year = {2026},
howpublished = {\url{https://pith.science/paper/UJK7YLOX}},
note = {Machine review of arXiv:2607.28625}
}
read the original abstract
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.
Reference graph
Works this paper leans on
-
[1]
Psychology press, 2014
James J Gibson.The ecological approach to visual perception: classic edition. Psychology press, 2014
2014
-
[2]
MIT press, 2006
Rolf Pfeifer and Josh Bongard.How the body shapes the way we think: a new view of intelligence. MIT press, 2006
2006
-
[3]
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. InCVPR, 2009
2009
-
[4]
LAION-5B: an open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. LAION-5B: an open large-scale dataset for training next generation image-text models. InNeurIPS, 2022
2022
-
[5]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InICML, 2021
2021
-
[6]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhari- wal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Lit...
2020
-
[7]
Ego-Exo4D: understanding skilled human activity from first- and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyl- los Afouras, et al. Ego-Exo4D: understanding skilled human activity from first- and third-person perspectives. InCVPR, 2024
2024
-
[8]
Rescal- ing egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100.IJCV, 130 (1):33–55, 2022
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Rescal- ing egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100.IJCV, 130 (1):33–55, 2022
2022
-
[9]
Xperience-10M: A large-scale egocentric multimodal dataset with structured 3D/4D annotations
Ropedia. Xperience-10M: A large-scale egocentric multimodal dataset with structured 3D/4D annotations. Hugging Face dataset, 2026. URLhttps://huggingface.co/datasets/ ropedia-ai/xperience-10m
2026
-
[10]
Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll
Bharat Lal Bhatnagar, Xianghui Xie, Ilya A. Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. BEHA VE: dataset and method for tracking human object interactions. In CVPR, 2022
2022
-
[11]
Black, and Dimitrios Tzionas
Omid Taheri, Nima Ghorbani, Michael J. Black, and Dimitrios Tzionas. GRAB: A dataset of whole-body human grasping of objects. InECCV, 2020
2020
-
[12]
Black, and Otmar Hilliges
Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, and Otmar Hilliges. ARCTIC: A dataset for dexterous bimanual hand-object manipulation. InCVPR, 2023
2023
-
[13]
OAKINK2: A dataset of bimanual hands-object manipulation in complex task completion
Xinyu Zhan, Lixin Yang, Yifei Zhao, Kangrui Mao, Hanlin Xu, Zenan Lin, Kailin Li, and Cewu Lu. OAKINK2: A dataset of bimanual hands-object manipulation in complex task completion. InCVPR, 2024. 29
2024
-
[14]
HOT3D: hand and object tracking in 3D from egocentric multi-view videos
Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Shangchen Han, Fan Zhang, Linguang Zhang, Jade Fountain, Edward Miller, Selen Basol, Richard Newcombe, Robert Wang, Jakob Julian Engel, and Tomas Hodan. HOT3D: hand and object tracking in 3D from egocentric multi-view videos. InCVPR, 2025
2025
-
[15]
Twigg, Charles C
Samarth Brahmbhatt, Chengcheng Tang, Christopher D. Twigg, Charles C. Kemp, and James Hays. ContactPose: A dataset of grasps with object contact and hand pose. InECCV, 2020
2020
-
[16]
Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Dieter Fox
Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S. Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Dieter Fox. DexYCB: A benchmark for capturing hand grasping of objects. InCVPR, 2021
2021
-
[17]
GigaHands: A massive annotated dataset of bimanual hand activities
Rao Fu, Dingxi Zhang, Alex Jiang, Wanjia Fu, Austin Funk, Daniel Ritchie, and Srinath Sridhar. GigaHands: A massive annotated dataset of bimanual hand activities. InCVPR, 2025
2025
-
[18]
H2O: two hands manipulating objects for first person interaction recognition
Taein Kwon, Bugra Tekin, Jan Stuhmer, Federica Bogo, and Marc Pollefeys. H2O: two hands manipulating objects for first person interaction recognition. InICCV, 2021
2021
-
[19]
World action models are zero-shot policies
Seonghyeon Ye, Yunhao Ge, Kaiyuan Zheng, Shenyuan Gao, Sihyun Yu, George Kurian, Suneel Indupuru, You Liang Tan, Chuning Zhu, Jiannan Xiang, Ayaan Malik, Kyungmin Lee, William Liang, Nadun Ranawaka, Jiasheng Gu, Yinzhen Xu, Guanzhi Wang, Fengyuan Hu, Avnish Narayan, Johan Bjorck, Jing Wang, Gwanghyun Kim, Dantong Niu, Ruijie Zheng, Yuqi Xie, Jimmy Wu, Qi ...
Pith/arXiv arXiv 2026
-
[20]
EgoVLA: learning vision-language-action models from egocentric human videos
Ruihan Yang, Qinxi Yu, Yecheng Wu, Rui Yan, Borui Li, An-Chieh Cheng, Xueyan Zou, Yunhao Fang, Xuxin Cheng, Ri-Zhao Qiu, Hongxu Yin, Sifei Liu, Song Han, Yao Lu, and Xiaolong Wang. EgoVLA: learning vision-language-action models from egocentric human videos. InCoRL, 2025
2025
-
[21]
Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Robert Equi, Chelsea Finn, Niccolo Fusai, Manuel Y . Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, brian ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Z. Ren,...
2025
-
[22]
World Guidance: world modeling in condi- tion space for action generation.arXiv 2602.22010, 2026
Yue Su, Sijin Chen, Haixin Shi, Mingyu Liu, Zhengshen Zhang, Ningyuan Huang, Weiheng Zhong, Zhengbang Zhu, Yuxiao Liu, and Xihui Liu. World Guidance: world modeling in condi- tion space for action generation.arXiv 2602.22010, 2026
arXiv 2026
-
[23]
Ego4D: around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, et al. Ego4D: around the world in 3,000 hours of egocentric video. InCVPR, 2022
2022
-
[24]
HoloAssist: an egocentric human interaction dataset for interactive AI assistants in the real world
Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bugra Tekin, Felipe Vieira Frujeri, Neel Joshi, and Marc Pollefeys. HoloAssist: an egocentric human interaction dataset for interactive AI assistants in the real world. InICCV, 2023
2023
-
[25]
EgoLife: towards egocentric life assistant
Jingkang Yang, Shuai Liu, Hongming Guo, Yuhao Dong, Xiamengwei Zhang, Sicheng Zhang, Pengyun Wang, Zitang Zhou, Binzhu Xie, Ziyue Wang, Bei Ouyang, Zhengyu Lin, Marco Cominelli, Zhongang Cai, Bo Li, Yuanhan Zhang, Peiyuan Zhang, Fangzhou Hong, Joerg Wid- mer, Francesco Gringoli, Lei Yang, and Ziwei Liu. EgoLife: towards egocentric life assistant. In CVPR,...
2025
-
[26]
HD-EPIC: A highly-detailed egocentric video dataset
Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Parida, Kait- ing Liu, Prajwal Gatti, Siddhant Bansal, Kevin Flanagan, Jacob Chalk, Zhifan Zhu, Rhodri Guer- rier, Fahd Abdelazim, Bin Zhu, Davide Moltisanti, Michael Wray, Hazel Doughty, and Dima Damen. HD-EPIC: A highly-detailed egocentric video dataset. InCVPR, 2025
2025
-
[27]
HOnnotate: A method for 3D annotation of hand and object poses
Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. HOnnotate: A method for 3D annotation of hand and object poses. InCVPR, 2020
2020
-
[28]
OakInk: A large-scale knowledge repository for understanding hand-object interaction
Lixin Yang, Kailin Li, Xinyu Zhan, Fei Wu, Anran Xu, Liu Liu, and Cewu Lu. OakInk: A large-scale knowledge repository for understanding hand-object interaction. InCVPR, 2022
2022
-
[29]
HOI4D: A 4D egocentric dataset for category-level human-object interaction
Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4D egocentric dataset for category-level human-object interaction. InCVPR, 2022
2022
-
[30]
TACO: benchmarking generalizable bimanual Tool-ACtion-Object understanding
Yun Liu, Haolin Yang, Xu Si, Ling Liu, Zipeng Li, Yuxiang Zhang, Yebin Liu, and Li Yi. TACO: benchmarking generalizable bimanual Tool-ACtion-Object understanding. InCVPR, 2024
2024
-
[31]
Black, and Dimitrios Tzionas
Yinghao Huang, Omid Taheri, Michael J. Black, and Dimitrios Tzionas. InterCap: joint marker- less 3D tracking of humans and objects in interaction.IJCV, 132(7):2551–2566, 2024
2024
-
[32]
Full-body articulated human-object interaction
Nan Jiang, Tengyu Liu, Zhexuan Cao, Jieming Cui, Zhiyuan Zhang, Yixin Chen, He Wang, Yixin Zhu, and Siyuan Huang. Full-body articulated human-object interaction. InICCV, 2023
2023
-
[33]
EgoBody: human body shape and motion of interacting people from head- mounted devices
Siwei Zhang, Qianli Ma, Yan Zhang, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, and Siyu Tang. EgoBody: human body shape and motion of interacting people from head- mounted devices. InECCV, 2022
2022
-
[34]
Parkhi, Richard A
Xiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters, Thomas Whelan, Chen Kong, Omkar M. Parkhi, Richard A. Newcombe, and Carl Yuheng Ren. Aria Digital Twin: A new benchmark dataset for egocentric 3D machine perception. InICCV, 2023
2023
-
[35]
Karen Liu
Jiaman Li, Jiajun Wu, and C. Karen Liu. Object motion guided human motion synthesis.ACM TOG, 42(6):197:1–197:11, 2023
2023
-
[36]
HIMO: A new benchmark for full-body human interacting with multiple objects
Xintao Lv, Liang Xu, Yichao Yan, Xin Jin, Congsheng Xu, Shuwen Wu, Yifan Liu, Lincheng Li, Mengxiao Bi, Wenjun Zeng, and Xiaokang Yang. HIMO: A new benchmark for full-body human interacting with multiple objects. InECCV, 2024
2024
-
[37]
Scaling up dynamic human-scene interaction modeling
Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, and Siyuan Huang. Scaling up dynamic human-scene interaction modeling. InCVPR, 2024
2024
-
[38]
ParaHome: parameterizing ev- eryday home activities towards 3D generative modeling of human-object interactions
Jeonghwan Kim, Jisoo Kim, Jeonghyeon Na, and Hanbyul Joo. ParaHome: parameterizing ev- eryday home activities towards 3D generative modeling of human-object interactions. InCVPR, 2024
2024
-
[39]
Karen Liu, Ziwei Liu, Jakob J
Lingni Ma, Yuting Ye, Fangzhou Hong, Vladimir Guzov, Yifeng Jiang, Rowan Postyeni, Luis Pesqueira, Alexander Gamino, Vijay Baiyya, Hyo Jin Kim, Kevin Bailey, David Soriano Fosas, C. Karen Liu, Ziwei Liu, Jakob J. Engel, Renzo De Nardi, and Richard A. Newcombe. Nymeria: A massive collection of multimodal egocentric daily motion in the wild. InECCV, 2024
2024
-
[40]
HUMOTO: A 4D dataset of mocap human object interactions
Jiaxin Lu, Chun-Hao Paul Huang, Uttaran Bhattacharya, Qixing Huang, and Yi Zhou. HUMOTO: A 4D dataset of mocap human object interactions. InICCV, 2025
2025
-
[41]
Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen- Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, Abraham Lee, Kuan Fang, Chelsea Finn, and Sergey Levine
Homer Rich Walke, Kevin Black, Tony Z. Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen- Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, Abraham Lee, Kuan Fang, Chelsea Finn, and Sergey Levine. BridgeData V2: A dataset for robot learning at scale. InCoRL, 2023. 31
2023
-
[42]
RH20T: A comprehensive robotic dataset for learning diverse skills in one-shot
Haoshu Fang, Hongjie Fang, Zhenyu Tang, Jirong Liu, Chenxi Wang, Junbo Wang, Haoyi Zhu, and Cewu Lu. RH20T: A comprehensive robotic dataset for learning diverse skills in one-shot. InICRA, 2024
2024
-
[43]
Karen Liu
Chen Wang, Haochen Shi, Weizhuo Wang, Ruohan Zhang, Li Fei-Fei, and C. Karen Liu. DexCap: scalable and portable mocap data collection system for dexterous manipulation. InRSS, 2024
2024
-
[44]
DROID: A large-scale in-the-wild robot manipulation dataset
Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, et al. DROID: A large-scale in-the-wild robot manipulation dataset. InRSS, 2024
2024
-
[45]
Ag- iBot World Colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems
AgiBot-World-Contributors, Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, et al. Ag- iBot World Colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems. InRSS, 2025
2025
-
[46]
Yoon, Mouli Sivapurapu, and Jian Zhang
Ryan Hoque, Peide Huang, David J. Yoon, Mouli Sivapurapu, and Jian Zhang. EgoDex: learning dexterous manipulation from large-scale egocentric video. InICLR, 2026
2026
-
[47]
Galaxea open-world dataset and G0 dual-system VLA model.arXiv 2509.00576, 2025
Tao Jiang, Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Jianning Cui, Xiao Liu, Shuiqi Cheng, Jiyang Gao, Huazhe Xu, and Hang Zhao. Galaxea open-world dataset and G0 dual-system VLA model.arXiv 2509.00576, 2025
Pith/arXiv arXiv 2025
-
[48]
EgoScale: scaling dexterous manipulation with diverse egocentric human data.arXiv 2602.16710, 2026
Ruijie Zheng, Dantong Niu, Yuqi Xie, Jing Wang, Mengda Xu, Yunfan Jiang, Fernando Casta˜neda, Fengyuan Hu, You Liang Tan, Letian Fu, Trevor Darrell, Furong Huang, Yuke Zhu, Danfei Xu, and Linxi Fan. EgoScale: scaling dexterous manipulation with diverse egocentric human data.arXiv 2602.16710, 2026
arXiv 2026
-
[49]
Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and cap- turing hands and bodies together.ACM TOG, 36(6):245:1–245:17, 2017
2017
-
[50]
Keypoint Transformer: solving joint identification in challenging hands and object interactions for accurate 3D pose estimation
Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Keypoint Transformer: solving joint identification in challenging hands and object interactions for accurate 3D pose estimation. InCVPR, 2022
2022
-
[51]
Reconstructing hands in 3D with transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Ji- tendra Malik. Reconstructing hands in 3D with transformers. InCVPR, 2024
2024
-
[52]
HaWoR: world- space hand motion reconstruction from egocentric videos
Jinglei Zhang, Jiankang Deng, Chao Ma, and Rolandos Alexandros Potamias. HaWoR: world- space hand motion reconstruction from egocentric videos. InCVPR, 2025
2025
-
[53]
Yufei Ye, Yao Feng, Omid Taheri, Haiwen Feng, Shubham Tulsiani, and Michael J. Black. Pre- dicting 4D hand trajectory from monocular videos. In3DV, 2026
2026
-
[54]
HORT: monoc- ular hand-held objects reconstruction with transformers
Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, and Cordelia Schmid. HORT: monoc- ular hand-held objects reconstruction with transformers. InICCV, 2025
2025
-
[55]
PhySIC: physically plausible 3D human-scene interaction and contact from a single image
Pradyumna Yalandur Muralidhar, Yuxuan Xue, Xianghui Xie, Margaret Kostyrko, and Gerard Pons-Moll. PhySIC: physically plausible 3D human-scene interaction and contact from a single image. InSIGGRAPH Asia, 2025
2025
-
[56]
Reconstructing 4D spatial intelligence: A survey.arXiv 2507.21045, 2025
Yukang Cao, Jiahao Lu, Zhisheng Huang, Zhuowen Shen, Chengfeng Zhao, Fangzhou Hong, Zhaoxi Chen, Xin Li, Wenping Wang, Yuan Liu, and Ziwei Liu. Reconstructing 4D spatial intelligence: A survey.arXiv 2507.21045, 2025
Pith/arXiv arXiv 2025
-
[57]
Humanoid policy∼human policy
Ri-Zhao Qiu, Shiqi Yang, Xuxin Cheng, Chaitanya Chawla, Jialong Li, Tairan He, Ge Yan, Lars Paulsen, Ge Yang, Sha Yi, Guanya Shi, and Xiaolong Wang. Humanoid policy∼human policy. InCoRL, 2025. 32
2025
-
[58]
ActionNet: A dataset for dexterous bimanual manipulation
Fourier ActionNet Team and Yao Mu. ActionNet: A dataset for dexterous bimanual manipulation. Dataset website, 2025. URLhttps://action-net.org/
2025
-
[59]
DexMV: imitation learning for dexterous manipulation from human videos
Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang, Ruihan Yang, Yang Fu, and Xiaolong Wang. DexMV: imitation learning for dexterous manipulation from human videos. InECCV, 2022
2022
-
[60]
EgoMimic: scaling imitation learning via egocentric video
Simar Kareer, Dhruv Patel, Ryan Punamiya, Pranay Mathur, Shuo Cheng, Chen Wang, Judy Hoffman, and Danfei Xu. EgoMimic: scaling imitation learning via egocentric video. InICRA, 2025
2025
-
[61]
Zhu, Simar Kareer, Judy Hoffman, and Danfei Xu
Ryan Punamiya, Dhruv Patel, Patcharapong Aphiwetsa, Pranav Kuppili, Lawrence Y . Zhu, Simar Kareer, Judy Hoffman, and Danfei Xu. EgoBridge: domain adaptation for generalizable imitation from egocentric human data. InNeurIPS, 2025
2025
-
[62]
UniVLA: learning to act anywhere with task-centric latent actions
Qingwen Bu, Yanting Yang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Yao, Ping Luo, and Hongyang Li. UniVLA: learning to act anywhere with task-centric latent actions. InRSS, 2025
2025
-
[63]
In-N-On: scaling egocentric manipulation with in-the-wild and on-task data
Xiongyi Cai, Ri-Zhao Qiu, Geng Chen, Lai Wei, Isabella Liu, Tianshu Huang, Xuxin Cheng, and Xiaolong Wang. In-N-On: scaling egocentric manipulation with in-the-wild and on-task data. arXiv 2511.15704, 2025
arXiv 2025
-
[64]
Emergence of human to robot transfer in vision-language-action models
Simar Kareer, Karl Pertsch, James Darpinian, Judy Hoffman, Danfei Xu, Sergey Levine, Chelsea Finn, and Suraj Nair. Emergence of human to robot transfer in vision-language-action models. arXiv 2512.22414, 2025
arXiv 2025
-
[65]
Do as I can, not as I say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, et al. Do as I can, not as I say: Grounding language in robotic affordances. InCoRL, 2022
2022
-
[66]
Marcel Torne, Karl Pertsch, Homer Walke, Kyle Vedder, Suraj Nair, Brian Ichter, Allen Z. Ren, Haohuan Wang, Jiaming Tang, Kyle Stachowicz, Karan Dhabalia, Michael Equi, Quan Vuong, Jost Tobias Springenberg, Sergey Levine, Chelsea Finn, and Danny Driess. MEM: multi-scale embodied memory for vision language action models.arXiv 2603.03596, 2026
arXiv 2026
-
[67]
Chenyang Gu, Mingyuan Zhang, Haozhe Xie, Zhongang Cai, Lei Yang, and Ziwei Liu. Bridg- ing semantic and kinematic conditions with diffusion-based discrete motion tokenizer.arXiv 2603.19227, 2026
arXiv 2026
-
[68]
InfiniteDance: scalable 3D dance generation towards in-the-wild general- ization
Ronghui Li, Zhongyuan Hu, Li Siyao, Youliang Zhang, Haozhe Xie, Mingyuan Zhang, Jie Guo, Xiu Li, and Ziwei Liu. InfiniteDance: scalable 3D dance generation towards in-the-wild general- ization. InECCV, 2026
2026
-
[69]
Wong, and Ziwei Liu
Yukang Cao, Liang Pan, Kai Han, Kwan-Yee K. Wong, and Ziwei Liu. AvatarGO: zero-shot 4D human-object interaction generation and animation. InICLR, 2025
2025
-
[70]
Zhao, and Chelsea Finn
Zipeng Fu, Tony Z. Zhao, and Chelsea Finn. Mobile ALOHA: Learning bimanual mobile manip- ulation with low-cost whole-body teleoperation. InCoRL, 2024
2024
-
[71]
Shihan Wu, Xuecheng Liu, Shaoxuan Xie, Pengwei Wang, Xinghang Li, et al. RoboCOIN: an open-sourced bimanual robotic data COllection for INtegrated manipulation.arXiv 2511.17441, 2025
Pith/arXiv arXiv 2025
-
[72]
Chengkai Hou, Kun Wu, Jiaming Liu, Zhengping Che, Di Wu, et al. RoboMIND 2.0: A mul- timodal, bimanual mobile manipulation dataset for generalizable embodied intelligence.arXiv 2512.24653, 2025. 33
arXiv 2025
-
[73]
HSImul3R: physics-in-the-loop reconstruction of simulation-ready human-scene interactions
Yukang Cao, Haozhe Xie, Fangzhou Hong, Long Zhuo, Zhaoxi Chen, Liang Pan, and Ziwei Liu. HSImul3R: physics-in-the-loop reconstruction of simulation-ready human-scene interactions. In ECCV, 2026
2026
-
[74]
WholeBodyVLA: towards unified latent VLA for whole-body loco-manipulation control
Haoran Jiang, Jin Chen, Qingwen Bu, Li Chen, Modi Shi, Yanjie Zhang, Delong Li, Chuanzhe Suo, Chuang Wang, Zhihui Peng, and Hongyang Li. WholeBodyVLA: towards unified latent VLA for whole-body loco-manipulation control. InICLR, 2026
2026
-
[75]
DynamicVLA: A vision-language-action model for dynamic object manipulation.arXiv 2601.22153, 2026
Haozhe Xie, Beichen Wen, Jiarui Zheng, Zhaoxi Chen, Fangzhou Hong, Haiwen Diao, and Ziwei Liu. DynamicVLA: A vision-language-action model for dynamic object manipulation.arXiv 2601.22153, 2026
arXiv 2026
-
[76]
Zhu, Pranav Kuppili, Ryan Punamiya, Patcharapong Aphiwetsa, Dhruv Patel, Simar Kareer, Sehoon Ha, and Danfei Xu
Lawrence Y . Zhu, Pranav Kuppili, Ryan Punamiya, Patcharapong Aphiwetsa, Dhruv Patel, Simar Kareer, Sehoon Ha, and Danfei Xu. EMMA: scaling mobile manipulation via egocentric human data.IEEE RA-L, 11(3):3087–3094, 2025
2025
-
[77]
HoMMI: learning whole-body mobile manipulation from human demonstra- tions.arXiv 2603.03243, 2026
Xiaomeng Xu, Jisang Park, Han Zhang, Eric Cousineau, Aditya Bhat, Jose Barreiros, Dian Wang, and Shuran Song. HoMMI: learning whole-body mobile manipulation from human demonstra- tions.arXiv 2603.03243, 2026
Pith/arXiv arXiv 2026
-
[78]
2D Gaussian splat- ting for geometrically accurate radiance fields
Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2D Gaussian splat- ting for geometrically accurate radiance fields. InSIGGRAPH, 2024
2024
-
[79]
Garrido-Jurado, R
S. Garrido-Jurado, R. Mu ˜noz-Salinas, F.J. Madrid-Cuevas, and M.J. Mar´ın-Jim´enez. Automatic generation and detection of highly reliable fiducial markers under occlusion.PR, 47(6):2280– 2292, 2014
2014
-
[80]
Tsai and Reimar Lenz
Roger Y . Tsai and Reimar Lenz. A new technique for fully autonomous and efficient 3D robotics hand/eye calibration.IEEE T-RA, 5(3):345–358, 1989
1989
This paper was first reviewed by grok-4.5 on July 31, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.