Pith. sign in

REVIEW 3 major objections 6 minor 116 references

Everyday household skill can be recorded as one time-aligned stream of vision, body, hands, objects, sound, and touch in real homes.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

ACE-Data-0 is a 150-hour home HOI dataset with millisecond-synced ego/exo video, mocap body/hands, object 6-DoF, audio, and tactile signals, plus a three-level benchmark exposing large SOTA gaps.

T0 review reviewed 2026-07-31 challenge →

load-bearing objection Solid dual-scale home capture with true multi-modal GT and honest SOTA gaps; limitations are real but already owned, and the paper deserves referees. the 3 major comments →

arxiv 2607.28625 v1 pith:UJK7YLOX submitted 2026-07-30 cs.CV

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

classification cs.CV
keywords embodied AIhuman-object interactionegocentric visionmultimodal capturemotion capturetactile sensinglong-horizon household activityvision-language-action
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Embodied AI is stuck because the data it needs—how first-person sight, whole-body motion, dexterous hands, object state, sound, and contact evolve together while people chase household goals—has never been recorded as one registered stream. Existing datasets split that loop across viewpoints, modalities, or lab-scale setups, so models never see the full perception–action cycle. This paper’s answer is the Ambient Capture Engine: turn real homes into calibrated studios at two scales (table for fine hand–object work; room for locomotion and whole-body activity) and capture egocentric plus multi-view video, full-body and hand motion, object 6-DoF trajectories, audio, and tactile pressure on a shared clock and world frame. From that engine comes ACE-Data-0—150 hours, 17M frames, 75,000 episodes, 200 task categories, 50 people, goal-level instructions rather than scripted steps—plus a hierarchical benchmark from tactile-from-vision up through body pose to hand–object motion. State-of-the-art methods fail under contact, occlusion, egomotion, and long horizons, which is exactly why the authors argue this synchronized supervision is a foundation for imitation, world models, and vision-language-action systems.

Core claim

The full perception–action loop of everyday home interaction can be captured holistically: dual-scale ambient instrumentation in real furnished homes yields a single metrically registered, millisecond-aligned multisensory stream (ego and multi-exo video, body and hand motion, object geometry and 6-DoF, audio, tactile), and on the resulting long-horizon ACE-Data-0 benchmark current methods show large, systematic gaps under contact, occlusion, egomotion, and extended time.

What carries the argument

Ambient Capture Engine (ACE): two complementary home configurations (table-scale dense close-range capture; room-scale full-apartment coverage) that share optical-clock temporal sync to the motion-capture timeline and marker-bridged spatial calibration into one world frame, so every modality describes the same physical moment.

Load-bearing premise

Two instrumented homes, a few dozen pre-scanned markered objects, and people wearing suits, gloves, headsets, and markers still produce behavior and visuals representative enough to train general embodied agents.

What would settle it

Train imitation or VLA policies on ACE-Data-0 and test zero-shot or few-shot transfer in an uninstrumented third home with novel layouts, lighting, and unmarked everyday objects; if performance collapses relative to in-distribution homes, the representativeness claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Imitation and policy learning can supervise contact, kinematics, and multi-view vision from the same physical instant rather than stitching mismatched sources.
  • World models and VLA systems gain long-horizon household chains with natural subtask order, hesitation, and recovery under goal-level instructions.
  • Benchmarks can diagnose failure mode (missed contact vs bad object state vs trajectory drift) instead of only final task success.
  • Egocentric and exocentric hand/body estimators can be compared and fused under identical measured ground truth.
  • Tactile-from-vision becomes a learnable mapping because every frame is paired with glove pressure on a shared clock.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If suit and marker appearance dominate learned features, domain randomization or appearance stripping may be required before robot transfer works.
  • Extending ground truth beyond rigid marked objects to fluids, cloth, and articulated appliances is the natural next bottleneck the dataset itself surfaces.
  • Providing measured headset motion as an oracle input would cleanly separate hand-reconstruction error from egomotion error in future ego baselines.
  • Scaling sites and unmarked objects may matter more than scaling hours inside the same two layouts for open-world generalization.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The manuscript introduces the Ambient Capture Engine (ACE), a dual-scale (table-scale and room-scale) capture system that turns furnished homes into calibrated, time-synchronized multi-sensor studios, and releases ACE-Data-0: ~150 hours / 17M frames / 75,000 episodes across 200 task categories from 50 participants in 2 sites. Streams include egocentric and multi-view exocentric video, full-body and hand motion, object meshes and 6-DoF trajectories, audio, and tactile pressure, registered to a common OptiTrack-referenced timeline and world frame. Tasks are goal-level (atomic HOI, long-horizon HOI chains, HSI) rather than scripted atomic actions. A hierarchical benchmark evaluates tactile-from-vision, human motion recovery, and hand motion from ego/exo views; zero-shot tests of 30+ SOTA methods report large gaps under contact, occlusion, egomotion, and long horizons. The authors position the corpus as supervision for imitation learning, world models, and VLA systems.

Significance. If the capture fidelity, scale, and released annotations hold as described, this is a substantial systems and resource contribution for embodied AI. Prior HOI and egocentric datasets typically fragment viewpoint, modality, or spatial scale; ACE’s explicit dual-scale design plus measured (not purely estimated) body/hand/object/tactile alignment in real homes is a clear advance over lab-only mocap HOI and uninstrumented egocentric video. Strengths include a concrete synchronization and calibration pipeline (§4.1: OptiTrack reference, QR optical clock, hand-eye ego calibration, reported ms residuals and <3 px median reprojection), goal-level long-horizon collection (§4.2), rich measured annotations (§4.3), and unusually broad zero-shot tables (Tables 3–6). For a dataset/benchmark paper, that package is significant even without new learning algorithms.

major comments (3)
  1. [§5; Abstract; §1; §6] The abstract, introduction, and conclusion repeatedly claim ACE-Data-0 as a “scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI,” yet §5 only reports zero-shot evaluation of external pretrained checkpoints. There is no fine-tuning, imitation, world-model, or VLA experiment that uses ACE-Data-0 as training supervision. For a resource paper this is not fatal, but the training-utility claim is load-bearing in the framing: either add a minimal transfer/finetune study (e.g., hand or body estimator fine-tuned on a train split, or a small behavior-cloning baseline) or substantially tone down claims that currently outrun the evidence in §5.
  2. [§3.2.2; §5.3.1; Tables 5–6] Hand ground truth is heterogeneous across scales: Manus gloves at room scale versus RANSAC triangulation of 2D keypoints from eight exo cameras plus manual refinement at table scale (§3.2.2, §5.3.1), with joints masked when under-observed. Tables 5–6 and the cross-view analysis (§5.3.3) treat these as a single comparable GT. The paper should quantify residual error of the table-scale pipeline (e.g., against a glove subset or multi-view consistency), state which benchmark splits use which GT source, and confirm that masked joints do not selectively remove the hardest contact frames—otherwise ego/exo and method rankings may partly reflect GT quality rather than viewpoint difficulty.
  3. [Abstract; §1; Table 1; Limitations (§6)] Domain breadth is a real limit on the “real home / general embodied” claim: only two sites, ~50 pre-scanned markered rigid objects, and participants in mocap suits, gloves, and headsets (Limitations; §3.2; §4.2). The Limitations section already states this, but the main text and Table 1 still present ACE-Data-0 as closing the home-scene gap relative to in-the-wild ego corpora. Please make the scope explicit earlier (abstract/intro/Table 1 caption): instrumented rigid HOI/HSI in two instrumented homes, not open-world domestic diversity—and avoid implying coverage of fluids, articulations, or deformables that are explicitly unannotated.
minor comments (6)
  1. [Table 1] Table 1 is very dense; a short caption note defining Sync/LH and the checkmark variants (measured vs estimated vs partial) would help readers compare rows without hunting the legend paragraph.
  2. [§4.2.3; Fig. 8] Dataset statistics: abstract and §4.2.3 say 17M frames / 150 hours / 75,000 episodes; Fig. 8 uses “15M+” in one place. Please reconcile all headline numbers.
  3. [§5.1.1; Table 3] §5.1.1 metrics (Temp Acc., C-IoU, V-IoU, CoP) need precise definitions or a short appendix (thresholds, min-max aggregation, palm/fingertip mask). Reproducibility of Table 3 depends on them.
  4. [§4.3.1–4.3.2] Textual annotations rely on Gemini-3.1-pro-preview with human correction (§4.3.2). State approximate correction rate and whether language labels are used in any benchmark track (they appear unused in §5).
  5. [Project page / §6] Release plan: project and HF links are given; please state what will be public at acceptance (raw vs processed streams, meshes, sync tables, train/test splits, licenses) so the “foundation” claim is checkable.
  6. [Table 4; References] Minor polish: “W A-MPJPE” spacing in Table 4; a few repeated figure callouts; ensure all baseline citations in Tables 3–6 match the reference list versions used.

Circularity Check

0 steps flagged

No significant circularity: dataset/capture paper with measured GT and external pretrained baselines.

full rationale

ACE-Data-0 is a systems, capture, and benchmark paper, not a first-principles derivation that recovers fitted constants as predictions. Load-bearing claims are engineering and empirical: dual-scale ambient capture in real homes, hardware/software sync and calibration to a common spatio-temporal frame, goal-level long-horizon collection, and a hierarchical benchmark on which released pretrained SOTA methods fail under contact, occlusion, egomotion, and long horizons. Ground truth is measured (OptiTrack markers, Manus gloves, tactile pressure, scanned/2DGS object meshes), not inferred from the models under test; annotations such as pose reprojections and boxes follow by projection from tracked states. Benchmark tables compare external checkpoints to held-out mocap/tactile GT. Text captions use an external VLM plus human correction—ordinary labeling, not a circular proof of multi-modal alignment. Author self-citations appear in related work and future-application context and are not load-bearing uniqueness or ansatz imports for the central dataset claim. No self-definitional loop, fitted-input-as-prediction, or renaming of a known result as a forced derivation was found.

Axiom & Free-Parameter Ledger

3 free parameters · 6 axioms · 3 invented entities

Load-bearing commitments are engineering and sampling assumptions of the capture recipe, not free parameters in a fitted theory. The central scientific claim (unified ambient capture closes a data gap and stresses SOTA) rests on trusting mocap/tactile as GT, two-site home ecology as ‘real home’ distribution, goal-level prompting as natural behavior, and zero-shot pretrained evaluation as diagnostic of field readiness.

free parameters (3)
  • Sync residual / timing model (offset + slow drift line fit to QR clock readings) = ms-level residuals; ego mutual misalignment <2 ms; calibration re-estimated offsets <4 ms
    Per-take linear time mapping from camera timestamps to OptiTrack; reported ms-level residuals. Operational calibration choice that defines ‘synchronized’ labels.
  • Held-out test split size = 10 hours
    Benchmark uses 10 hours held out; selection criteria beyond ‘test set’ are lightly specified and affect all reported gaps.
  • Contact/pressure metric thresholds (C-IoU, V-IoU aggregation, CoP region definition)
    Tactile evaluation depends on thresholding and aggregation choices that are conventional but not uniquely determined.
axioms (6)
  • domain assumption Optical mocap marker trajectories plus glove sensors constitute metric ground truth for body, hand, object pose, and contact pressure in evaluation.
    Entire benchmark hierarchy compares vision methods to these streams (§5); soft-tissue, marker error, and glove calibration limits are not quantified as uncertainty.
  • domain assumption Two furnished homes with dense truss-mounted sensors remain ecologically valid ‘real home’ interaction distributions despite instrumentation.
    Stated design principle §3.2.1 and Limitations (two sites; suit/gloves/markers visible).
  • domain assumption Goal-level verbal instructions (not step scripts) induce natural long-horizon household behavior suitable for imitation/VLA supervision.
    Task design §4.2.1; core differentiator vs atomic HOI datasets.
  • domain assumption Zero-shot evaluation of officially released pretrained checkpoints fairly exposes method gaps on home long-horizon data.
    Explicit protocol §5; does not test in-domain finetuning ceilings.
  • standard math Standard multi-view geometry, hand-eye calibration, bundle adjustment, and SMPL-X/MANO representations are adequate to register all streams into one spatio-temporal frame.
    Calibration §4.1.2 cites classical hand-eye and BA; errors reported as median reprojection px.
  • ad hoc to paper Gemini-generated then human-corrected language descriptions are acceptable semantic annotations aligned to measured activity.
    Annotation pipeline §4.3.2; only non-measured annotation family.
invented entities (3)
  • Ambient Capture Engine (ACE) dual-scale studio paradigm independent evidence
    purpose: Name the hardware+sync+calibration recipe that turns homes into multi-modal recording studios at table and room scale.
    Systems contribution; independently checkable via released data and described hardware, not a latent physical entity.
  • ACE-Data-0 corpus and hierarchical signal→component→interaction benchmark independent evidence
    purpose: Provide the concrete dataset and evaluation ladder claimed to ground embodied learning.
    Primary artifact; evidence is the release and reported baselines, external to any fitted theory.
  • ACE-Ego-Head-V02 Lite / ACE-Sense-Glove Lite no independent evidence
    purpose: Custom wearable ego and tactile sensing in the suite.
    Device names specific to the project; reproducibility depends on whether specs/data suffice without identical hardware.

reviewed 2026-07-31 · how reviews work

0 comments
Cite this review

Pith. "Pith review of ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine." pith.science (2026). https://pith.science/paper/UJK7YLOX

@misc{pith2026260728625,
  author       = {Pith},
  title        = {Pith review of: ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UJK7YLOX}},
  note         = {Machine review of arXiv:2607.28625}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

116 extracted references · 9 linked inside Pith

  1. [1]

    Psychology press, 2014

    James J Gibson.The ecological approach to visual perception: classic edition. Psychology press, 2014

  2. [2]

    MIT press, 2006

    Rolf Pfeifer and Josh Bongard.How the body shapes the way we think: a new view of intelligence. MIT press, 2006

  3. [3]

    ImageNet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. InCVPR, 2009

  4. [4]

    LAION-5B: an open large-scale dataset for training next generation image-text models

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. LAION-5B: an open large-scale dataset for training next generation image-text models. InNeurIPS, 2022

  5. [5]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InICML, 2021

  6. [6]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhari- wal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Lit...

  7. [7]

    Ego-Exo4D: understanding skilled human activity from first- and third-person perspectives

    Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyl- los Afouras, et al. Ego-Exo4D: understanding skilled human activity from first- and third-person perspectives. InCVPR, 2024

  8. [8]

    Rescal- ing egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100.IJCV, 130 (1):33–55, 2022

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Rescal- ing egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100.IJCV, 130 (1):33–55, 2022

  9. [9]

    Xperience-10M: A large-scale egocentric multimodal dataset with structured 3D/4D annotations

    Ropedia. Xperience-10M: A large-scale egocentric multimodal dataset with structured 3D/4D annotations. Hugging Face dataset, 2026. URLhttps://huggingface.co/datasets/ ropedia-ai/xperience-10m

  10. [10]

    Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll

    Bharat Lal Bhatnagar, Xianghui Xie, Ilya A. Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. BEHA VE: dataset and method for tracking human object interactions. In CVPR, 2022

  11. [11]

    Black, and Dimitrios Tzionas

    Omid Taheri, Nima Ghorbani, Michael J. Black, and Dimitrios Tzionas. GRAB: A dataset of whole-body human grasping of objects. InECCV, 2020

  12. [12]

    Black, and Otmar Hilliges

    Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, and Otmar Hilliges. ARCTIC: A dataset for dexterous bimanual hand-object manipulation. InCVPR, 2023

  13. [13]

    OAKINK2: A dataset of bimanual hands-object manipulation in complex task completion

    Xinyu Zhan, Lixin Yang, Yifei Zhao, Kangrui Mao, Hanlin Xu, Zenan Lin, Kailin Li, and Cewu Lu. OAKINK2: A dataset of bimanual hands-object manipulation in complex task completion. InCVPR, 2024. 29

  14. [14]

    HOT3D: hand and object tracking in 3D from egocentric multi-view videos

    Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Shangchen Han, Fan Zhang, Linguang Zhang, Jade Fountain, Edward Miller, Selen Basol, Richard Newcombe, Robert Wang, Jakob Julian Engel, and Tomas Hodan. HOT3D: hand and object tracking in 3D from egocentric multi-view videos. InCVPR, 2025

  15. [15]

    Twigg, Charles C

    Samarth Brahmbhatt, Chengcheng Tang, Christopher D. Twigg, Charles C. Kemp, and James Hays. ContactPose: A dataset of grasps with object contact and hand pose. InECCV, 2020

  16. [16]

    Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Dieter Fox

    Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S. Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Dieter Fox. DexYCB: A benchmark for capturing hand grasping of objects. InCVPR, 2021

  17. [17]

    GigaHands: A massive annotated dataset of bimanual hand activities

    Rao Fu, Dingxi Zhang, Alex Jiang, Wanjia Fu, Austin Funk, Daniel Ritchie, and Srinath Sridhar. GigaHands: A massive annotated dataset of bimanual hand activities. InCVPR, 2025

  18. [18]

    H2O: two hands manipulating objects for first person interaction recognition

    Taein Kwon, Bugra Tekin, Jan Stuhmer, Federica Bogo, and Marc Pollefeys. H2O: two hands manipulating objects for first person interaction recognition. InICCV, 2021

  19. [19]

    World action models are zero-shot policies

    Seonghyeon Ye, Yunhao Ge, Kaiyuan Zheng, Shenyuan Gao, Sihyun Yu, George Kurian, Suneel Indupuru, You Liang Tan, Chuning Zhu, Jiannan Xiang, Ayaan Malik, Kyungmin Lee, William Liang, Nadun Ranawaka, Jiasheng Gu, Yinzhen Xu, Guanzhi Wang, Fengyuan Hu, Avnish Narayan, Johan Bjorck, Jing Wang, Gwanghyun Kim, Dantong Niu, Ruijie Zheng, Yuqi Xie, Jimmy Wu, Qi ...

  20. [20]

    EgoVLA: learning vision-language-action models from egocentric human videos

    Ruihan Yang, Qinxi Yu, Yecheng Wu, Rui Yan, Borui Li, An-Chieh Cheng, Xueyan Zou, Yunhao Fang, Xuxin Cheng, Ri-Zhao Qiu, Hongxu Yin, Sifei Liu, Song Han, Yao Lu, and Xiaolong Wang. EgoVLA: learning vision-language-action models from egocentric human videos. InCoRL, 2025

  21. [21]

    Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Robert Equi, Chelsea Finn, Niccolo Fusai, Manuel Y . Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, brian ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Z. Ren,...

  22. [22]

    World Guidance: world modeling in condi- tion space for action generation.arXiv 2602.22010, 2026

    Yue Su, Sijin Chen, Haixin Shi, Mingyu Liu, Zhengshen Zhang, Ningyuan Huang, Weiheng Zhong, Zhengbang Zhu, Yuxiao Liu, and Xihui Liu. World Guidance: world modeling in condi- tion space for action generation.arXiv 2602.22010, 2026

  23. [23]

    Ego4D: around the world in 3,000 hours of egocentric video

    Kristen Grauman, Andrew Westbury, Eugene Byrne, et al. Ego4D: around the world in 3,000 hours of egocentric video. InCVPR, 2022

  24. [24]

    HoloAssist: an egocentric human interaction dataset for interactive AI assistants in the real world

    Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bugra Tekin, Felipe Vieira Frujeri, Neel Joshi, and Marc Pollefeys. HoloAssist: an egocentric human interaction dataset for interactive AI assistants in the real world. InICCV, 2023

  25. [25]

    EgoLife: towards egocentric life assistant

    Jingkang Yang, Shuai Liu, Hongming Guo, Yuhao Dong, Xiamengwei Zhang, Sicheng Zhang, Pengyun Wang, Zitang Zhou, Binzhu Xie, Ziyue Wang, Bei Ouyang, Zhengyu Lin, Marco Cominelli, Zhongang Cai, Bo Li, Yuanhan Zhang, Peiyuan Zhang, Fangzhou Hong, Joerg Wid- mer, Francesco Gringoli, Lei Yang, and Ziwei Liu. EgoLife: towards egocentric life assistant. In CVPR,...

  26. [26]

    HD-EPIC: A highly-detailed egocentric video dataset

    Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Parida, Kait- ing Liu, Prajwal Gatti, Siddhant Bansal, Kevin Flanagan, Jacob Chalk, Zhifan Zhu, Rhodri Guer- rier, Fahd Abdelazim, Bin Zhu, Davide Moltisanti, Michael Wray, Hazel Doughty, and Dima Damen. HD-EPIC: A highly-detailed egocentric video dataset. InCVPR, 2025

  27. [27]

    HOnnotate: A method for 3D annotation of hand and object poses

    Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. HOnnotate: A method for 3D annotation of hand and object poses. InCVPR, 2020

  28. [28]

    OakInk: A large-scale knowledge repository for understanding hand-object interaction

    Lixin Yang, Kailin Li, Xinyu Zhan, Fei Wu, Anran Xu, Liu Liu, and Cewu Lu. OakInk: A large-scale knowledge repository for understanding hand-object interaction. InCVPR, 2022

  29. [29]

    HOI4D: A 4D egocentric dataset for category-level human-object interaction

    Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4D egocentric dataset for category-level human-object interaction. InCVPR, 2022

  30. [30]

    TACO: benchmarking generalizable bimanual Tool-ACtion-Object understanding

    Yun Liu, Haolin Yang, Xu Si, Ling Liu, Zipeng Li, Yuxiang Zhang, Yebin Liu, and Li Yi. TACO: benchmarking generalizable bimanual Tool-ACtion-Object understanding. InCVPR, 2024

  31. [31]

    Black, and Dimitrios Tzionas

    Yinghao Huang, Omid Taheri, Michael J. Black, and Dimitrios Tzionas. InterCap: joint marker- less 3D tracking of humans and objects in interaction.IJCV, 132(7):2551–2566, 2024

  32. [32]

    Full-body articulated human-object interaction

    Nan Jiang, Tengyu Liu, Zhexuan Cao, Jieming Cui, Zhiyuan Zhang, Yixin Chen, He Wang, Yixin Zhu, and Siyuan Huang. Full-body articulated human-object interaction. InICCV, 2023

  33. [33]

    EgoBody: human body shape and motion of interacting people from head- mounted devices

    Siwei Zhang, Qianli Ma, Yan Zhang, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, and Siyu Tang. EgoBody: human body shape and motion of interacting people from head- mounted devices. InECCV, 2022

  34. [34]

    Parkhi, Richard A

    Xiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters, Thomas Whelan, Chen Kong, Omkar M. Parkhi, Richard A. Newcombe, and Carl Yuheng Ren. Aria Digital Twin: A new benchmark dataset for egocentric 3D machine perception. InICCV, 2023

  35. [35]

    Karen Liu

    Jiaman Li, Jiajun Wu, and C. Karen Liu. Object motion guided human motion synthesis.ACM TOG, 42(6):197:1–197:11, 2023

  36. [36]

    HIMO: A new benchmark for full-body human interacting with multiple objects

    Xintao Lv, Liang Xu, Yichao Yan, Xin Jin, Congsheng Xu, Shuwen Wu, Yifan Liu, Lincheng Li, Mengxiao Bi, Wenjun Zeng, and Xiaokang Yang. HIMO: A new benchmark for full-body human interacting with multiple objects. InECCV, 2024

  37. [37]

    Scaling up dynamic human-scene interaction modeling

    Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, and Siyuan Huang. Scaling up dynamic human-scene interaction modeling. InCVPR, 2024

  38. [38]

    ParaHome: parameterizing ev- eryday home activities towards 3D generative modeling of human-object interactions

    Jeonghwan Kim, Jisoo Kim, Jeonghyeon Na, and Hanbyul Joo. ParaHome: parameterizing ev- eryday home activities towards 3D generative modeling of human-object interactions. InCVPR, 2024

  39. [39]

    Karen Liu, Ziwei Liu, Jakob J

    Lingni Ma, Yuting Ye, Fangzhou Hong, Vladimir Guzov, Yifeng Jiang, Rowan Postyeni, Luis Pesqueira, Alexander Gamino, Vijay Baiyya, Hyo Jin Kim, Kevin Bailey, David Soriano Fosas, C. Karen Liu, Ziwei Liu, Jakob J. Engel, Renzo De Nardi, and Richard A. Newcombe. Nymeria: A massive collection of multimodal egocentric daily motion in the wild. InECCV, 2024

  40. [40]

    HUMOTO: A 4D dataset of mocap human object interactions

    Jiaxin Lu, Chun-Hao Paul Huang, Uttaran Bhattacharya, Qixing Huang, and Yi Zhou. HUMOTO: A 4D dataset of mocap human object interactions. InICCV, 2025

  41. [41]

    Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen- Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, Abraham Lee, Kuan Fang, Chelsea Finn, and Sergey Levine

    Homer Rich Walke, Kevin Black, Tony Z. Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen- Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, Abraham Lee, Kuan Fang, Chelsea Finn, and Sergey Levine. BridgeData V2: A dataset for robot learning at scale. InCoRL, 2023. 31

  42. [42]

    RH20T: A comprehensive robotic dataset for learning diverse skills in one-shot

    Haoshu Fang, Hongjie Fang, Zhenyu Tang, Jirong Liu, Chenxi Wang, Junbo Wang, Haoyi Zhu, and Cewu Lu. RH20T: A comprehensive robotic dataset for learning diverse skills in one-shot. InICRA, 2024

  43. [43]

    Karen Liu

    Chen Wang, Haochen Shi, Weizhuo Wang, Ruohan Zhang, Li Fei-Fei, and C. Karen Liu. DexCap: scalable and portable mocap data collection system for dexterous manipulation. InRSS, 2024

  44. [44]

    DROID: A large-scale in-the-wild robot manipulation dataset

    Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, et al. DROID: A large-scale in-the-wild robot manipulation dataset. InRSS, 2024

  45. [45]

    Ag- iBot World Colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems

    AgiBot-World-Contributors, Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, et al. Ag- iBot World Colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems. InRSS, 2025

  46. [46]

    Yoon, Mouli Sivapurapu, and Jian Zhang

    Ryan Hoque, Peide Huang, David J. Yoon, Mouli Sivapurapu, and Jian Zhang. EgoDex: learning dexterous manipulation from large-scale egocentric video. InICLR, 2026

  47. [47]

    Galaxea open-world dataset and G0 dual-system VLA model.arXiv 2509.00576, 2025

    Tao Jiang, Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Jianning Cui, Xiao Liu, Shuiqi Cheng, Jiyang Gao, Huazhe Xu, and Hang Zhao. Galaxea open-world dataset and G0 dual-system VLA model.arXiv 2509.00576, 2025

  48. [48]

    EgoScale: scaling dexterous manipulation with diverse egocentric human data.arXiv 2602.16710, 2026

    Ruijie Zheng, Dantong Niu, Yuqi Xie, Jing Wang, Mengda Xu, Yunfan Jiang, Fernando Casta˜neda, Fengyuan Hu, You Liang Tan, Letian Fu, Trevor Darrell, Furong Huang, Yuke Zhu, Danfei Xu, and Linxi Fan. EgoScale: scaling dexterous manipulation with diverse egocentric human data.arXiv 2602.16710, 2026

  49. [49]

    Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and cap- turing hands and bodies together.ACM TOG, 36(6):245:1–245:17, 2017

  50. [50]

    Keypoint Transformer: solving joint identification in challenging hands and object interactions for accurate 3D pose estimation

    Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Keypoint Transformer: solving joint identification in challenging hands and object interactions for accurate 3D pose estimation. InCVPR, 2022

  51. [51]

    Reconstructing hands in 3D with transformers

    Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Ji- tendra Malik. Reconstructing hands in 3D with transformers. InCVPR, 2024

  52. [52]

    HaWoR: world- space hand motion reconstruction from egocentric videos

    Jinglei Zhang, Jiankang Deng, Chao Ma, and Rolandos Alexandros Potamias. HaWoR: world- space hand motion reconstruction from egocentric videos. InCVPR, 2025

  53. [53]

    Yufei Ye, Yao Feng, Omid Taheri, Haiwen Feng, Shubham Tulsiani, and Michael J. Black. Pre- dicting 4D hand trajectory from monocular videos. In3DV, 2026

  54. [54]

    HORT: monoc- ular hand-held objects reconstruction with transformers

    Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, and Cordelia Schmid. HORT: monoc- ular hand-held objects reconstruction with transformers. InICCV, 2025

  55. [55]

    PhySIC: physically plausible 3D human-scene interaction and contact from a single image

    Pradyumna Yalandur Muralidhar, Yuxuan Xue, Xianghui Xie, Margaret Kostyrko, and Gerard Pons-Moll. PhySIC: physically plausible 3D human-scene interaction and contact from a single image. InSIGGRAPH Asia, 2025

  56. [56]

    Reconstructing 4D spatial intelligence: A survey.arXiv 2507.21045, 2025

    Yukang Cao, Jiahao Lu, Zhisheng Huang, Zhuowen Shen, Chengfeng Zhao, Fangzhou Hong, Zhaoxi Chen, Xin Li, Wenping Wang, Yuan Liu, and Ziwei Liu. Reconstructing 4D spatial intelligence: A survey.arXiv 2507.21045, 2025

  57. [57]

    Humanoid policy∼human policy

    Ri-Zhao Qiu, Shiqi Yang, Xuxin Cheng, Chaitanya Chawla, Jialong Li, Tairan He, Ge Yan, Lars Paulsen, Ge Yang, Sha Yi, Guanya Shi, and Xiaolong Wang. Humanoid policy∼human policy. InCoRL, 2025. 32

  58. [58]

    ActionNet: A dataset for dexterous bimanual manipulation

    Fourier ActionNet Team and Yao Mu. ActionNet: A dataset for dexterous bimanual manipulation. Dataset website, 2025. URLhttps://action-net.org/

  59. [59]

    DexMV: imitation learning for dexterous manipulation from human videos

    Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang, Ruihan Yang, Yang Fu, and Xiaolong Wang. DexMV: imitation learning for dexterous manipulation from human videos. InECCV, 2022

  60. [60]

    EgoMimic: scaling imitation learning via egocentric video

    Simar Kareer, Dhruv Patel, Ryan Punamiya, Pranay Mathur, Shuo Cheng, Chen Wang, Judy Hoffman, and Danfei Xu. EgoMimic: scaling imitation learning via egocentric video. InICRA, 2025

  61. [61]

    Zhu, Simar Kareer, Judy Hoffman, and Danfei Xu

    Ryan Punamiya, Dhruv Patel, Patcharapong Aphiwetsa, Pranav Kuppili, Lawrence Y . Zhu, Simar Kareer, Judy Hoffman, and Danfei Xu. EgoBridge: domain adaptation for generalizable imitation from egocentric human data. InNeurIPS, 2025

  62. [62]

    UniVLA: learning to act anywhere with task-centric latent actions

    Qingwen Bu, Yanting Yang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Yao, Ping Luo, and Hongyang Li. UniVLA: learning to act anywhere with task-centric latent actions. InRSS, 2025

  63. [63]

    In-N-On: scaling egocentric manipulation with in-the-wild and on-task data

    Xiongyi Cai, Ri-Zhao Qiu, Geng Chen, Lai Wei, Isabella Liu, Tianshu Huang, Xuxin Cheng, and Xiaolong Wang. In-N-On: scaling egocentric manipulation with in-the-wild and on-task data. arXiv 2511.15704, 2025

  64. [64]

    Emergence of human to robot transfer in vision-language-action models

    Simar Kareer, Karl Pertsch, James Darpinian, Judy Hoffman, Danfei Xu, Sergey Levine, Chelsea Finn, and Suraj Nair. Emergence of human to robot transfer in vision-language-action models. arXiv 2512.22414, 2025

  65. [65]

    Do as I can, not as I say: Grounding language in robotic affordances

    Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, et al. Do as I can, not as I say: Grounding language in robotic affordances. InCoRL, 2022

  66. [66]

    Ren, Haohuan Wang, Jiaming Tang, Kyle Stachowicz, Karan Dhabalia, Michael Equi, Quan Vuong, Jost Tobias Springenberg, Sergey Levine, Chelsea Finn, and Danny Driess

    Marcel Torne, Karl Pertsch, Homer Walke, Kyle Vedder, Suraj Nair, Brian Ichter, Allen Z. Ren, Haohuan Wang, Jiaming Tang, Kyle Stachowicz, Karan Dhabalia, Michael Equi, Quan Vuong, Jost Tobias Springenberg, Sergey Levine, Chelsea Finn, and Danny Driess. MEM: multi-scale embodied memory for vision language action models.arXiv 2603.03596, 2026

  67. [67]

    Bridg- ing semantic and kinematic conditions with diffusion-based discrete motion tokenizer.arXiv 2603.19227, 2026

    Chenyang Gu, Mingyuan Zhang, Haozhe Xie, Zhongang Cai, Lei Yang, and Ziwei Liu. Bridg- ing semantic and kinematic conditions with diffusion-based discrete motion tokenizer.arXiv 2603.19227, 2026

  68. [68]

    InfiniteDance: scalable 3D dance generation towards in-the-wild general- ization

    Ronghui Li, Zhongyuan Hu, Li Siyao, Youliang Zhang, Haozhe Xie, Mingyuan Zhang, Jie Guo, Xiu Li, and Ziwei Liu. InfiniteDance: scalable 3D dance generation towards in-the-wild general- ization. InECCV, 2026

  69. [69]

    Wong, and Ziwei Liu

    Yukang Cao, Liang Pan, Kai Han, Kwan-Yee K. Wong, and Ziwei Liu. AvatarGO: zero-shot 4D human-object interaction generation and animation. InICLR, 2025

  70. [70]

    Zhao, and Chelsea Finn

    Zipeng Fu, Tony Z. Zhao, and Chelsea Finn. Mobile ALOHA: Learning bimanual mobile manip- ulation with low-cost whole-body teleoperation. InCoRL, 2024

  71. [71]

    RoboCOIN: an open-sourced bimanual robotic data COllection for INtegrated manipulation.arXiv 2511.17441, 2025

    Shihan Wu, Xuecheng Liu, Shaoxuan Xie, Pengwei Wang, Xinghang Li, et al. RoboCOIN: an open-sourced bimanual robotic data COllection for INtegrated manipulation.arXiv 2511.17441, 2025

  72. [72]

    RoboMIND 2.0: A mul- timodal, bimanual mobile manipulation dataset for generalizable embodied intelligence.arXiv 2512.24653, 2025

    Chengkai Hou, Kun Wu, Jiaming Liu, Zhengping Che, Di Wu, et al. RoboMIND 2.0: A mul- timodal, bimanual mobile manipulation dataset for generalizable embodied intelligence.arXiv 2512.24653, 2025. 33

  73. [73]

    HSImul3R: physics-in-the-loop reconstruction of simulation-ready human-scene interactions

    Yukang Cao, Haozhe Xie, Fangzhou Hong, Long Zhuo, Zhaoxi Chen, Liang Pan, and Ziwei Liu. HSImul3R: physics-in-the-loop reconstruction of simulation-ready human-scene interactions. In ECCV, 2026

  74. [74]

    WholeBodyVLA: towards unified latent VLA for whole-body loco-manipulation control

    Haoran Jiang, Jin Chen, Qingwen Bu, Li Chen, Modi Shi, Yanjie Zhang, Delong Li, Chuanzhe Suo, Chuang Wang, Zhihui Peng, and Hongyang Li. WholeBodyVLA: towards unified latent VLA for whole-body loco-manipulation control. InICLR, 2026

  75. [75]

    DynamicVLA: A vision-language-action model for dynamic object manipulation.arXiv 2601.22153, 2026

    Haozhe Xie, Beichen Wen, Jiarui Zheng, Zhaoxi Chen, Fangzhou Hong, Haiwen Diao, and Ziwei Liu. DynamicVLA: A vision-language-action model for dynamic object manipulation.arXiv 2601.22153, 2026

  76. [76]

    Zhu, Pranav Kuppili, Ryan Punamiya, Patcharapong Aphiwetsa, Dhruv Patel, Simar Kareer, Sehoon Ha, and Danfei Xu

    Lawrence Y . Zhu, Pranav Kuppili, Ryan Punamiya, Patcharapong Aphiwetsa, Dhruv Patel, Simar Kareer, Sehoon Ha, and Danfei Xu. EMMA: scaling mobile manipulation via egocentric human data.IEEE RA-L, 11(3):3087–3094, 2025

  77. [77]

    HoMMI: learning whole-body mobile manipulation from human demonstra- tions.arXiv 2603.03243, 2026

    Xiaomeng Xu, Jisang Park, Han Zhang, Eric Cousineau, Aditya Bhat, Jose Barreiros, Dian Wang, and Shuran Song. HoMMI: learning whole-body mobile manipulation from human demonstra- tions.arXiv 2603.03243, 2026

  78. [78]

    2D Gaussian splat- ting for geometrically accurate radiance fields

    Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2D Gaussian splat- ting for geometrically accurate radiance fields. InSIGGRAPH, 2024

  79. [79]

    Garrido-Jurado, R

    S. Garrido-Jurado, R. Mu ˜noz-Salinas, F.J. Madrid-Cuevas, and M.J. Mar´ın-Jim´enez. Automatic generation and detection of highly reliable fiducial markers under occlusion.PR, 47(6):2280– 2292, 2014

  80. [80]

    Tsai and Reimar Lenz

    Roger Y . Tsai and Reimar Lenz. A new technique for fully autonomous and efficient 3D robotics hand/eye calibration.IEEE T-RA, 5(3):345–358, 1989

Showing first 80 references.

This paper was first reviewed by grok-4.5 on July 31, 2026.