Pith. sign in

REVIEW 3 major objections 6 minor 116 references

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

T0 review · 3 major / 6 minor · reviewed 2026-07-31 · grok-4.5

Pith's one-line read Everyday household skill can be recorded as one time-aligned stream of vision, body, hands, objects, sound, and touch in real homes.

desk verdict Solid dual-scale home capture with true multi-modal GT and honest SOTA gaps; limitations are real but already owned, and the paper deserves referees. read the letter →

arxiv 2607.28625 v1 pith:UJK7YLOX submitted 2026-07-30 cs.CV

classification cs.CV
keywords embodiedAIhuman-objectinteractionegocentricvisionmultimodalcapturemotiontactilesensinglong-horizonhouseholdactivityvision-language-action
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Embodied AI is stuck because the data it needs—how first-person sight, whole-body motion, dexterous hands, object state, sound, and contact evolve together while people chase household goals—has never been recorded as one registered stream. Existing datasets split that loop across viewpoints, modalities, or lab-scale setups, so models never see the full perception–action cycle. This paper’s answer is the Ambient Capture Engine: turn real homes into calibrated studios at two scales (table for fine hand–object work; room for locomotion and whole-body activity) and capture egocentric plus multi-view video, full-body and hand motion, object 6-DoF trajectories, audio, and tactile pressure on a shared clock and world frame. From that engine comes ACE-Data-0—150 hours, 17M frames, 75,000 episodes, 200 task categories, 50 people, goal-level instructions rather than scripted steps—plus a hierarchical benchmark from tactile-from-vision up through body pose to hand–object motion. State-of-the-art methods fail under contact, occlusion, egomotion, and long horizons, which is exactly why the authors argue this synchronized supervision is a foundation for imitation, world models, and vision-language-action systems.

What carries the argument

Ambient Capture Engine (ACE): two complementary home configurations (table-scale dense close-range capture; room-scale full-apartment coverage) that share optical-clock temporal sync to the motion-capture timeline and marker-bridged spatial calibration into one world frame, so every modality describes the same physical moment.

What would settle it

Train imitation or VLA policies on ACE-Data-0 and test zero-shot or few-shot transfer in an uninstrumented third home with novel layouts, lighting, and unmarked everyday objects; if performance collapses relative to in-distribution homes, the representativeness claim fails.

Watch

Extended reading notes

Core claim

The full perception–action loop of everyday home interaction can be captured holistically: dual-scale ambient instrumentation in real furnished homes yields a single metrically registered, millisecond-aligned multisensory stream (ego and multi-exo video, body and hand motion, object geometry and 6-DoF, audio, tactile), and on the resulting long-horizon ACE-Data-0 benchmark current methods show large, systematic gaps under contact, occlusion, egomotion, and extended time.

Load-bearing premise

Two instrumented homes, a few dozen pre-scanned markered objects, and people wearing suits, gloves, headsets, and markers still produce behavior and visuals representative enough to train general embodied agents.

Editorial extensions

If this is right

  • Imitation and policy learning can supervise contact, kinematics, and multi-view vision from the same physical instant rather than stitching mismatched sources.
  • World models and VLA systems gain long-horizon household chains with natural subtask order, hesitation, and recovery under goal-level instructions.
  • Benchmarks can diagnose failure mode (missed contact vs bad object state vs trajectory drift) instead of only final task success.
  • Egocentric and exocentric hand/body estimators can be compared and fused under identical measured ground truth.
  • Tactile-from-vision becomes a learnable mapping because every frame is paired with glove pressure on a shared clock.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If suit and marker appearance dominate learned features, domain randomization or appearance stripping may be required before robot transfer works.
  • Extending ground truth beyond rigid marked objects to fluids, cloth, and articulated appliances is the natural next bottleneck the dataset itself surfaces.
  • Providing measured headset motion as an oracle input would cleanly separate hand-reconstruction error from egomotion error in future ego baselines.
  • Scaling sites and unmarked objects may matter more than scaling hours inside the same two layouts for open-world generalization.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The manuscript introduces the Ambient Capture Engine (ACE), a dual-scale (table-scale and room-scale) capture system that turns furnished homes into calibrated, time-synchronized multi-sensor studios, and releases ACE-Data-0: ~150 hours / 17M frames / 75,000 episodes across 200 task categories from 50 participants in 2 sites. Streams include egocentric and multi-view exocentric video, full-body and hand motion, object meshes and 6-DoF trajectories, audio, and tactile pressure, registered to a common OptiTrack-referenced timeline and world frame. Tasks are goal-level (atomic HOI, long-horizon HOI chains, HSI) rather than scripted atomic actions. A hierarchical benchmark evaluates tactile-from-vision, human motion recovery, and hand motion from ego/exo views; zero-shot tests of 30+ SOTA methods report large gaps under contact, occlusion, egomotion, and long horizons. The authors position the corpus as supervision for imitation learning, world models, and VLA systems.

Significance. If the capture fidelity, scale, and released annotations hold as described, this is a substantial systems and resource contribution for embodied AI. Prior HOI and egocentric datasets typically fragment viewpoint, modality, or spatial scale; ACE’s explicit dual-scale design plus measured (not purely estimated) body/hand/object/tactile alignment in real homes is a clear advance over lab-only mocap HOI and uninstrumented egocentric video. Strengths include a concrete synchronization and calibration pipeline (§4.1: OptiTrack reference, QR optical clock, hand-eye ego calibration, reported ms residuals and <3 px median reprojection), goal-level long-horizon collection (§4.2), rich measured annotations (§4.3), and unusually broad zero-shot tables (Tables 3–6). For a dataset/benchmark paper, that package is significant even without new learning algorithms.

major comments (3)
  1. [§5; Abstract; §1; §6] The abstract, introduction, and conclusion repeatedly claim ACE-Data-0 as a “scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI,” yet §5 only reports zero-shot evaluation of external pretrained checkpoints. There is no fine-tuning, imitation, world-model, or VLA experiment that uses ACE-Data-0 as training supervision. For a resource paper this is not fatal, but the training-utility claim is load-bearing in the framing: either add a minimal transfer/finetune study (e.g., hand or body estimator fine-tuned on a train split, or a small behavior-cloning baseline) or substantially tone down claims that currently outrun the evidence in §5.
  2. [§3.2.2; §5.3.1; Tables 5–6] Hand ground truth is heterogeneous across scales: Manus gloves at room scale versus RANSAC triangulation of 2D keypoints from eight exo cameras plus manual refinement at table scale (§3.2.2, §5.3.1), with joints masked when under-observed. Tables 5–6 and the cross-view analysis (§5.3.3) treat these as a single comparable GT. The paper should quantify residual error of the table-scale pipeline (e.g., against a glove subset or multi-view consistency), state which benchmark splits use which GT source, and confirm that masked joints do not selectively remove the hardest contact frames—otherwise ego/exo and method rankings may partly reflect GT quality rather than viewpoint difficulty.
  3. [Abstract; §1; Table 1; Limitations (§6)] Domain breadth is a real limit on the “real home / general embodied” claim: only two sites, ~50 pre-scanned markered rigid objects, and participants in mocap suits, gloves, and headsets (Limitations; §3.2; §4.2). The Limitations section already states this, but the main text and Table 1 still present ACE-Data-0 as closing the home-scene gap relative to in-the-wild ego corpora. Please make the scope explicit earlier (abstract/intro/Table 1 caption): instrumented rigid HOI/HSI in two instrumented homes, not open-world domestic diversity—and avoid implying coverage of fluids, articulations, or deformables that are explicitly unannotated.
minor comments (6)
  1. [Table 1] Table 1 is very dense; a short caption note defining Sync/LH and the checkmark variants (measured vs estimated vs partial) would help readers compare rows without hunting the legend paragraph.
  2. [§4.2.3; Fig. 8] Dataset statistics: abstract and §4.2.3 say 17M frames / 150 hours / 75,000 episodes; Fig. 8 uses “15M+” in one place. Please reconcile all headline numbers.
  3. [§5.1.1; Table 3] §5.1.1 metrics (Temp Acc., C-IoU, V-IoU, CoP) need precise definitions or a short appendix (thresholds, min-max aggregation, palm/fingertip mask). Reproducibility of Table 3 depends on them.
  4. [§4.3.1–4.3.2] Textual annotations rely on Gemini-3.1-pro-preview with human correction (§4.3.2). State approximate correction rate and whether language labels are used in any benchmark track (they appear unused in §5).
  5. [Project page / §6] Release plan: project and HF links are given; please state what will be public at acceptance (raw vs processed streams, meshes, sync tables, train/test splits, licenses) so the “foundation” claim is checkable.
  6. [Table 4; References] Minor polish: “W A-MPJPE” spacing in Table 4; a few repeated figure callouts; ensure all baseline citations in Tables 3–6 match the reference list versions used.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: dataset/capture paper with measured GT and external pretrained baselines.

full rationale

ACE-Data-0 is a systems, capture, and benchmark paper, not a first-principles derivation that recovers fitted constants as predictions. Load-bearing claims are engineering and empirical: dual-scale ambient capture in real homes, hardware/software sync and calibration to a common spatio-temporal frame, goal-level long-horizon collection, and a hierarchical benchmark on which released pretrained SOTA methods fail under contact, occlusion, egomotion, and long horizons. Ground truth is measured (OptiTrack markers, Manus gloves, tactile pressure, scanned/2DGS object meshes), not inferred from the models under test; annotations such as pose reprojections and boxes follow by projection from tracked states. Benchmark tables compare external checkpoints to held-out mocap/tactile GT. Text captions use an external VLM plus human correction—ordinary labeling, not a circular proof of multi-modal alignment. Author self-citations appear in related work and future-application context and are not load-bearing uniqueness or ansatz imports for the central dataset claim. No self-definitional loop, fitted-input-as-prediction, or renaming of a known result as a forced derivation was found.

Assumptions & free parameters 3 free parameters · 6 assumptions · 3 invented entities

Load-bearing commitments are engineering and sampling assumptions of the capture recipe, not free parameters in a fitted theory. The central scientific claim (unified ambient capture closes a data gap and stresses SOTA) rests on trusting mocap/tactile as GT, two-site home ecology as ‘real home’ distribution, goal-level prompting as natural behavior, and zero-shot pretrained evaluation as diagnostic of field readiness.

free parameters (3)
  • Sync residual / timing model (offset + slow drift line fit to QR clock readings) = ms-level residuals; ego mutual misalignment <2 ms; calibration re-estimated offsets <4 ms
    Per-take linear time mapping from camera timestamps to OptiTrack; reported ms-level residuals. Operational calibration choice that defines ‘synchronized’ labels.
  • Held-out test split size = 10 hours
    Benchmark uses 10 hours held out; selection criteria beyond ‘test set’ are lightly specified and affect all reported gaps.
  • Contact/pressure metric thresholds (C-IoU, V-IoU aggregation, CoP region definition)
    Tactile evaluation depends on thresholding and aggregation choices that are conventional but not uniquely determined.
assumptions (6)
  • domain assumption Optical mocap marker trajectories plus glove sensors constitute metric ground truth for body, hand, object pose, and contact pressure in evaluation.
    Entire benchmark hierarchy compares vision methods to these streams (§5); soft-tissue, marker error, and glove calibration limits are not quantified as uncertainty.
  • domain assumption Two furnished homes with dense truss-mounted sensors remain ecologically valid ‘real home’ interaction distributions despite instrumentation.
    Stated design principle §3.2.1 and Limitations (two sites; suit/gloves/markers visible).
  • domain assumption Goal-level verbal instructions (not step scripts) induce natural long-horizon household behavior suitable for imitation/VLA supervision.
    Task design §4.2.1; core differentiator vs atomic HOI datasets.
  • domain assumption Zero-shot evaluation of officially released pretrained checkpoints fairly exposes method gaps on home long-horizon data.
    Explicit protocol §5; does not test in-domain finetuning ceilings.
  • standard math Standard multi-view geometry, hand-eye calibration, bundle adjustment, and SMPL-X/MANO representations are adequate to register all streams into one spatio-temporal frame.
    Calibration §4.1.2 cites classical hand-eye and BA; errors reported as median reprojection px.
  • ad hoc to paper Gemini-generated then human-corrected language descriptions are acceptable semantic annotations aligned to measured activity.
    Annotation pipeline §4.3.2; only non-measured annotation family.
invented entities (3)
  • Ambient Capture Engine (ACE) dual-scale studio paradigm independent evidence
    purpose: Name the hardware+sync+calibration recipe that turns homes into multi-modal recording studios at table and room scale.
    Systems contribution; independently checkable via released data and described hardware, not a latent physical entity.
  • ACE-Data-0 corpus and hierarchical signal→component→interaction benchmark independent evidence
    purpose: Provide the concrete dataset and evaluation ladder claimed to ground embodied learning.
    Primary artifact; evidence is the release and reported baselines, external to any fitted theory.
  • ACE-Ego-Head-V02 Lite / ACE-Sense-Glove Lite
    purpose: Custom wearable ego and tactile sensing in the suite.
    Device names specific to the project; reproducibility depends on whether specs/data suffice without identical hardware.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine." pith.science (2026). https://pith.science/paper/UJK7YLOX

@misc{pith2026260728625,
  author       = {Pith},
  title        = {Pith review of: ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UJK7YLOX}},
  note         = {Machine review of arXiv:2607.28625}
}
read the original abstract

Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

116 extracted references · 9 linked inside Pith

  1. [1]

    Psychology press, 2014

    James J Gibson.The ecological approach to visual perception: classic edition. Psychology press, 2014

  2. [2]

    MIT press, 2006

    Rolf Pfeifer and Josh Bongard.How the body shapes the way we think: a new view of intelligence. MIT press, 2006

  3. [3]

    ImageNet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. InCVPR, 2009

  4. [4]

    LAION-5B: an open large-scale dataset for training next generation image-text models

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. LAION-5B: an open large-scale dataset for training next generation image-text models. InNeurIPS, 2022

  5. [5]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InICML, 2021

  6. [6]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhari- wal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Lit...

  7. [7]

    Ego-Exo4D: understanding skilled human activity from first- and third-person perspectives

    Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyl- los Afouras, et al. Ego-Exo4D: understanding skilled human activity from first- and third-person perspectives. InCVPR, 2024

  8. [8]

    Rescal- ing egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100.IJCV, 130 (1):33–55, 2022

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Rescal- ing egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100.IJCV, 130 (1):33–55, 2022

Show all 116 references
  1. [9]

    Xperience-10M: A large-scale egocentric multimodal dataset with structured 3D/4D annotations

    Ropedia. Xperience-10M: A large-scale egocentric multimodal dataset with structured 3D/4D annotations. Hugging Face dataset, 2026. URLhttps://huggingface.co/datasets/ ropedia-ai/xperience-10m

  2. [10]

    Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll

    Bharat Lal Bhatnagar, Xianghui Xie, Ilya A. Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. BEHA VE: dataset and method for tracking human object interactions. In CVPR, 2022

  3. [11]

    Black, and Dimitrios Tzionas

    Omid Taheri, Nima Ghorbani, Michael J. Black, and Dimitrios Tzionas. GRAB: A dataset of whole-body human grasping of objects. InECCV, 2020

  4. [12]

    Black, and Otmar Hilliges

    Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, and Otmar Hilliges. ARCTIC: A dataset for dexterous bimanual hand-object manipulation. InCVPR, 2023

  5. [13]

    OAKINK2: A dataset of bimanual hands-object manipulation in complex task completion

    Xinyu Zhan, Lixin Yang, Yifei Zhao, Kangrui Mao, Hanlin Xu, Zenan Lin, Kailin Li, and Cewu Lu. OAKINK2: A dataset of bimanual hands-object manipulation in complex task completion. InCVPR, 2024. 29

  6. [14]

    HOT3D: hand and object tracking in 3D from egocentric multi-view videos

    Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Shangchen Han, Fan Zhang, Linguang Zhang, Jade Fountain, Edward Miller, Selen Basol, Richard Newcombe, Robert Wang, Jakob Julian Engel, and Tomas Hodan. HOT3D: hand and object tracking in 3D from egocentric ...

  7. [15]

    Twigg, Charles C

    Samarth Brahmbhatt, Chengcheng Tang, Christopher D. Twigg, Charles C. Kemp, and James Hays. ContactPose: A dataset of grasps with object contact and hand pose. InECCV, 2020

  8. [16]

    Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Dieter Fox

    Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S. Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Dieter Fox. DexYCB: A benchmark for capturing hand grasping of objects. InCVPR, 2021

  9. [17]

    GigaHands: A massive annotated dataset of bimanual hand activities

    Rao Fu, Dingxi Zhang, Alex Jiang, Wanjia Fu, Austin Funk, Daniel Ritchie, and Srinath Sridhar. GigaHands: A massive annotated dataset of bimanual hand activities. InCVPR, 2025

  10. [18]

    H2O: two hands manipulating objects for first person interaction recognition

    Taein Kwon, Bugra Tekin, Jan Stuhmer, Federica Bogo, and Marc Pollefeys. H2O: two hands manipulating objects for first person interaction recognition. InICCV, 2021

  11. [19]

    World action models are zero-shot policies

    Seonghyeon Ye, Yunhao Ge, Kaiyuan Zheng, Shenyuan Gao, Sihyun Yu, George Kurian, Suneel Indupuru, You Liang Tan, Chuning Zhu, Jiannan Xiang, Ayaan Malik, Kyungmin Lee, William Liang, Nadun Ranawaka, Jiasheng Gu, Yinzhen Xu, Guanzhi Wang, Fengyuan Hu, Avnish Narayan, Johan Bjor...

  12. [20]

    EgoVLA: learning vision-language-action models from egocentric human videos

    Ruihan Yang, Qinxi Yu, Yecheng Wu, Rui Yan, Borui Li, An-Chieh Cheng, Xueyan Zou, Yunhao Fang, Xuxin Cheng, Ri-Zhao Qiu, Hongxu Yin, Sifei Liu, Song Han, Yao Lu, and Xiaolong Wang. EgoVLA: learning vision-language-action models from egocentric human videos. InCoRL, 2025

  13. [21]

    Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Robert Equi, Chelsea Finn, Niccolo Fusai, Manuel Y . Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, brian ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey L...

  14. [22]

    World Guidance: world modeling in condi- tion space for action generation.arXiv 2602.22010, 2026

    Yue Su, Sijin Chen, Haixin Shi, Mingyu Liu, Zhengshen Zhang, Ningyuan Huang, Weiheng Zhong, Zhengbang Zhu, Yuxiao Liu, and Xihui Liu. World Guidance: world modeling in condi- tion space for action generation.arXiv 2602.22010, 2026

  15. [23]

    Ego4D: around the world in 3,000 hours of egocentric video

    Kristen Grauman, Andrew Westbury, Eugene Byrne, et al. Ego4D: around the world in 3,000 hours of egocentric video. InCVPR, 2022

  16. [24]

    HoloAssist: an egocentric human interaction dataset for interactive AI assistants in the real world

    Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bugra Tekin, Felipe Vieira Frujeri, Neel Joshi, and Marc Pollefeys. HoloAssist: an egocentric human interaction dataset for interactive AI assistants in the real world. In...

  17. [25]

    EgoLife: towards egocentric life assistant

    Jingkang Yang, Shuai Liu, Hongming Guo, Yuhao Dong, Xiamengwei Zhang, Sicheng Zhang, Pengyun Wang, Zitang Zhou, Binzhu Xie, Ziyue Wang, Bei Ouyang, Zhengyu Lin, Marco Cominelli, Zhongang Cai, Bo Li, Yuanhan Zhang, Peiyuan Zhang, Fangzhou Hong, Joerg Wid- mer, Francesco Gringol...

  18. [26]

    HD-EPIC: A highly-detailed egocentric video dataset

    Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Parida, Kait- ing Liu, Prajwal Gatti, Siddhant Bansal, Kevin Flanagan, Jacob Chalk, Zhifan Zhu, Rhodri Guer- rier, Fahd Abdelazim, Bin Zhu, Davide Moltisanti, Michael Wray, Hazel Doughty, and Dima ...

  19. [27]

    HOnnotate: A method for 3D annotation of hand and object poses

    Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. HOnnotate: A method for 3D annotation of hand and object poses. InCVPR, 2020

  20. [28]

    OakInk: A large-scale knowledge repository for understanding hand-object interaction

    Lixin Yang, Kailin Li, Xinyu Zhan, Fei Wu, Anran Xu, Liu Liu, and Cewu Lu. OakInk: A large-scale knowledge repository for understanding hand-object interaction. InCVPR, 2022

  21. [29]

    HOI4D: A 4D egocentric dataset for category-level human-object interaction

    Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4D egocentric dataset for category-level human-object interaction. InCVPR, 2022

  22. [30]

    TACO: benchmarking generalizable bimanual Tool-ACtion-Object understanding

    Yun Liu, Haolin Yang, Xu Si, Ling Liu, Zipeng Li, Yuxiang Zhang, Yebin Liu, and Li Yi. TACO: benchmarking generalizable bimanual Tool-ACtion-Object understanding. InCVPR, 2024

  23. [31]

    Black, and Dimitrios Tzionas

    Yinghao Huang, Omid Taheri, Michael J. Black, and Dimitrios Tzionas. InterCap: joint marker- less 3D tracking of humans and objects in interaction.IJCV, 132(7):2551–2566, 2024

  24. [32]

    Full-body articulated human-object interaction

    Nan Jiang, Tengyu Liu, Zhexuan Cao, Jieming Cui, Zhiyuan Zhang, Yixin Chen, He Wang, Yixin Zhu, and Siyuan Huang. Full-body articulated human-object interaction. InICCV, 2023

  25. [33]

    EgoBody: human body shape and motion of interacting people from head- mounted devices

    Siwei Zhang, Qianli Ma, Yan Zhang, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, and Siyu Tang. EgoBody: human body shape and motion of interacting people from head- mounted devices. InECCV, 2022

  26. [34]

    Parkhi, Richard A

    Xiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters, Thomas Whelan, Chen Kong, Omkar M. Parkhi, Richard A. Newcombe, and Carl Yuheng Ren. Aria Digital Twin: A new benchmark dataset for egocentric 3D machine perception. InICCV, 2023

  27. [35]

    Karen Liu

    Jiaman Li, Jiajun Wu, and C. Karen Liu. Object motion guided human motion synthesis.ACM TOG, 42(6):197:1–197:11, 2023

  28. [36]

    HIMO: A new benchmark for full-body human interacting with multiple objects

    Xintao Lv, Liang Xu, Yichao Yan, Xin Jin, Congsheng Xu, Shuwen Wu, Yifan Liu, Lincheng Li, Mengxiao Bi, Wenjun Zeng, and Xiaokang Yang. HIMO: A new benchmark for full-body human interacting with multiple objects. InECCV, 2024

  29. [37]

    Scaling up dynamic human-scene interaction modeling

    Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, and Siyuan Huang. Scaling up dynamic human-scene interaction modeling. InCVPR, 2024

  30. [38]

    ParaHome: parameterizing ev- eryday home activities towards 3D generative modeling of human-object interactions

    Jeonghwan Kim, Jisoo Kim, Jeonghyeon Na, and Hanbyul Joo. ParaHome: parameterizing ev- eryday home activities towards 3D generative modeling of human-object interactions. InCVPR, 2024

  31. [39]

    Karen Liu, Ziwei Liu, Jakob J

    Lingni Ma, Yuting Ye, Fangzhou Hong, Vladimir Guzov, Yifeng Jiang, Rowan Postyeni, Luis Pesqueira, Alexander Gamino, Vijay Baiyya, Hyo Jin Kim, Kevin Bailey, David Soriano Fosas, C. Karen Liu, Ziwei Liu, Jakob J. Engel, Renzo De Nardi, and Richard A. Newcombe. Nymeria: A massi...

  32. [40]

    HUMOTO: A 4D dataset of mocap human object interactions

    Jiaxin Lu, Chun-Hao Paul Huang, Uttaran Bhattacharya, Qixing Huang, and Yi Zhou. HUMOTO: A 4D dataset of mocap human object interactions. InICCV, 2025

  33. [41]

    Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen- Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, Abraham Lee, Kuan Fang, Chelsea Finn, and Sergey Levine

    Homer Rich Walke, Kevin Black, Tony Z. Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen- Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, Abraham Lee, Kuan Fang, Chelsea Finn, and Sergey Levine. BridgeData V2: A dataset for robot learning at scale. InCoRL, 2023. 31

  34. [42]

    RH20T: A comprehensive robotic dataset for learning diverse skills in one-shot

    Haoshu Fang, Hongjie Fang, Zhenyu Tang, Jirong Liu, Chenxi Wang, Junbo Wang, Haoyi Zhu, and Cewu Lu. RH20T: A comprehensive robotic dataset for learning diverse skills in one-shot. InICRA, 2024

  35. [43]

    Karen Liu

    Chen Wang, Haochen Shi, Weizhuo Wang, Ruohan Zhang, Li Fei-Fei, and C. Karen Liu. DexCap: scalable and portable mocap data collection system for dexterous manipulation. InRSS, 2024

  36. [44]

    DROID: A large-scale in-the-wild robot manipulation dataset

    Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, et al. DROID: A large-scale in-the-wild robot manipulation dataset. InRSS, 2024

  37. [45]

    Ag- iBot World Colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems

    AgiBot-World-Contributors, Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, et al. Ag- iBot World Colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems. InRSS, 2025

  38. [46]

    Yoon, Mouli Sivapurapu, and Jian Zhang

    Ryan Hoque, Peide Huang, David J. Yoon, Mouli Sivapurapu, and Jian Zhang. EgoDex: learning dexterous manipulation from large-scale egocentric video. InICLR, 2026

  39. [47]

    Galaxea open-world dataset and G0 dual-system VLA model.arXiv 2509.00576, 2025

    Tao Jiang, Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Jianning Cui, Xiao Liu, Shuiqi Cheng, Jiyang Gao, Huazhe Xu, and Hang Zhao. Galaxea open-world dataset and G0 dual-system VLA model.arXiv 2509.00576, 2025

  40. [48]

    EgoScale: scaling dexterous manipulation with diverse egocentric human data.arXiv 2602.16710, 2026

    Ruijie Zheng, Dantong Niu, Yuqi Xie, Jing Wang, Mengda Xu, Yunfan Jiang, Fernando Casta˜neda, Fengyuan Hu, You Liang Tan, Letian Fu, Trevor Darrell, Furong Huang, Yuke Zhu, Danfei Xu, and Linxi Fan. EgoScale: scaling dexterous manipulation with diverse egocentric human data.ar...

  41. [49]

    Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and cap- turing hands and bodies together.ACM TOG, 36(6):245:1–245:17, 2017

  42. [50]

    Keypoint Transformer: solving joint identification in challenging hands and object interactions for accurate 3D pose estimation

    Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Keypoint Transformer: solving joint identification in challenging hands and object interactions for accurate 3D pose estimation. InCVPR, 2022

  43. [51]

    Reconstructing hands in 3D with transformers

    Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Ji- tendra Malik. Reconstructing hands in 3D with transformers. InCVPR, 2024

  44. [52]

    HaWoR: world- space hand motion reconstruction from egocentric videos

    Jinglei Zhang, Jiankang Deng, Chao Ma, and Rolandos Alexandros Potamias. HaWoR: world- space hand motion reconstruction from egocentric videos. InCVPR, 2025

  45. [53]

    Yufei Ye, Yao Feng, Omid Taheri, Haiwen Feng, Shubham Tulsiani, and Michael J. Black. Pre- dicting 4D hand trajectory from monocular videos. In3DV, 2026

  46. [54]

    HORT: monoc- ular hand-held objects reconstruction with transformers

    Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, and Cordelia Schmid. HORT: monoc- ular hand-held objects reconstruction with transformers. InICCV, 2025

  47. [55]

    PhySIC: physically plausible 3D human-scene interaction and contact from a single image

    Pradyumna Yalandur Muralidhar, Yuxuan Xue, Xianghui Xie, Margaret Kostyrko, and Gerard Pons-Moll. PhySIC: physically plausible 3D human-scene interaction and contact from a single image. InSIGGRAPH Asia, 2025

  48. [56]

    Reconstructing 4D spatial intelligence: A survey.arXiv 2507.21045, 2025

    Yukang Cao, Jiahao Lu, Zhisheng Huang, Zhuowen Shen, Chengfeng Zhao, Fangzhou Hong, Zhaoxi Chen, Xin Li, Wenping Wang, Yuan Liu, and Ziwei Liu. Reconstructing 4D spatial intelligence: A survey.arXiv 2507.21045, 2025

  49. [57]

    Humanoid policy∼human policy

    Ri-Zhao Qiu, Shiqi Yang, Xuxin Cheng, Chaitanya Chawla, Jialong Li, Tairan He, Ge Yan, Lars Paulsen, Ge Yang, Sha Yi, Guanya Shi, and Xiaolong Wang. Humanoid policy∼human policy. InCoRL, 2025. 32

  50. [58]

    ActionNet: A dataset for dexterous bimanual manipulation

    Fourier ActionNet Team and Yao Mu. ActionNet: A dataset for dexterous bimanual manipulation. Dataset website, 2025. URLhttps://action-net.org/

  51. [59]

    DexMV: imitation learning for dexterous manipulation from human videos

    Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang, Ruihan Yang, Yang Fu, and Xiaolong Wang. DexMV: imitation learning for dexterous manipulation from human videos. InECCV, 2022

  52. [60]

    EgoMimic: scaling imitation learning via egocentric video

    Simar Kareer, Dhruv Patel, Ryan Punamiya, Pranay Mathur, Shuo Cheng, Chen Wang, Judy Hoffman, and Danfei Xu. EgoMimic: scaling imitation learning via egocentric video. InICRA, 2025

  53. [61]

    Zhu, Simar Kareer, Judy Hoffman, and Danfei Xu

    Ryan Punamiya, Dhruv Patel, Patcharapong Aphiwetsa, Pranav Kuppili, Lawrence Y . Zhu, Simar Kareer, Judy Hoffman, and Danfei Xu. EgoBridge: domain adaptation for generalizable imitation from egocentric human data. InNeurIPS, 2025

  54. [62]

    UniVLA: learning to act anywhere with task-centric latent actions

    Qingwen Bu, Yanting Yang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Yao, Ping Luo, and Hongyang Li. UniVLA: learning to act anywhere with task-centric latent actions. InRSS, 2025

  55. [63]

    In-N-On: scaling egocentric manipulation with in-the-wild and on-task data

    Xiongyi Cai, Ri-Zhao Qiu, Geng Chen, Lai Wei, Isabella Liu, Tianshu Huang, Xuxin Cheng, and Xiaolong Wang. In-N-On: scaling egocentric manipulation with in-the-wild and on-task data. arXiv 2511.15704, 2025

  56. [64]

    Emergence of human to robot transfer in vision-language-action models

    Simar Kareer, Karl Pertsch, James Darpinian, Judy Hoffman, Danfei Xu, Sergey Levine, Chelsea Finn, and Suraj Nair. Emergence of human to robot transfer in vision-language-action models. arXiv 2512.22414, 2025

  57. [65]

    Do as I can, not as I say: Grounding language in robotic affordances

    Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, et al. Do as I can, not as I say: Grounding language in robotic affordances. InCoRL, 2022

  58. [66]

    Ren, Haohuan Wang, Jiaming Tang, Kyle Stachowicz, Karan Dhabalia, Michael Equi, Quan Vuong, Jost Tobias Springenberg, Sergey Levine, Chelsea Finn, and Danny Driess

    Marcel Torne, Karl Pertsch, Homer Walke, Kyle Vedder, Suraj Nair, Brian Ichter, Allen Z. Ren, Haohuan Wang, Jiaming Tang, Kyle Stachowicz, Karan Dhabalia, Michael Equi, Quan Vuong, Jost Tobias Springenberg, Sergey Levine, Chelsea Finn, and Danny Driess. MEM: multi-scale embodi...

  59. [67]

    Bridg- ing semantic and kinematic conditions with diffusion-based discrete motion tokenizer.arXiv 2603.19227, 2026

    Chenyang Gu, Mingyuan Zhang, Haozhe Xie, Zhongang Cai, Lei Yang, and Ziwei Liu. Bridg- ing semantic and kinematic conditions with diffusion-based discrete motion tokenizer.arXiv 2603.19227, 2026

  60. [68]

    InfiniteDance: scalable 3D dance generation towards in-the-wild general- ization

    Ronghui Li, Zhongyuan Hu, Li Siyao, Youliang Zhang, Haozhe Xie, Mingyuan Zhang, Jie Guo, Xiu Li, and Ziwei Liu. InfiniteDance: scalable 3D dance generation towards in-the-wild general- ization. InECCV, 2026

  61. [69]

    Wong, and Ziwei Liu

    Yukang Cao, Liang Pan, Kai Han, Kwan-Yee K. Wong, and Ziwei Liu. AvatarGO: zero-shot 4D human-object interaction generation and animation. InICLR, 2025

  62. [70]

    Zhao, and Chelsea Finn

    Zipeng Fu, Tony Z. Zhao, and Chelsea Finn. Mobile ALOHA: Learning bimanual mobile manip- ulation with low-cost whole-body teleoperation. InCoRL, 2024

  63. [71]

    RoboCOIN: an open-sourced bimanual robotic data COllection for INtegrated manipulation.arXiv 2511.17441, 2025

    Shihan Wu, Xuecheng Liu, Shaoxuan Xie, Pengwei Wang, Xinghang Li, et al. RoboCOIN: an open-sourced bimanual robotic data COllection for INtegrated manipulation.arXiv 2511.17441, 2025

  64. [72]

    RoboMIND 2.0: A mul- timodal, bimanual mobile manipulation dataset for generalizable embodied intelligence.arXiv 2512.24653, 2025

    Chengkai Hou, Kun Wu, Jiaming Liu, Zhengping Che, Di Wu, et al. RoboMIND 2.0: A mul- timodal, bimanual mobile manipulation dataset for generalizable embodied intelligence.arXiv 2512.24653, 2025. 33

  65. [73]

    HSImul3R: physics-in-the-loop reconstruction of simulation-ready human-scene interactions

    Yukang Cao, Haozhe Xie, Fangzhou Hong, Long Zhuo, Zhaoxi Chen, Liang Pan, and Ziwei Liu. HSImul3R: physics-in-the-loop reconstruction of simulation-ready human-scene interactions. In ECCV, 2026

  66. [74]

    WholeBodyVLA: towards unified latent VLA for whole-body loco-manipulation control

    Haoran Jiang, Jin Chen, Qingwen Bu, Li Chen, Modi Shi, Yanjie Zhang, Delong Li, Chuanzhe Suo, Chuang Wang, Zhihui Peng, and Hongyang Li. WholeBodyVLA: towards unified latent VLA for whole-body loco-manipulation control. InICLR, 2026

  67. [75]

    DynamicVLA: A vision-language-action model for dynamic object manipulation.arXiv 2601.22153, 2026

    Haozhe Xie, Beichen Wen, Jiarui Zheng, Zhaoxi Chen, Fangzhou Hong, Haiwen Diao, and Ziwei Liu. DynamicVLA: A vision-language-action model for dynamic object manipulation.arXiv 2601.22153, 2026

  68. [76]

    Zhu, Pranav Kuppili, Ryan Punamiya, Patcharapong Aphiwetsa, Dhruv Patel, Simar Kareer, Sehoon Ha, and Danfei Xu

    Lawrence Y . Zhu, Pranav Kuppili, Ryan Punamiya, Patcharapong Aphiwetsa, Dhruv Patel, Simar Kareer, Sehoon Ha, and Danfei Xu. EMMA: scaling mobile manipulation via egocentric human data.IEEE RA-L, 11(3):3087–3094, 2025

  69. [77]

    HoMMI: learning whole-body mobile manipulation from human demonstra- tions.arXiv 2603.03243, 2026

    Xiaomeng Xu, Jisang Park, Han Zhang, Eric Cousineau, Aditya Bhat, Jose Barreiros, Dian Wang, and Shuran Song. HoMMI: learning whole-body mobile manipulation from human demonstra- tions.arXiv 2603.03243, 2026

  70. [78]

    2D Gaussian splat- ting for geometrically accurate radiance fields

    Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2D Gaussian splat- ting for geometrically accurate radiance fields. InSIGGRAPH, 2024

  71. [79]

    Garrido-Jurado, R

    S. Garrido-Jurado, R. Mu ˜noz-Salinas, F.J. Madrid-Cuevas, and M.J. Mar´ın-Jim´enez. Automatic generation and detection of highly reliable fiducial markers under occlusion.PR, 47(6):2280– 2292, 2014

  72. [80]

    Tsai and Reimar Lenz

    Roger Y . Tsai and Reimar Lenz. A new technique for fully autonomous and efficient 3D robotics hand/eye calibration.IEEE T-RA, 5(3):345–358, 1989

  73. [81]

    McLauchlan, Richard I

    Bill Triggs, Philip F. McLauchlan, Richard I. Hartley, and Andrew W. Fitzgibbon. Bundle adjust- ment — A modern synthesis. InVision Algorithms: Theory and Practice, 2000

  74. [82]

    Unified temporal and spatial calibration for multi-sensor systems

    Paul Furgale, Joern Rehder, and Roland Siegwart. Unified temporal and spatial calibration for multi-sensor systems. InIROS, 2013

  75. [83]

    Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. InCVPR, 2019

  76. [84]

    Gemini: A family of highly capable multimodal models.arXiv 2312.11805, 2025

    Gemini Team. Gemini: A family of highly capable multimodal models.arXiv 2312.11805, 2025

  77. [85]

    Twigg, Chengde Wan, James Hays, and Charles C

    Patrick Grady, Chengcheng Tang, Samarth Brahmbhatt, Christopher D. Twigg, Chengde Wan, James Hays, and Charles C. Kemp. PressureVision: estimating hand pressure from a single RGB image. InECCV, 2022

  78. [86]

    EgoTactile: learning grasp pressure for everyday objects from egocentric video

    Yuan Zeng, Yujia Shi, Tiao Tan, Xingting Li, Yaqi Qin, Zongqing Lu, Wenming Yang, Jing-Hao Xue, and Qingmin Liao. EgoTactile: learning grasp pressure for everyday objects from egocentric video. InICML, 2026

  79. [87]

    TouchAnything: A dataset and framework for bimanual tactile estimation from ego- centric video.arXiv 2605.13083, 2026

    Jianyi Zhou, Ziteng Gao, Feiyang Hong, Zirui Liu, Guannan Zhang, Weisheng Dai, Ruichen Zhen, Chuqiao Lyu, Haotian Wu, Yinian Mao, Xushi Wang, Yuxiang Jiang, Wenbo Ding, and Shuo Yang. TouchAnything: A dataset and framework for bimanual tactile estimation from ego- centric vide...

  80. [88]

    FoundationPose: unified 6D pose esti- mation and tracking of novel objects

    Bowen Wen, Wei Yang, Jan Kautz, and Stan Birchfield. FoundationPose: unified 6D pose esti- mation and tracking of novel objects. InCVPR, 2024. 34

  81. [89]

    Multi-HMR: multi-person whole-body human mesh recovery in a single shot

    Fabien Baradel, Matthieu Armando, Salma Galaaoui, Romain Br ´egier, Philippe Weinzaepfel, Gr´egory Rogez, and Thomas Lucas. Multi-HMR: multi-person whole-body human mesh recovery in a single shot. InECCV, 2024

  82. [90]

    Multi-HMR 2: multi- person camera-centric human detection, mesh recovery and tracking.arXiv 2606.14841, 2026

    Gu ´enol´e Fiche, Philippe Weinzaepfel, Romain Br´egier, and Fabien Baradel. Multi-HMR 2: multi- person camera-centric human detection, mesh recovery and tracking.arXiv 2606.14841, 2026

  83. [91]

    SAM 3D Body: robust full-body human mesh recovery.arXiv 2602.15989, 2026

    Xitong Yang, Devansh Kukreja, Don Pinkus, Anushka Sagar, Taosha Fan, Jinhyung Park, Soyong Shin, Jinkun Cao, Jiawei Liu, Nicolas Ugrinovic, Matt Feiszli, Jitendra Malik, Piotr Doll ´ar, and Kris Kitani. SAM 3D Body: robust full-body human mesh recovery.arXiv 2602.15989, 2026

  84. [92]

    PyMAF-X: towards well-aligned full-body model regression from monocular images.IEEE TPAMI, 45(10):12287–12303, 2023

    Hongwen Zhang, Yating Tian, Yuxiang Zhang, Mengcheng Li, Liang An, Zhenan Sun, and Yebin Liu. PyMAF-X: towards well-aligned full-body model regression from monocular images.IEEE TPAMI, 45(10):12287–12303, 2023

  85. [93]

    Huang, Otmar Hilliges, and Michael J

    Muhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, and Michael J. Black. PARE: part attention regressor for 3D human body estimation. InICCV, 2021

  86. [94]

    Priyanka Patel and Michael J. Black. CameraHMR: aligning people with perspective. In3DV, 2025

  87. [95]

    One-stage 3D whole-body mesh recovery with component aware transformer

    Jing Lin, Ailing Zeng, Haoqian Wang, Lei Zhang, and Yu Li. One-stage 3D whole-body mesh recovery with component aware transformer. InCVPR, 2023

  88. [96]

    SMPLer-X: scaling up expressive human pose and shape estimation

    Zhongang Cai, Wanqi Yin, Ailing Zeng, Chen Wei, Qingping Sun, Wang Yanjun, Hui En Pang, Haiyi Mei, Mingyuan Zhang, Lei Zhang, Chen Change Loy, Lei Yang, and Ziwei Liu. SMPLer-X: scaling up expressive human pose and shape estimation. InNeurIPS, 2023

  89. [97]

    SMPLest-X: ultimate scaling for expressive human pose and shape estimation.IEEE TPAMI, 48(2):1778–1794, 2026

    Wanqi Yin, Zhongang Cai, Ruisi Wang, Ailing Zeng, Chen Wei, Qingping Sun, Haiyi Mei, Yan- jun Wang, Hui En Pang, Mingyuan Zhang, Lei Zhang, Chen Change Loy, Atsushi Yamashita, Lei Yang, and Ziwei Liu. SMPLest-X: ultimate scaling for expressive human pose and shape estimation.I...

  90. [98]

    Humans in 4D: Reconstructing and tracking humans with transformers

    Shubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa, and Jitendra Ma- lik. Humans in 4D: Reconstructing and tracking humans with transformers. InICCV, 2023

  91. [99]

    World-grounded human motion recovery via gravity-view coordinates

    Zehong Shen, Huaijin Pi, Yan Xia, Zhi Cen, Sida Peng, Zechen Hu, Hujun Bao, Ruizhen Hu, and Xiaowei Zhou. World-grounded human motion recovery via gravity-view coordinates. In SIGGRAPH Asia, 2024

  92. [100]

    Soyong Shin, Juyong Kim, Eni Halilaj, and Michael J. Black. WHAM: reconstructing world- grounded humans with accurate 3D motion. InCVPR, 2024

  93. [101]

    GitHub, 2021

    EasyMoCap: make human motion capture easier. GitHub, 2021. URLhttps://github. com/zju3dv/EasyMocap

  94. [102]

    Novel view synthesis of human interactions from sparse multi-view videos

    Qing Shuai, Chen Geng, Qi Fang, Sida Peng, Wenhao Shen, Xiaowei Zhou, and Hujun Bao. Novel view synthesis of human interactions from sparse multi-view videos. InSIGGRAPH, 2022

  95. [103]

    UniSH: unifying scene and human reconstruction in a feed-forward pass.arXiv 2601.01222, 2026

    Mengfei Li, Peng Li, Zheng Zhang, Jiahao Lu, Chengfeng Zhao, Wei Xue, Qifeng Liu, Sida Peng, Wenxiao Zhang, Wenhan Luo, Yuan Liu, and Yike Guo. UniSH: unifying scene and human reconstruction in a feed-forward pass.arXiv 2601.01222, 2026

  96. [104]

    Joint optimization for 4D human-scene reconstruction in the wild.arXiv 2501.02158, 2025

    Zhizheng Liu, Joe Lin, Wayne Wu, and Bolei Zhou. Joint optimization for 4D human-scene reconstruction in the wild.arXiv 2501.02158, 2025

  97. [105]

    Hu- man3R: everyone everywhere all at once

    Yue Chen, Xingyu Chen, Yuxuan Xue, Anpei Chen, Yuliang Xiu, and Gerard Pons-Moll. Hu- man3R: everyone everywhere all at once. InICLR, 2026. 35

  98. [106]

    Hanz Cuevas-Velasquez, Anastasios Yiannakidis, Soyong Shin, Giorgio Becherini, Markus H¨oschle, Joachim Tesch, Taylor Obersat, Tsvetelina Alexiadis, and Michael J. Black. MAMMA: markerless and automatic multi-person motion action capture. InCVPR, 2026

  99. [107]

    Human mesh recovery from arbitrary multi-view images.arXiv 2403.12434, 2024

    Xiaoben Li, Mancheng Meng, Ziyan Wu, Terrence Chen, Fan Yang, and Dinggang Shen. Human mesh recovery from arbitrary multi-view images.arXiv 2403.12434, 2024

  100. [108]

    Reconstructing people, places, and cameras

    Lea M ¨uller, Hongsuk Choi, Anthony Zhang, Brent Yi, Jitendra Malik, and Angjoo Kanazawa. Reconstructing people, places, and cameras. InCVPR, 2025

  101. [109]

    Karen Liu, and Jiajun Wu

    Jiaman Li, C. Karen Liu, and Jiajun Wu. Ego-body pose estimation via ego-head pose estimation. InCVPR, 2023

  102. [110]

    Estimating body and hand motion in an ego-sensed world

    Brent Yi, Vickie Ye, Maya Zheng, Yunqi Li, Lea M ¨uller, Georgios Pavlakos, Yi Ma, Jitendra Malik, and Angjoo Kanazawa. Estimating body and hand motion in an ego-sensed world. In CVPR, 2025

  103. [111]

    Human3.6M: large scale datasets and predictive methods for 3D human sensing in natural environments.IEEE TPAMI, 36 (7):1325–1339, 2014

    Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6M: large scale datasets and predictive methods for 3D human sensing in natural environments.IEEE TPAMI, 36 (7):1325–1339, 2014

  104. [112]

    Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. SMPL: a skinned multi-person linear model.ACM TOG, 34(6):248:1–248:16, 2015

  105. [113]

    3D hand pose estimation in everyday egocentric images

    Aditya Prakash, Ruisen Tu, Matthew Chang, and Saurabh Gupta. 3D hand pose estimation in everyday egocentric images. InECCV, 2024

  106. [114]

    Dyn-HaMR: recovering 4D interacting hand motion from a dynamic camera

    Zhengdi Yu, Stefanos Zafeiriou, and Tolga Birdal. Dyn-HaMR: recovering 4D interacting hand motion from a dynamic camera. InCVPR, 2025

  107. [115]

    WiLoR: end-to-end 3D hand localization and reconstruction in-the-wild

    Rolandos Alexandros Potamias, Jinglei Zhang, Jiankang Deng, and Stefanos Zafeiriou. WiLoR: end-to-end 3D hand localization and reconstruction in-the-wild. InCVPR, 2025

  108. [116]

    OmniHands: robust motion capture of interactive hands via a versatile trans- former.ACM TOG, 42(6):197:1–197:11, 2026

    Dixuan Lin, Yuxiang Zhang, Mengcheng Li, Yebin Liu, Wei Jing, Qi Yan, Qianying Wang, and Hongwen Zhang. OmniHands: robust motion capture of interactive hands via a versatile trans- former.ACM TOG, 42(6):197:1–197:11, 2026. 36

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.