Pith. sign in

REVIEW 3 major objections 6 minor 2 cited by

A new dataset pairs human and multi-robot dexterous grasps on the same objects so grasp strategies can transfer across hand designs.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 19:55 UTC pith:XUTPNABU

load-bearing objection Solid multi-view paired human–multi-dexterous-robot grasp resource; the real soft spot is unquantified semantic pairing, not the capture stack. the 3 major comments →

arxiv 2604.14944 v2 pith:XUTPNABU submitted 2026-04-16 cs.RO cs.CV

HRDexDB: A Paired Human-Robot Dataset for Cross-Embodiment Dexterous Grasping

classification cs.RO cs.CV
keywords cross-embodimentdexterous graspinghuman-robot datasetcontact map transfergrasp retrievalmulti-view captureobject 6D posetactile sensing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper introduces HRDexDB, a real-world dataset that records comparable grasping trials by human hands and four different multi-finger robotic hands on a shared set of 100 objects. Prior collections usually cover only humans, only robots, or human–robot pairs with simple grippers, so they cannot show how contact and motion should change when morphology and kinematics change. HRDexDB supplies synchronized multi-view RGB, reconstructed 3D hand and robot motion, object 6D trajectories, success labels, and fingertip force signals for tactile-enabled hands. The authors use the pairs to learn robot-specific contact maps from human contacts and to retrieve robot grasp priors from human queries, and they also test hand and object pose estimators under heavy occlusion. A reader who wants robots that can use human-like tools cares because the dataset makes cross-embodiment transfer a measurable problem instead of a guess across mismatched sources.

Core claim

HRDexDB is presented as the first markerless dataset that pairs human dexterous grasps with grasps from multiple robotic hand embodiments on the same objects under comparable motions, delivering high-precision 3D agent and object trajectories, multi-view RGB, success labels, and tactile signals so that human-to-robot and robot-to-robot transfer can be trained and evaluated directly.

What carries the argument

The paired multi-modal capture and reconstruction pipeline: a calibrated 21-camera exocentric rig plus egocentric views, teleoperated robot trials matched to human demonstrations, MANO-based human hand fitting, and multi-view-consistent object 6D tracking that places human and robot sequences in one world frame over shared objects.

Load-bearing premise

The work assumes that a teleoperator who watches a human grasp and then performs a “semantically corresponding” robot grasp produces pairs aligned enough to teach transfer, even though dense motion correspondence across different hand shapes is still undefined.

What would settle it

Train contact-transfer or grasp-retrieval models on the human–robot pairs and compare real-world grasp success on held-out objects and hand embodiments against models that use only raw human contacts or unpaired robot data; if the paired models do not win, the claim that these pairs form a useful transfer benchmark fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Robot-specific contact maps learned from paired human grasps can raise grasp success over optimizing against human contact maps alone.
  • A shared embedding space can rank feasible robot grasp priors from a human hand–object query, including across objects.
  • Hand pose and object 6D pose methods can be stress-tested under the severe occlusions of real dexterous contact.
  • Cross-embodiment learning can train and evaluate on paired real trajectories rather than only synthetic grasps or unpaired sources.
  • Scaling the same protocol toward many more objects would widen the benchmark for transfer and perception.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Semantic teleoperation pairing may hide how hard true trajectory-level imitation is when finger count, joint limits, and timing differ.
  • Because tactile signals stay embodiment-specific, cross-hand force transfer may need a shared contact intermediate rather than raw force matching.
  • Perception models trained only on human hand–object data may systematically fail on robot hands; the paired views can measure that domain gap.
  • The same multi-view reconstruction setup could support bimanual or tool-use extensions if the collection protocol were broadened.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces HRDexDB, a paired human–robot dexterous grasping dataset with human hands and four robotic embodiments (Allegro V4/V5 Plus, Inspire RH56DFTP/RH56F1) over 100 shared objects. It reports 2.1K trials / ~24M frames with 21 exocentric + 2 egocentric RGB views, reconstructed MANO hand motion, robot states, object 6D poses, success labels, and tactile signals for tactile-enabled hands, all in a unified world frame. Capture uses a multi-camera rig and teleoperation (Xsens + MANUS); reconstruction combines HaMeR keypoints, triangulation, MANO fitting with SAM3 silhouettes, and FoundationStereo/FoundationPose with multi-view silhouette consistency. Downstream experiments include human-to-robot contact-map transfer (with CEDex-style optimization) and latent grasp retrieval, plus perception benchmarks under occlusion. The central claim is that this is the first markerless paired multi-robot dexterous dataset over shared objects and thus a foundational cross-embodiment benchmark.

Significance. If the resource holds up under scrutiny, it fills a genuine gap between human HOI datasets and robot-only or gripper-centric HROI collections (Table 1). Multi-view markerless RGB, multi-embodiment pairing on shared objects, object 6D trajectories, and optional tactile are practically useful for contact transfer, retrieval, and perception under occlusion. The contact-transfer setup usefully isolates the contact objective (same optimizer; Human-Contact vs Transferred-Contact) and reports both simulation and real-hardware lift success, which is stronger evidence than sim-only ablations. The main value is as a community resource rather than a new algorithmic theory; public release and expansion toward 1,000 objects would amplify impact.

major comments (3)
  1. §3.2 Paired Acquisition Protocol and Limitations ('Defining Trajectory Correspondence'): The foundational-benchmark claim (Abstract; §1; Table 1) depends on pairs being more than same-object co-captures. Pairing is defined only as a teleoperator watching the human demo and executing a 'semantically corresponding' grasp that preserves 'grasp intent' while allowing morphology, kinematics, and timing differences. No quantitative pair-quality metric is reported (contact-map IoU / part agreement between human and robot, wrist–object trajectory DTW, success-conditioned alignment, or inter-annotator agreement on intent). Without this, Table 2 gains could be driven by same-object geometry plus any successful robot grasp rather than by the advertised human–robot pairing. A short quantitative characterization of pair alignment (even on a subset) is load-bearing for the central claim.
  2. Table 2 (and surrounding §4.1 text): Transferred-Contact is reported as superior to Human-Contact in sim and real for both hands, but the numeric cells for Transferred are redacted/placeholder in the manuscript as provided (shown as boxes rather than percentages). Real-trial counts are modest (60/30 Inspire/Allegro). For a resource paper whose main empirical support for 'cross-embodiment transfer' is this table, the final success rates, confidence intervals or binomial uncertainty, and a clearer statement of train/test object/pose splits must be present and reproducible. Incomplete primary results undermine evaluation of the transfer claim.
  3. §4.2 Latent-Space Robot Grasp Retrieval: The task is well motivated, but the manuscript text as provided does not include the quantitative retrieval metrics (e.g., Recall@K, rank of true pair, same-object vs cross-object numbers) that would substantiate the shared-latent claim. Qualitative Figure 4 alone is insufficient. Either report the missing metrics with clear baselines (geometry-only / random / unpaired) or demote this section to a qualitative illustration so the paper's claims match the evidence.
minor comments (6)
  1. Table 1: Several comparison cells use dashes or incomplete frame/resolution entries for large robot datasets; a short caption note on what is unknown vs not applicable would reduce ambiguity when claiming uniqueness.
  2. §3.1 Eqs. (1)–(2): Clarify whether y (success) is defined identically for human and robot trials, and how failure modes are labeled (slip, collision, incomplete lift).
  3. §3.3 Object 6D Tracking: Report approximate tracking error or multi-view silhouette residual on a held-out subset so users know the ground-truth quality under heavy occlusion.
  4. Figure 3 / §4.1: State explicitly how human contact maps C_h are obtained from the reconstructed MANO + object mesh (threshold, proximity, or thermal/other), since supervision quality depends on it.
  5. Limitations already note heterogeneous tactile and undefined dense correspondence; cross-reference these earlier when claiming 'closely aligned captures' in the Abstract and §1 so claims and caveats stay consistent.
  6. Minor presentation: some figure text in the source appears garbled (e.g., 'MA!O'); ensure camera counts, hand names, and URLs are consistent between Abstract, Table 1, and §3.

Circularity Check

0 steps flagged

No circular derivation: HRDexDB is an empirical dataset-and-benchmark paper whose transfer and perception results are measured against held-out grasps and real hardware, not forced by definitional identities or fitted constants renamed as predictions.

full rationale

The paper’s central claim is the construction and release of a paired multi-embodiment grasping resource (Abstract; §1; Table 1), not a first-principles derivation. Contact-map transfer (§4.1) trains a supervised map from human (C_h, P_h) to robot (C_r, P_r) on paired trials, then freezes a fixed CEDex-style optimizer and compares Human-Contact vs Transferred-Contact success rates in simulation and on real lifts (Table 2); those rates are external empirical outcomes, not identities of the training loss. Latent grasp retrieval (§4.2) likewise ranks held-out robot candidates by embedding similarity and reports retrieval quality—again an evaluation, not a tautology. Reconstruction uses standard external tools (MANO, HaMeR, SAM3, FoundationStereo/Pose) with multi-view geometric constraints; none of these steps redefine the reported success or retrieval metrics. Self-citations (e.g., lab capture platforms) are ordinary engineering context and do not underwrite a uniqueness theorem or force the transfer gains. Semantic teleop pairing quality is a weak-assumption / correctness concern, not circularity: the paper does not claim the pairs are dense trajectory matches, and Limitations explicitly leaves dense correspondence open. No step reduces a claimed prediction to its own fitted input by construction.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 1 invented entities

As a dataset-and-benchmark paper, load-bearing content is mostly engineering assumptions about capture fidelity and semantic pairing rather than free physical constants. The central claim rests on the adequacy of the multi-view reconstruction stack, the teleop pairing protocol, and the claim of uniqueness versus prior datasets.

free parameters (2)
  • Contact-weighted L1 / CE loss weights and PointNet++ architecture choices for contact transfer
    Training details for the human-to-robot contact map model are only sketched; any loss weights, learning rates, or capacity choices are free design parameters that affect reported transfer success.
  • Grasp optimizer hyperparameters (CEDex contact/penetration/self-collision terms; BODex/CuRobo motion generation)
    Sim and real success rates depend on fixed but paper-external optimizer settings; they are not derived from first principles in this work.
axioms (4)
  • domain assumption Multi-view HaMeR keypoints + MANO fitting + SAM3 silhouettes yield sufficiently accurate 3D human hand ground truth for benchmark use.
    §3.3 Human Hand Reconstruction treats this pipeline as producing high-precision spatiotemporal GT without independent marker-based validation reported in the text.
  • domain assumption FoundationStereo + FoundationPose + multi-view silhouette consistency produce accurate object 6D trajectories under severe occlusion.
    §3.3 Object 6D Tracking; accuracy of the 'high-precision' claim depends on these foundation models.
  • ad hoc to paper A teleoperator-executed grasp that preserves 'grasp intent' after watching a human demo is a valid paired supervision signal for cross-embodiment learning.
    §3.2 Paired Acquisition Protocol; later admitted as an open problem for trajectory correspondence.
  • domain assumption No prior public dataset matches the combination of markerless multi-view paired human + multi-dexterous-robot grasps on shared objects with tactile.
    Table 1 and §1 'to the best of our knowledge' uniqueness claim; depends on completeness of the comparison set.
invented entities (1)
  • HRDexDB paired trial schema (T_robot / T_human with unified world frame, success label y, optional F_tactile) no independent evidence
    purpose: Defines the multi-modal unit of the dataset and what counts as a paired grasp episode.
    Standard dataset schema invention; useful but not an independent physical entity. independent_evidence is false until public release and external re-use.

pith-pipeline@v1.1.0-grok45 · 15087 in / 3178 out tokens · 38639 ms · 2026-07-12T19:55:31.963349+00:00 · methodology

0 comments
read the original abstract

We present HRDexDB, a paired cross-embodiment dexterous grasping dataset of high-fidelity dexterous grasping sequences featuring both human and diverse robotic hands. Unlike existing datasets, HRDexDB provides a comprehensive collection of grasping trajectories across human hands and multiple robot hand embodiments, spanning 100 diverse objects. Leveraging state-of-the-art vision methods and a dedicated multi-camera system, HRDexDB offers high-precision spatiotemporal 3D ground-truth motion for both the agent and the manipulated object. The dataset comprises 2.1K grasping trials, each enriched with synchronized visual and kinematic modalities, with contact-force signals available for tactile-enabled robotic hands. By providing closely aligned captures of human dexterity and robotic execution on the same target objects under comparable grasping motions, HRDexDB serves as a foundational benchmark for cross-embodiment dexterous manipulation.

Figures

Figures reproduced from arXiv: 2604.14944 by Byungjun Kim, Hanbyul Joo, Jisoo Kim, Jongbin Lim, Mingi Choi, Subin Jeon, Taeyun Ha.

Figure 1
Figure 1. Figure 1: Overview of HRDexDB. HRDexDB is a large-scale multimodal dataset containing 1.4K high-fidelity grasping episodes across 100 objects, with 4 different embodiments. Using a unified multi-camera capture system, we record paired human and robotic manipulation sequences with syn￾chronized modalities, including 3D hand and robot trajectories, object 6D poses, egocentric RGBD streams, tactile sensing, and success… view at source ↗
Figure 2
Figure 2. Figure 2: Visual Overview of Paired Human-Robot Grasping and Contact Maps. We visualize 48 representative objects from our collection, illustrating the diversity in geometry and functional categories. Each entry displays a paired grasping motion between a human hand (left) and a robotic hand (right), highlighting the similarity in grasping poses and contact patterns across different em￾bodiments. The color-coded con… view at source ↗
Figure 3
Figure 3. Figure 3: Capture System Overview. (Left) System architectures; (Middle) Capture protocol for human hand grasping; (Right) Capture protocol for robot hand via teleoperation using an IMU-based wearable motion capture device (Xsens and Manus Gloves). 3.2.2 Robotic Platform and Teleoperation System. The robotic platform consists of a 6-DOF xArm6 manipulator equipped with interchangeable end￾effectors. We use three dext… view at source ↗
Figure 4
Figure 4. Figure 4: An Example of Paired Grasp Capture Data. (Top) Human hand grasp. (Bottom) Robot [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Object 6D Pose Annotation Pipelines. Object 6D pose estimation from a calibrated stereo pair [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Hand Pose Annotation Pipelines. Hand pose reconstruction from multiview and silhouette-based optimization for MANO shape parameters. For each captured sequence, we fix β and optimize pose parameters frame-wise using the trian￾gulated joints. Temporal consistency is encouraged by initializing each frame from the previous solution and applying a One-Euro filter [32] to suppress high-frequency jitter. 3.5.2 O… view at source ↗
Figure 7
Figure 7. Figure 7: Pose consistency improves with increasing camera views. (Left) Overlay of object pose projections from two independent runs of the same tracking pipeline. With four cameras noticeable boundary discrepancies appear, while projections nearly coincide with 21 cameras. (Right) Mean Vertex Distance (MVD) across 20 static objects decreases as the number of views increases from 4 to 21 [PITH_FULL_IMAGE:figures/f… view at source ↗
Figure 8
Figure 8. Figure 8: Visualization of Geometric Affordance. We visualize the contact patterns by computing the spatial proximity between 3D meshes. The sub-millimeter tracking precision enables the capture of high-fidelity contact heatmaps for both human and robot hands across various objects. Note the consistent affordance patterns across different embodiments, demonstrating the semantic alignment of our paired demonstrations… view at source ↗
Figure 9
Figure 9. Figure 9: Embodiment-Specific Grasping Outcomes. (Left) Inspire F1 achieves stable force clo￾sure (71% success). (Right) Allegro hand fails (0% success) due to gravitational slippage despite initial contact. These results highlight that grasp success is strictly contingent on per-embodiment physical limits, such as actuation strength and friction dynamics. illustrated in [PITH_FULL_IMAGE:figures/full_fig_p011_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PhiZero: A World Model Built Around Physical Language

    cs.CV 2026-07 conditional novelty 6.0

    A self-supervised discrete physical-language bottleneck plus a VLM reasoner lets a world model predict state transitions before rendering video, improving physical coherence and enabling zero-shot motion transfer.

  2. Robot-Factored World Models via Robot Rendering

    cs.RO 2026-07 conditional novelty 6.0

    Conditioning a video world model on rendered nominal robot trajectories (URDF mesh + depth) instead of raw actions or logged future states improves action-following and enables zero-shot embodiment change.

Reference graph

Works this paper leans on

53 extracted references · 8 linked inside Pith · cited by 2 Pith papers

  1. [1]

    Moon, S.-I

    G. Moon, S.-I. Yu, H. Wen, T. Shiratori, and K. M. Lee. Interhand2. 6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image. InECCV, 2020

  2. [2]

    Garcia-Hernando, S

    G. Garcia-Hernando, S. Yuan, S. Baek, and T.-K. Kim. First-person hand action benchmark with rgb-d videos and 3d hand pose annotations. InCVPR, 2018

  3. [3]

    Brahmbhatt, C

    S. Brahmbhatt, C. Ham, C. C. Kemp, and J. Hays. Contactdb: Analyzing and predicting grasp contact via thermal imaging. InCVPR, 2019

  4. [4]

    Zimmermann, D

    C. Zimmermann, D. Ceylan, J. Yang, B. Russell, M. Argus, and T. Brox. Freihand: A dataset for markerless capture of hand pose and shape from single rgb images. InICCV, 2019

  5. [5]

    Y .-W. Chao, W. Yang, Y . Xiang, P. Molchanov, A. Handa, J. Tremblay, Y . S. Narang, K. Van Wyk, U. Iqbal, S. Birchfield, et al. Dexycb: A benchmark for capturing hand grasping of objects. InCVPR, 2021

  6. [6]

    Y . Liu, Y . Liu, C. Jiang, K. Lyu, W. Wan, H. Shen, B. Liang, Z. Fu, H. Wang, and L. Yi. Hoi4d: A 4d egocentric dataset for category-level human-object interaction. InCVPR, 2022

  7. [7]

    Z. Fan, O. Taheri, D. Tzionas, M. Kocabas, M. Kaufmann, M. J. Black, and O. Hilliges. Arctic: A dataset for dexterous bimanual hand-object manipulation. InCVPR, 2023

  8. [8]

    X. Zhan, L. Yang, Y . Zhao, K. Mao, H. Xu, Z. Lin, K. Li, and C. Lu. Oakink2: A dataset of bimanual hands-object manipulation in complex task completion. InCVPR, 2024

  9. [9]

    Hampali, M

    S. Hampali, M. Rad, M. Oberweger, and V . Lepetit. Honnotate: A method for 3d annotation of hand and object poses. InCVPR, 2020

  10. [10]

    Grauman, A

    K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. InCVPR, 2022

  11. [11]

    Damen, H

    D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al. The epic-kitchens dataset: Collection, challenges and baselines.IEEE TPAMI, 43(11):4125–4141, 2020

  12. [12]

    Q. Bu, J. Cai, L. Chen, X. Cui, Y . Ding, S. Feng, S. Gao, X. He, X. Hu, X. Huang, et al. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems.arXiv preprint arXiv:2503.06669, 2025

  13. [13]

    O’Neill, A

    A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. InICRA, 2024

  14. [14]

    S. Wu, X. Liu, S. Xie, P. Wang, X. Li, B. Yang, Z. Li, K. Zhu, H. Wu, Y . Liu, et al. Robocoin: An open-sourced bimanual robotic data collection for integrated manipulation.arXiv preprint arXiv:2511.17441, 2025

  15. [15]

    Khazatsky, K

    A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset. InRSSW, 2024

  16. [16]

    K. Wu, C. Hou, J. Liu, Z. Che, X. Ju, Z. Yang, M. Li, Y . Zhao, Z. Xu, G. Yang, et al. Robomind: Benchmark on multi-embodiment intelligence normative data for robot manipulation.arXiv preprint arXiv:2412.13877, 2024

  17. [17]

    H.-S. Fang, H. Fang, Z. Tang, J. Liu, C. Wang, J. Wang, H. Zhu, and C. Lu. Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot. InICRA, 2024. 10

  18. [18]

    T. Tao, M. K. Srirama, J. J. Liu, K. Shaw, and D. Pathak. Dexwild: Dexterous human interac- tions for in-the-wild robot policies.RSS, 2025

  19. [19]

    S. Xie, H. Cao, Z. Weng, Z. Xing, H. Chen, S. Shen, J. Leng, Z. Wu, and Y .-G. Jiang. Hu- man2robot: Learning robot actions from paired human-robot videos. InAAAI, 2026

  20. [20]

    Y . Liu, H. Yang, X. Si, L. Liu, Z. Li, Y . Zhang, Y . Liu, and L. Yi. Taco: Benchmarking generalizable bimanual tool-action-object understanding. InCVPR, 2024

  21. [21]

    J.-T. Song, J. Kim, J. Cao, Y . Lei, T. Yagi, and K. Kitani. Contact4d: A video dataset for whole-body human motion and finger contact in dexterous operations. In3DV, 2026

  22. [22]

    Banerjee, S

    P. Banerjee, S. Shkodrani, P. Moulon, S. Hampali, S. Han, F. Zhang, L. Zhang, J. Fountain, E. Miller, S. Basol, et al. Hot3d: Hand and object tracking in 3d from egocentric multi-view videos. InCVPR, 2025

  23. [23]

    R. Fu, D. Zhang, A. Jiang, W. Fu, A. Funk, D. Ritchie, and S. Sridhar. Gigahands: A massive annotated dataset of bimanual hand activities. InCVPR, 2025

  24. [24]

    Y . Liu, Y . Yang, Y . Wang, X. Wu, J. Wang, Y . Yao, S. Schwertfeger, S. Yang, W. Wang, J. Yu, et al. Realdex: Towards human-like grasping for robotic dexterous hand.arXiv preprint arXiv:2402.13853, 2024

  25. [25]

    R. Wang, J. Zhang, J. Chen, Y . Xu, P. Li, T. Liu, and H. Wang. Dexgraspnet: A large-scale robotic dexterous grasp dataset for general objects based on simulation. InICRA, 2023

  26. [26]

    P. Li, T. Liu, Y . Li, Y . Geng, Y . Zhu, Y . Yang, and S. Huang. Gendexgrasp: Generalizable dexterous grasping. InICRA, 2023

  27. [27]

    Zhang, S

    H. Zhang, S. Christen, Z. Fan, O. Hilliges, and J. Song. GraspXL: Generating grasping motions for diverse objects at scale. InECCV, 2024

  28. [28]

    Romero, D

    J. Romero, D. Tzionas, and M. J. Black. Embodied hands: Modeling and capturing hands and bodies together.ACM TOG (SIGGRAPH Asia), 36(6):245:1–245:17, Nov. 2017

  29. [29]

    Pavlakos, D

    G. Pavlakos, D. Shan, I. Radosavovic, A. Kanazawa, D. Fouhey, and J. Malik. Reconstructing hands in 3d with transformers. InCVPR, 2024

  30. [30]

    Carion, L

    N. Carion, L. Gustafson, Y .-T. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V . Alwala, H. Khedr, A. Huang, et al. Sam 3: Segment anything with concepts.arXiv preprint arXiv:2511.16719, 2025

  31. [31]

    B. Wen, M. Trepte, J. Aribido, J. Kautz, O. Gallo, and S. Birchfield. Foundationstereo: Zero- shot stereo matching. InCVPR, 2025

  32. [32]

    B. Wen, W. Yang, J. Kautz, and S. Birchfield. Foundationpose: Unified 6d pose estimation and tracking of novel objects. InCVPR, 2024

  33. [33]

    Z. Wu, R. A. Potamias, X. Zhang, Z. Zhang, J. Deng, and S. Luo. Cedex: Cross-embodiment dexterous grasp generation at scale from human-like contact representations. InICRA, 2026

  34. [34]

    C. R. Qi, L. Yi, H. Su, and L. J. Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. InNeurIPS, 2017

  35. [35]

    Makoviychuk, L

    V . Makoviychuk, L. Wawrzyniak, Y . Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, and G. State. Isaac gym: High performance gpu-based physics simula- tion for robot learning.arXiv preprint arXiv:2108.10470, 2021

  36. [36]

    J. Chen, Y . Ke, and H. Wang. Bodex: Scalable and efficient robotic dexterous grasp synthesis using bilevel optimization. InICRA, 2025. 11

  37. [37]

    Sundaralingam, A

    B. Sundaralingam, A. Murali, and S. Birchfield. curobov2: Dynamics-aware motion generation with depth-fused distance fields for high-dof robots.arxiv preprint arXiv:2603.05493, 2026

  38. [38]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervi- sion. InICML, 2021

  39. [39]

    Y . Qin, W. Yang, B. Huang, K. Van Wyk, H. Su, X. Wang, Y .-W. Chao, and D. Fox. Anyteleop: A general vision-based dexterous robot arm-hand teleoperation system. InRSS, 2023

  40. [40]

    R. A. Potamias, J. Zhang, J. Deng, and S. Zafeiriou. Wilor: End-to-end 3d hand localization and reconstruction in-the-wild. InCVPR, 2025

  41. [41]

    H. Dong, A. Chharia, W. Gou, F. Vicente Carrasco, and F. D. De la Torre. Hamba: Single-view 3d hand reconstruction with graph-guided bi-scanning mamba. InNeurIPS, 2024

  42. [42]

    K. Lin, L. Wang, and Z. Liu. Mesh graphormer. InICCV, 2021

  43. [43]

    Y . Rong, T. Shiratori, and H. Joo. Frankmocap: A monocular 3d whole-body pose estimation system via regression and integration. InICCVW, 2021

  44. [44]

    Xiang, H

    D. Xiang, H. Joo, and Y . Sheikh. Monocular total capture: Posing face, body, and hands in the wild. InCVPR, 2019

  45. [45]

    S. Jin, L. Xu, J. Xu, C. Wang, W. Liu, C. Qian, W. Ouyang, and P. Luo. Whole-body human pose estimation in the wild. InECCV, 2020

  46. [46]

    Zimmermann and T

    C. Zimmermann and T. Brox. Learning to estimate 3d hand pose from single rgb images. In ICCV, 2017

  47. [47]

    H.-S. Fang, J. Li, H. Tang, C. Xu, H. Zhu, Y . Xiu, Y .-L. Li, and C. Lu. Alphapose: Whole- body regional multi-person pose estimation and tracking in real-time.IEEE TPAMI, 45(6): 7157–7173, 2022

  48. [48]

    Simon, H

    T. Simon, H. Joo, I. Matthews, and Y . Sheikh. Hand keypoint detection in single images using multiview bootstrapping. InCVPR, 2017

  49. [49]

    E. P. ¨Ornek, Y . Labb´e, B. Tekin, L. Ma, C. Keskin, C. Forster, and T. Hoda ˇn. Foundpose: Unseen object pose estimation with foundation features. InECCV, 2024

  50. [50]

    V . N. Nguyen, T. Groueix, M. Salzmann, and V . Lepetit. Gigapose: Fast and robust novel object pose estimation via one correspondence. InCVPR, 2024

  51. [51]

    L. Liu, J. Lin, Z. Liu, and K. Jia. Picopose: Progressive pixel-to-pixel correspondence learning for novel object pose estimation. InCoRL, 2025

  52. [52]

    Labb´e, L

    Y . Labb´e, L. Manuelli, A. Mousavian, S. Tyree, S. Birchfield, J. Tremblay, J. Carpentier, M. Aubry, D. Fox, and J. Sivic. Megapose: 6d pose estimation of novel objects via render & compare. InCoRL, 2022

  53. [53]

    J. Lee, S. Han, J. Kim, I. Lee, M. Choi, J. Kim, W. Woo, and H. Joo. Omnirobothome: A multi-camera platform for real-time multiadic human-robot interaction.arXiv preprint arXiv:2604.28197, 2026. 12