REVIEW 4 major objections 6 minor 78 references
Guiding only the wrist, not the fingers, is enough for full-body humanoid object-manipulation retargeting.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 16:01 UTC pith:VTHQZT6K
load-bearing objection Solid wrist-as-gate decoupling for sim HOI retargeting that beats full-finger baselines on the reported sequences, with real ablations, but still per-scene and metric-sensitive. the 4 major comments →
WristMimic: Full-Body Humanoid Control with Wrist-Guided Manipulation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
WristMimic establishes that whole-body humanoid control for human–object interaction can be decoupled into kinematic pose guidance for contact-free body parts (including the wrists) and object-and-contact-driven learning for the fingers, with the wrist as the structural bridge that places fingers in grasp affordances. With wrist-specific resets and reward prioritization, this finger-agnostic formulation matches or surpasses full finger-pose supervision on object-trajectory metrics and supports retargeting across diverse hand embodiments.
What carries the argument
Wrist-guided decoupling: inside a contact window, time-varying reward weights zero proximal arm joints and prioritize wrists, while phase-specific wrist reset thresholds (tightest in the grasping phase) terminate episodes that leave the feasible wrist region, so fingers can be shaped only by object-pose and contact rewards.
Load-bearing premise
That keeping the wrist inside tight phase-specific pose bounds during contact is enough for reinforcement learning to discover working finger grasps from object motion and contact rewards alone.
What would settle it
Train the same pipeline on the same sequences after removing phase-specific wrist resets (or after deliberately offsetting wrist motion-capture outside the grasping-phase bounds) while leaving object and contact rewards unchanged; if success rate and object errors stay near the reported WristMimic numbers, the claim that wrist guidance is the critical bridge fails.
If this is right
- Many HOI retargeting tasks do not need dense finger kinematic capture if wrist placement and object trajectories are reliable.
- The same policy recipe can retarget across hand sizes, joint lengths, and joint limits without finger-specific motion data.
- Contact-free body imitation plus object outcomes suffice once the wrist is constrained as the gate into grasp affordances.
- Non-hand contacts such as sitting or foot-pushing need no wrist-like design and fall back to ordinary whole-body tracking.
- Fine in-hand reorientation remains out of scope; the method targets affordance-aware grasp and object-trajectory tracking.
Where Pith is reading between the lines
- Dense multi-finger MoCap pipelines may be overbuilt for whole-body retargeting when wrist and object signals are high quality.
- Teleoperation interfaces that stream only wrist pose plus object goals could adopt the same gate idea and drop finger streaming.
- Scene-general policies may need a shared wrist-affordance prior rather than denser finger labels.
- Tasks where object outcomes under-determine finger pre-shaping (e.g., multi-finger in-hand rotation) are the natural stress test of the claim.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. WristMimic proposes a physics-based whole-body retargeting framework for human–object interaction that decouples contact-free body/wrist control from contact-rich finger manipulation. Body and wrist joints are supervised by kinematic pose targets, while the 30 finger joints receive no finger pose targets and are shaped only by object-pose tracking and contact rewards. The wrist is treated as the structural gate between regimes; within a contact window the method applies time-varying reward weights (relaxing proximal arm joints, preserving wrist weights) and phase-specific wrist reset thresholds (default 7 cm / 0.2 rad in the grasping phase). Policies are trained per scene with PPO in Isaac Gym. On 40 sequences from OMOMO and ParaHome, WristMimic reports substantially higher success and lower object pose error than InterMimic and SkillMimicV2, and ablations attribute gains to the wrist constraints. A hand-morphology study shows successful grasping across three hand models without finger kinematic data.
Significance. If the result holds under the stated scope, the paper offers a practical and well-motivated alternative to dense finger MoCap for physics-based HOI retargeting: wrist placement plus object/contact outcomes can be sufficient for affordance-consistent grasping and object trajectory tracking. The contribution is concrete—joint partition J_cf/J_cr, contact-window reward modulation (Eqs. 4–5), phase-specific wrist resets (Alg. 1), and supporting ablations (Tabs. 2, 4, 5 and Supp. Table A)—and is of clear interest to humanoid control and dexterous retargeting. Strengths include head-to-head numbers on two datasets, isolation of wrist design choices, an explicit negative result that adding finger guidance can hurt, and demonstrated transfer across hand morphologies without finger-specific data. The main significance is methodological rather than a universal controller: it reframes which degrees of freedom need kinematic supervision for contact-rich HOI.
major comments (4)
- [Abstract; §4.1; §4.2; §4.5] Abstract and §1 claim that WristMimic “matches or surpasses methods using full finger pose supervision” and enables “finger-agnostic retargeting.” All reported policies are trained per scene (§4.1, §4.5), including the InterMimic and SkillMimicV2 baselines (§4.2). InterMimic was originally designed for broader HOI coverage; under a strictly per-scene protocol the comparison does not establish superiority of a general wrist-guided controller over a general full-supervision controller. Either multi-scene / multi-object training results, or a clearly scoped claim limited to per-scene retargeting, is needed for the abstract claim to be load-bearing.
- [§4.1; §4.3; Table 1] The success definition in §4.1 (sequence completion without early termination, average object position error <10 cm, contact for ≥80% of reference contact duration) does not require fidelity of grasp type or finger configuration to the demonstration. §4.3 explicitly notes that in Move Book the policy keeps a two-hand support strategy that differs from the human release pattern because of stability. Combined with object-centric rewards and no finger pose terms, this metric can credit stable but non-reference grasps. The headline margin in Table 1 should be accompanied by stricter secondary metrics (e.g., per-frame object error distributions, grasp-type consistency, or a tighter position threshold) so that the advantage is not partly an artifact of the success criterion.
- [Table 1; §4.2] Table 1 reports InterMimic success of 0.1% on ParaHome versus 86.6% on OMOMO, while WristMimic remains high on both. A near-total failure of a full-supervision baseline on half the evaluation set is central to the claimed superiority. The manuscript should document that the per-scene InterMimic reimplementation used comparable compute, hyperparameter search, and contact/reward settings, or analyze why dense finger supervision collapses on ParaHome’s more dexterous handles. Without that, the large average gap (91.1% vs 43.3%) is hard to interpret as a pure effect of removing finger targets.
- [§3.3; Alg. 1; Table 4; Fig. 1] The weakest operational assumption is that phase-specific wrist kinematic bounds (default 7 cm / 0.2 rad in the grasping phase; Alg. 1 lines 21–25, Tab. 4) plus object/contact rewards suffice for PPO to discover successful finger configurations when wrist MoCap is imperfect or when pre-shaping is required. Tab. 4 only varies thresholds on four ParaHome scenes; there is no controlled injection of wrist noise, no evaluation when first-contact labels are uncertain, and no failure-mode analysis for grasps that need finger pre-shape before contact. A short sensitivity study on wrist reference noise and on contact-label timing would make the central “wrist as gate” claim more robust.
minor comments (6)
- [§3.2 Goal State] In §3.2 the goal state is said to exclude 30 finger joints, yet contact targets use 2 hand-level indicators; a one-sentence clarification that hand-level contact is not a finger pose target would avoid confusion.
- [§3.3; Eq. (5)] Eq. (5) introduces w_red without a numeric value in the main text; the value appears only in the supplement. Stating the default w_red in §3.3 would improve reproducibility.
- [Fig. 3; Fig. 5] Fig. 3 and Fig. 5 would benefit from a short caption note on which frames are first-contact versus mid-manipulation, so qualitative claims about wrist placement are easier to verify.
- [§2.1] Related work cites Chen et al. [9] as a hierarchical wrist-guided hand-level method; a clearer sentence on how whole-body single-policy decoupling differs from that hierarchical hand-only setup would sharpen novelty.
- [Supplementary A.1; §4.4] Supplementary Table A (finger guidance hurts) is important supporting evidence; consider promoting a one-row summary into the main ablation section.
- [Abstract; §4.2] Minor typography: “position only hand trajectories” and similar compounds in the abstract would read better with hyphens; “bone-vector” vs “bone vector” is inconsistent.
Circularity Check
No circularity: empirical RL design with wrist resets/rewards evaluated against external baselines and object-trajectory metrics.
full rationale
WristMimic is a standard physics-based RL method paper. The core design (decouple kinematic targets to body+wrist only; shape fingers via object-pose and contact rewards; enforce wrist placement via phase-specific reset thresholds and time-varying reward weights inside a contact window) is a set of explicit engineering choices (Alg. 1, Eqs. 3–5, §3.3). These are then trained with PPO and measured on independent success criteria (completion without early termination + average object position error <10 cm + contact maintained for ≥80 % of reference duration) and against external baselines (InterMimic, SkillMimicV2) on held-out sequences from OMOMO and ParaHome. Nothing is fitted to a quantity that is later reported as a “prediction,” no uniqueness theorem is imported from the authors’ prior work, and self-citations appear only as related-work background or as the baselines being compared against. The supplementary finger-guidance ablation further shows that adding the very kinematic finger targets the method deliberately omits actually degrades the same metrics, confirming that the reported gains are not definitional. Evaluation caveats (per-scene policies, success definition that tolerates non-reference grasps) affect claims of generality or superiority but do not create circularity by construction.
Axiom & Free-Parameter Ledger
free parameters (4)
- grasping-phase wrist reset thresholds =
7 cm / 0.2 rad
- contact-window extents and phase splits =
tb=10, ta=15, τ1=-2, τ2=12
- time-varying joint reward weights (w_arm=0, w_red, wrist=1) =
arm=0, reduced body, wrist=1 in window
- multiplicative reward component scales λ =
see supp. Tab. F
axioms (4)
- domain assumption The wrist is largely free from object contact, can be tracked kinematically, and largely determines reachable finger grasp affordances.
- domain assumption Object kinematics plus contact labels supply sufficient outcome-level supervision for fingers to learn successful manipulation once the wrist is correctly placed.
- domain assumption Isaac Gym / PhysX with the listed PD gains, contact offsets and convex-hull settings adequately captures the contact dynamics needed for the reported grasps.
- standard math PPO with multiplicative tracking rewards optimizes the intended multi-objective imitation objective.
invented entities (2)
-
Phase-specific wrist reset thresholds inside a contact window
no independent evidence
-
J_cf / J_cr joint partition with wrist as the explicit gate
no independent evidence
Cite this review
Pith. "Pith review of WristMimic: Full-Body Humanoid Control with Wrist-Guided Manipulation." pith.science (2026). https://pith.science/paper/VTHQZT6K
@misc{pith2026260706438,
author = {Pith},
title = {Pith review of: WristMimic: Full-Body Humanoid Control with Wrist-Guided Manipulation},
year = {2026},
howpublished = {\url{https://pith.science/paper/VTHQZT6K}},
note = {Machine review of arXiv:2607.06438}
}
read the original abstract
Retargeting human object interaction demonstrations to physics based simulation requires reproducing not only body motion but also the object motion and contacts that make manipulation succeed. However, position only hand trajectories do not specify the contact forces needed to manipulate objects, and directly tracking them can overconstrain contact rich finger behavior. We introduce WristMimic, a wrist guided whole body control framework that explicitly separates contact free body motion from contact rich hand manipulation. The contact free body and wrist are guided by kinematic pose targets, whereas the fingers are not directly supervised by human hand pose. Instead, they learn grasping and manipulation behaviors from object tracking and contact outcomes. Our key insight is that the wrist is the natural gate between these two regimes. It is largely free from contact and can be tracked kinematically, yet it determines the global hand configuration and places the fingers within reachable grasp affordances. To ensure reliable wrist placement during interaction, we introduce wrist specific reset constraints and reward prioritization. Experiments show that WristMimic matches or surpasses methods using full finger pose supervision while enabling finger agnostic retargeting across diverse hand embodiments.
Figures
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:1910.07113 (2019)
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al.: Solving rubik’s cube with a robot hand. arXiv preprint arXiv:1910.07113 (2019)
Pith/arXiv arXiv 1910
-
[2]
The International Journal of Robotics Research39(1), 3–20 (2020)
Andrychowicz, O.M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pa- chocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al.: Learning dexterous in-hand manipulation. The International Journal of Robotics Research39(1), 3–20 (2020)
2020
-
[3]
In: International Conference on Computer Graphics and Interactive Techniques (2023)
Bae, J., Won, J., Lim, D., Min, C.H., Kim, Y.: Pmp: Learning to physically interact with environments using part-wise motion priors. In: International Conference on Computer Graphics and Interactive Techniques (2023)
2023
-
[4]
In: Proceedings of the Computer Vision and Pattern Recognition Conference
Banerjee, P., Shkodrani, S., Moulon, P., Hampali, S., Han, S., Zhang, F., Zhang, L., Fountain, J., Miller, E., Basol, S., et al.: Hot3d: Hand and object tracking in 3d from egocentric multi-view videos. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 7061–7071 (2025)
2025
-
[5]
In: 2024 International Conference on 3D Vision (3DV)
Braun, J., Christen, S., Kocabas, M., Aksan, E., Hilliges, O.: Physically plausible full-body hand-object interaction synthesis. In: 2024 International Conference on 3D Vision (3DV). pp. 464–473. IEEE (2024)
2024
-
[6]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Chao,Y.W.,Yang,W.,Xiang, Y.,Molchanov,P.,Handa,A., Tremblay, J.,Narang, Y.S., Van Wyk, K., Iqbal, U., Birchfield, S., et al.: Dexycb: A benchmark for capturing hand grasping of objects. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 9044–9053 (2021)
2021
-
[7]
arXiv preprint arXiv:2504.18829 (2025)
Chen, J., Ke, Y., Peng, L., Wang, H.: Dexonomy: Synthesizing all dexterous grasp types in a grasp taxonomy. arXiv preprint arXiv:2504.18829 (2025)
Pith/arXiv arXiv 2025
-
[8]
arXiv preprint arXiv:2309.00987 (2023)
Chen,Y.,Wang,C.,Fei-Fei,L.,Liu,C.K.:Sequentialdexterity:Chainingdexterous policies for long-horizon manipulation. arXiv preprint arXiv:2309.00987 (2023)
Pith/arXiv arXiv 2023
-
[9]
arXiv preprint arXiv:2411.04005 (2024)
Chen, Y., Wang, C., Yang, Y., Liu, C.K.: Object-centric dexterous manipulation from human motion data. arXiv preprint arXiv:2411.04005 (2024)
Pith/arXiv arXiv 2024
-
[10]
In: 2025 IEEE International Conference on Robotics and Automation (ICRA)
Chen, Z., Chen, S., Arlaud, E., Laptev, I., Schmid, C.: Vividex: Learning vision- based dexterous manipulation from human videos. In: 2025 IEEE International Conference on Robotics and Automation (ICRA). pp. 3336–3343. IEEE (2025)
2025
-
[11]
arXiv preprint arXiv:2506.14770 (2025)
Chen, Z., Ji, M., Cheng, X., Peng, X., Peng, X.B., Wang, X.: Gmt: General motion tracking for humanoid whole-body control. arXiv preprint arXiv:2506.14770 (2025)
Pith/arXiv arXiv 2025
-
[12]
arXiv preprint arXiv:2402.16796 (2024)
Cheng, X., Ji, Y., Chen, J., Yang, R., Yang, G., Wang, X.: Expressive whole-body control for humanoid robots. arXiv preprint arXiv:2402.16796 (2024)
Pith/arXiv arXiv 2024
-
[13]
IEEE transactions on robotics26(1), 1–20 (2009)
Dahiya, R.S., Metta, G., Valle, M., Sandini, G.: Tactile sensing—from humans to humanoids. IEEE transactions on robotics26(1), 1–20 (2009)
2009
-
[14]
In: Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition
Fan, Z., Taheri, O., Tzionas, D., Kocabas, M., Kaufmann, M., Black, M.J., Hilliges, O.: Arctic: A dataset for dexterous bimanual hand-object manipulation. In: Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 12943–12954 (2023)
2023
-
[15]
arXiv preprint arXiv:2406.10454 (2024)
Fu, Z., Zhao, Q., Wu, Q., Wetzstein, G., Finn, C.: Humanplus: Humanoid shad- owing and imitation from humans. arXiv preprint arXiv:2406.10454 (2024)
Pith/arXiv arXiv 2024
-
[16]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Hampali, S., Rad, M., Oberweger, M., Lepetit, V.: Honnotate: A method for 3d annotation of hand and object poses. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 3196–3206 (2020)
2020
-
[17]
In: 2023 IEEE 18 W
Handa, A., Allshire, A., Makoviychuk, V., Petrenko, A., Singh, R., Liu, J., Makovi- ichuk, D., Van Wyk, K., Zhurkevich, A., Sundaralingam, B., et al.: Dextreme: Transfer of agile in-hand manipulation from simulation to reality. In: 2023 IEEE 18 W. Yu et al. International Conference on Robotics and Automation (ICRA). pp. 5977–5984. IEEE (2023)
2023
-
[18]
In: ACM SIGGRAPH 2023 Conference Pro- ceedings
Hassan, M., Guo, Y., Wang, T., Black, M., Fidler, S., Peng, X.B.: Synthesizing physical character-scene interactions. In: ACM SIGGRAPH 2023 Conference Pro- ceedings. pp. 1–9 (2023)
2023
-
[19]
arXiv preprint arXiv:2507.02747 (2025)
He, J., Li, D., Yu, X., Qi, Z., Zhang, W., Chen, J., Zhang, Z., Zhang, Z., Yi, L., Wang, H.: Dexvlg: Dexterous vision-language-grasp model at scale. arXiv preprint arXiv:2507.02747 (2025)
Pith/arXiv arXiv 2025
-
[20]
arXiv preprint arXiv:2406.08858 (2024)
He,T.,Luo,Z.,He,X.,Xiao,W.,Zhang,C.,Zhang,W.,Kitani,K.,Liu,C.,Shi,G.: Omnih2o: Universal and dexterous human-to-humanoid whole-body teleoperation and learning. arXiv preprint arXiv:2406.08858 (2024)
Pith/arXiv arXiv 2024
-
[21]
Hendrycks, D., Gimpel, K.: Gaussian error linear units (GELUs) (Jun 2016)
2016
-
[22]
arXiv preprint arXiv:2410.24091 (2024)
Huang, B., Wang, Y., Yang, X., Luo, Y., Li, Y.: 3d-vitac: Learning fine-grained manipulation with visuo-tactile sensing. arXiv preprint arXiv:2410.24091 (2024)
Pith/arXiv arXiv 2024
-
[23]
In: 2025 IEEE International Conference on Robotics and Automation (ICRA)
Kareer, S., Patel, D., Punamiya, R., Mathur, P., Cheng, S., Wang, C., Hoffman, J., Xu, D.: Egomimic: Scaling imitation learning via egocentric video. In: 2025 IEEE International Conference on Robotics and Automation (ICRA). pp. 13226–13233. IEEE (2025)
2025
-
[24]
In: Proceedings of the Computer Vision and Pattern Recognition Conference
Kim, J., Kim, J., Na, J., Joo, H.: Parahome: Parameterizing everyday home activi- ties towards 3d generative modeling of human-object interactions. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 1816–1828 (2025)
2025
-
[25]
In: Proceedings of the IEEE/CVF international conference on computer vision
Kwon, T., Tekin, B., Stühmer, J., Bogo, F., Pollefeys, M.: H2o: Two hands ma- nipulating objects for first person interaction recognition. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 10138–10148 (2021)
2021
-
[26]
arXiv preprint arXiv:2601.16046 (2026)
Lee, J., Park, E., Cho, M.: Dexter: Language-driven dexterous grasp generation with embodied reasoning. arXiv preprint arXiv:2601.16046 (2026)
Pith/arXiv arXiv 2026
-
[27]
ACM Transactions on Graphics (TOG)42(6), 1–11 (2023)
Li, J., Wu, J., Liu, C.K.: Object motion guided human motion synthesis. ACM Transactions on Graphics (TOG)42(6), 1–11 (2023)
2023
-
[28]
arXiv preprint arXiv:2410.11792 (2024)
Li, J., Zhu, Y., Xie, Y., Jiang, Z., Seo, M., Pavlakos, G., Zhu, Y.: Okami: Teaching humanoid robots manipulation skills through single video imitation. arXiv preprint arXiv:2410.11792 (2024)
Pith/arXiv arXiv 2024
-
[29]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Li, K., Li, P., Liu, T., Li, Y., Huang, S.: Maniptrans: Efficient dexterous bimanual manipulation transfer via residual learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 6991–7003 (2025)
2025
-
[30]
In: European Conference on Computer Vision
Li, K., Wang, J., Yang, L., Lu, C., Dai, B.: Semgrasp: Semantic grasp generation via language aligned discretization. In: European Conference on Computer Vision. pp. 109–127. Springer (2024)
2024
-
[31]
arXiv preprint arXiv:2508.08241 (2025)
Liao, Q., Truong, T.E., Huang, X., Gao, Y., Tevet, G., Sreenath, K., Liu, C.K.: Beyondmimic: From motion tracking to versatile humanoid control via guided dif- fusion. arXiv preprint arXiv:2508.08241 (2025)
Pith/arXiv arXiv 2025
-
[32]
In: European Conference on Computer Vision
Lu, J., Kang, H., Li, H., Liu, B., Yang, Y., Huang, Q., Hua, G.: Ugg: Unified generative grasping. In: European Conference on Computer Vision. pp. 414–433. Springer (2024)
2024
-
[33]
Advances in Neural Information Processing Systems37, 2161–2184 (2024)
Luo, Z., Cao, J., Christen, S., Winkler, A., Kitani, K., Xu, W.: Omnigrasp: Grasp- ing diverse objects with simulated humanoids. Advances in Neural Information Processing Systems37, 2161–2184 (2024)
2024
-
[34]
Luo, Z., Cao, J., Kitani, K., Xu, W., et al.: Perpetual humanoid control for real- timesimulatedavatars.In:ProceedingsoftheIEEE/CVFInternationalConference on Computer Vision. pp. 10895–10904 (2023) WristMimic 19
2023
-
[35]
arXiv preprint arXiv:2310.04582 (2023)
Luo, Z., Cao, J., Merel, J., Winkler, A., Huang, J., Kitani, K., Xu, W.: Uni- versal humanoid motion representations for physics-based control. arXiv preprint arXiv:2310.04582 (2023)
Pith/arXiv arXiv 2023
-
[36]
arXiv preprint arXiv:2511.07820 (2025)
Luo, Z., Yuan, Y., Wang, T., Li, C., Chen, S., Castaneda, F., Cao, Z.A., Li, J., Minor, D., Ben, Q., et al.: Sonic: Supersizing motion tracking for natural humanoid whole-body control. arXiv preprint arXiv:2511.07820 (2025)
Pith/arXiv arXiv 2025
-
[37]
In: Proceedings of the IEEE/CVF international conference on computer vision
Mahmood, N., Ghorbani, N., Troje, N.F., Pons-Moll, G., Black, M.J.: Amass: Archive of motion capture as surface shapes. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 5442–5451 (2019)
2019
-
[38]
In: NeurIPS Datasets and Benchmarks (2021)
Makoviychuk, V., Wawrzyniak, L., Guo, Y., Lu, M., Storey, K., Macklin, M., Hoeller, D., Rudin, N., Allshire, A., Handa, A., State, G.: Isaac gym: High perfor- mance gpu-based physics simulation for robot learning. In: NeurIPS Datasets and Benchmarks (2021)
2021
-
[39]
arXiv preprint arXiv:2505.24853 (2025)
Mandi, Z., Hou, Y., Fox, D., Narang, Y., Mandlekar, A., Song, S.: Dexmachina: Functional retargeting for bimanual dexterous manipulation. arXiv preprint arXiv:2505.24853 (2025)
Pith/arXiv arXiv 2025
-
[40]
In: European Conference on Computer Vision
Meng, H., Jin, S., Liu, W., Qian, C., Lin, M., Ouyang, W., Luo, P.: 3d interacting hand pose estimation by hand de-occlusion and removal. In: European Conference on Computer Vision. pp. 380–397. Springer (2022)
2022
-
[41]
6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image
Moon, G., Yu, S.I., Wen, H., Shiratori, T., Lee, K.M.: Interhand2. 6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image. In: European Conference on Computer Vision. pp. 548–564. Springer (2020)
2020
-
[42]
In: Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition
Ohkawa, T., He, K., Sener, F., Hodan, T., Tran, L., Keskin, C.: Assemblyhands: Towards egocentric activity understanding via 3d hand pose estimation. In: Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 12999–13008 (2023)
2023
-
[43]
arXiv preprint arXiv:2511.09484 (2025)
Pan, C., Wang, C., Qi, H., Liu, Z., Bharadhwaj, H., Sharma, A., Wu, T., Shi, G., Malik, J., Hogan, F.: Spider: Scalable physics-informed dexterous retargeting. arXiv preprint arXiv:2511.09484 (2025)
arXiv 2025
-
[44]
In: Proceedings of the Computer Vision and Pattern Recognition Conference
Pan, L., Yang, Z., Dou, Z., Wang, W., Huang, B., Dai, B., Komura, T., Wang, J.: Tokenhsi: Unified synthesis of physical human-scene interactions through task tokenization. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 5379–5391 (2025)
2025
-
[45]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Pavlakos, G., Choutas, V., Ghorbani, N., Bolkart, T., Osman, A.A., Tzionas, D., Black, M.J.: Expressive body capture: 3d hands, face, and body from a single im- age. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 10975–10985 (2019)
2019
-
[46]
ACM Transactions On Graphics (TOG)37(4), 1–14 (2018)
Peng, X.B., Abbeel, P., Levine, S., Van de Panne, M.: Deepmimic: Example-guided deep reinforcement learning of physics-based character skills. ACM Transactions On Graphics (TOG)37(4), 1–14 (2018)
2018
-
[47]
Peng, X.B., Guo, Y., Halper, L., Levine, S., Fidler, S.: Ase: Large-scale reusable adversarialskillembeddingsforphysicallysimulatedcharacters.ACMTransactions On Graphics (TOG)41(4), 1–17 (2022)
2022
-
[48]
ACM Transactions on Graphics (TOG)40(4), 1–20 (2021)
Peng, X.B., Ma, Z., Abbeel, P., Levine, S., Kanazawa, A.: Amp: Adversarial motion priors for stylized physics-based character control. ACM Transactions on Graphics (TOG)40(4), 1–20 (2021)
2021
-
[49]
In: Conference on Robot Learning
Qi, H., Yi, B., Suresh, S., Lambeta, M., Ma, Y., Calandra, R., Malik, J.: General in-hand object rotation with vision and touch. In: Conference on Robot Learning. pp. 2549–2564. PMLR (2023) 20 W. Yu et al
2023
-
[50]
IEEE Robotics and Automation Letters7(4), 10873–10881 (2022)
Qin, Y., Su, H., Wang, X.: From one hand to multiple hands: Imitation learning for dexterous manipulation from single-camera teleoperation. IEEE Robotics and Automation Letters7(4), 10873–10881 (2022)
2022
-
[51]
arXiv preprint arXiv:1707.06347 (2017)
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., Klimov, O.: Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017)
Pith/arXiv arXiv 2017
-
[52]
arXiv preprint arXiv:2403.08716 (2024)
Si, Z., Zhang, G., Ben, Q., Romero, B., Xian, Z., Liu, C., Gan, C.: Difftactile: A physics-based differentiable tactile simulator for contact-rich robotic manipulation. arXiv preprint arXiv:2403.08716 (2024)
Pith/arXiv arXiv 2024
-
[53]
In: European conference on computer vision
Taheri, O., Ghorbani, N., Black, M.J., Tzionas, D.: Grab: A dataset of whole- body human grasping of objects. In: European conference on computer vision. pp. 581–600. Springer (2020)
2020
-
[54]
arXiv preprint arXiv:2505.19086 (2025)
Tessler, C., Jiang, Y., Coumans, E., Luo, Z., Chechik, G., Peng, X.B.: Masked- manipulator: Versatile whole-body control for loco-manipulation. arXiv preprint arXiv:2505.19086 (2025)
arXiv 2025
-
[55]
arXiv preprint arXiv:2410.03441 (2024)
Tevet, G., Raab, S., Cohan, S., Reda, D., Luo, Z., Peng, X.B., Bermano, A.H., van de Panne, M.: Closd: Closing the loop between simulation and diffusion for multi-task character control. arXiv preprint arXiv:2410.03441 (2024)
Pith/arXiv arXiv 2024
-
[56]
In: European Conference on Computer Vision
Turpin, D., Wang, L., Heiden, E., Chen, Y.C., Macklin, M., Tsogkas, S., Dickinson, S., Garg, A.: Grasp’d: Differentiable contact-rich grasp synthesis for multi-fingered hands. In: European Conference on Computer Vision. pp. 201–221. Springer (2022)
2022
-
[57]
In: IEEE International Conference on Robotics and Automation (ICRA)
Wang, R., Zhang, J., Chen, J., Xu, Y., Li, P., Liu, T., Wang, H.: Dexgraspnet: A large-scale robotic dexterous grasp dataset for general objects based on simula- tion. In: IEEE International Conference on Robotics and Automation (ICRA). pp. 11359–11366 (2023)
2023
-
[58]
Wang, Y., Zhao, Q., Yu, R., Tsui, H.W., Zeng, A., Lin, J., Luo, Z., Yu, J., Li, X., Chen, Q., et al.: Skillmimic: Learning basketball interaction skills from demonstra- tions.In:ProceedingsoftheComputerVisionandPatternRecognitionConference. pp. 17540–17549 (2025)
2025
-
[59]
Advances in Neural Information Processing Systems37, 46881–46907 (2024)
Wei, Y.L., Jiang, J.J., Xing, C., Tan, X.T., Wu, X.M., Li, H., Cutkosky, M., Zheng, W.S.: Grasp as you say: Language-guided dexterous grasp generation. Advances in Neural Information Processing Systems37, 46881–46907 (2024)
2024
-
[60]
In: Proceedings of the IEEE/CVF International Conference on Com- puter Vision
Wei, Y.L., Lin, M., Lin, Y., Jiang, J.J., Wu, X.M., Zeng, L.A., Zheng, W.S.: Afford- dexgrasp: Open-set language-guided dexterous grasp with generalizable-instructive affordance. In: Proceedings of the IEEE/CVF International Conference on Com- puter Vision. pp. 11818–11828 (2025)
2025
-
[61]
IEEE Robotics and Automation Letters9(12), 11834–11840 (2024)
Weng, Z., Lu, H., Kragic, D., Lundell, J.: Dexdiffuser: Generating dexterous grasps with diffusion models. IEEE Robotics and Automation Letters9(12), 11834–11840 (2024)
2024
-
[62]
Won, J., Gopinath, D., Hodgins, J.: A scalable approach to control diverse be- haviors for physically simulated characters. ACM Trans. Graph.39(4) (2020), https://doi.org/10.1145/3386569.3392381
-
[63]
In: Proceedings of the IEEE/CVF International Conference on Com- puter Vision
Wu, Z., Li, J., Xu, P., Liu, C.K.: Human-object interaction from human-level in- structions. In: Proceedings of the IEEE/CVF International Conference on Com- puter Vision. pp. 11176–11186 (2025)
2025
-
[64]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Xu, G.H., Wei, Y.L., Zheng, D., Wu, X.M., Zheng, W.S.: Dexterous grasp trans- former. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 17933–17942 (2024)
2024
-
[65]
arXiv preprint arXiv:2505.21864 (2025) WristMimic 21
Xu, M., Zhang, H., Hou, Y., Xu, Z., Fan, L., Veloso, M., Song, S.: Dexumi: Using human hand as the universal manipulation interface for dexterous manipulation. arXiv preprint arXiv:2505.21864 (2025) WristMimic 21
arXiv 2025
-
[66]
In: Proceedings of the Computer Vision and Pattern Recognition Conference
Xu, S., Ling, H.Y., Wang, Y.X., Gui, L.Y.: Intermimic: Towards universal whole- body control for physics-based human-object interactions. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 12266–12277 (2025)
2025
-
[67]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Xu, Y., Wan, W., Zhang, J., Liu, H., Shan, Z., Shen, H., Wang, R., Geng, H., Weng, Y., Chen, J., et al.: Unidexgrasp: Universal robotic dexterous grasping via learning diverse proposal generation and goal-conditioned policy. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 4737–4746 (2023)
2023
-
[68]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Yang, L., Li, K., Zhan, X., Wu, F., Xu, A., Liu, L., Lu, C.: Oakink: A large-scale knowledge repository for understanding hand-object interaction. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 20953–20962 (2022)
2022
-
[69]
arXiv preprint arXiv:2507.07356 (2025)
Yin, K., Zeng, W., Fan, K., Dai, M., Wang, Z., Zhang, Q., Tian, Z., Wang, J., Pang, J., Zhang, W.: Unitracker: Learning universal whole-body motion tracker for humanoid robots. arXiv preprint arXiv:2507.07356 (2025)
arXiv 2025
-
[70]
arXiv preprint arXiv:2310.16917 (2023)
Yu, K., Han, Y., Wang, Q., Saxena, V., Xu, D., Zhao, Y.: Mimictouch: Leveraging multi-modal human tactile demonstrations for contact-rich manipulation. arXiv preprint arXiv:2310.16917 (2023)
Pith/arXiv arXiv 2023
-
[71]
In: Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers
Yu,R.,Wang,Y.,Zhao,Q.,Tsui,H.W.,Wang,J.,Tan,P.,Chen,Q.:Skillmimic-v2: Learning robust and generalizable interaction skills from sparse and noisy demon- strations. In: Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers. pp. 1–11 (2025)
2025
-
[72]
arXiv preprint arXiv:2508.20085 (2025)
Yuan, Z., Wei, T., Gu, L., Hua, P., Liang, T., Chen, Y., Xu, H.: Hermes: Human- to-robot embodied learning from multi-source motion data for mobile dexterous manipulation. arXiv preprint arXiv:2508.20085 (2025)
Pith/arXiv arXiv 2025
-
[73]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Zhan, X., Yang, L., Zhao, Y., Mao, K., Xu, H., Lin, Z., Li, K., Lu, C.: Oakink2: A dataset of bimanual hands-object manipulation in complex task completion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 445–456 (2024)
2024
-
[74]
In: European Conference on Computer Vision
Zhang,H.,Christen,S.,Fan,Z.,Hilliges,O.,Song,J.:Graspxl:Generatinggrasping motions for diverse objects at scale. In: European Conference on Computer Vision. pp. 386–403. Springer (2024)
2024
-
[75]
arXiv preprint arXiv:2509.13833 (2025)
Zhang, Z., Guo, J., Chen, C., Wang, J., Lin, C., Lian, Y., Xue, H., Wang, Z., Liu, M., Lyu, J., et al.: Track any motions under any disturbances. arXiv preprint arXiv:2509.13833 (2025)
arXiv 2025
-
[76]
arXiv preprint arXiv:2602.16710 (2026)
Zheng, R., Niu, D., Xie, Y., Wang, J., Xu, M., Jiang, Y., Castañeda, F., Hu, F., Tan, Y.L., Fu, L., et al.: Egoscale: Scaling dexterous manipulation with diverse egocentric human data. arXiv preprint arXiv:2602.16710 (2026)
arXiv 2026
-
[77]
arXiv preprint arXiv:2502.20900 (2025)
Zhong, Y., Huang, X., Li, R., Zhang, C., Chen, Z., Guan, T., Zeng, F., Lui, K.N., Ye, Y., Liang, Y., et al.: Dexgraspvla: A vision-language-action framework towards general dexterous grasping. arXiv preprint arXiv:2502.20900 (2025)
arXiv 2025
-
[78]
In: Proceedings of the IEEE/CVF international conference on computer vision
Zimmermann, C., Ceylan, D., Yang, J., Russell, B., Argus, M., Brox, T.: Freihand: A dataset for markerless capture of hand pose and shape from single rgb images. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 813–822 (2019) 22 W. Yu et al. Supplementary Material A Additional Results A.1 Ablation on Finger Guidance ParaHom...
2019
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.