Pith. sign in

REVIEW 44 cited by

Embodied Hands: Modeling and Capturing Hands and Bodies Together

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.02610 v1 pith:YZL232Z3 submitted 2022-01-07 cs.GR cs.CV

classification cs.GRcs.CV
keywords handshandmodelbodymanobodiesposeshape
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Humans move their hands and bodies together to communicate and solve tasks. Capturing and replicating such coordinated activity is critical for virtual characters that behave realistically. Surprisingly, most methods treat the 3D modeling and tracking of bodies and hands separately. Here we formulate a model of hands and bodies interacting together and fit it to full-body 4D sequences. When scanning or capturing the full body in 3D, hands are small and often partially occluded, making their shape and pose hard to recover. To cope with low-resolution, occlusion, and noise, we develop a new model called MANO (hand Model with Articulated and Non-rigid defOrmations). MANO is learned from around 1000 high-resolution 3D scans of hands of 31 subjects in a wide variety of hand poses. The model is realistic, low-dimensional, captures non-rigid shape changes with pose, is compatible with standard graphics packages, and can fit any human hand. MANO provides a compact mapping from hand poses to pose blend shape corrections and a linear manifold of pose synergies. We attach MANO to a standard parameterized 3D body shape model (SMPL), resulting in a fully articulated body and hand model (SMPL+H). We illustrate SMPL+H by fitting complex, natural, activities of subjects captured with a 4D scanner. The fitting is fully automatic and results in full body models that move naturally with detailed hand motions and a realism not seen before in full body performance capture. The models and data are freely available for research purposes in our website (http://mano.is.tue.mpg.de).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 44 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. C2Dex: Contact-Consistent Reconstruction and Retargeting for Dexterous Manipulation from Monocular Video

    cs.RO 2026-08 conditional novelty 7.0 of 10

    C2Dex converts monocular human videos into executable dexterous robot manipulation trajectories by using stable object-side contacts as a shared representation for reconstruction and retargeting, achieving 57.78% and ...

  2. InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A first large multimodal 4D human–dog interaction dataset (6.8M frames) plus an autoregressive model that generates dog motion from human body/hand gestures and audio.

  3. CustomDance: Customized 3D Dance Generation with Coarse-to-Fine Human-Centered Interactive Control

    cs.HC 2026-08 conditional novelty 6.0 of 10

    CustomDance combines an MLLM-based choreographic planner, multimodal dance-phrase retrieval, and diffusion inpainting into one three-stage interactive system for user-customized 3D dance generation.

  4. JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A dual action alignment framework (latent action world model plus canonical action space) converts heterogeneous human, simulated, and robot data into transferable supervision, with task scores rising monotonically as...

  5. BODIESReg: An Open-Source Pipeline for Registering 3D Body Scans Using Pose-Aligned Initialization

    q-bio.QM 2026-07 conditional novelty 6.0 of 10

    BODIESReg automatically registers 3D body scans to SMPL-family models using pose-aligned initialization, achieving 82.9% success on CHI3D and 100% on MorphoMotion with mean surface-fit error below 10mm.

  6. TactiDex: A Real-World Tactile-Guided Benchmark for Human-Like Dexterous Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A tactile-rich HOI dataset plus a tri-component force reward improves contact fidelity and success of human-to-robot dexterous transfer over kinematic imitation alone.

  7. Generative Relightable Avatars

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    GRA combines UV-space material optimization and physics rendering with feed-forward texture refinement and a fine-tuned video-to-video diffusion model to achieve controllable, high-detail relighting of full-body avatars.

  8. Grasp to Act: Dexterous Grasping for Tool Use in Dynamic Settings

    cs.RO 2026-02 conditional novelty 6.0 of 10

    Combining wrench-tested grasp optimization with real-time RL finger adjustments lets a 16-DoF robot hand keep tools stable during hammering, sawing, cutting, stirring, and scooping.

  9. SimGenHOI: Physically Realistic Whole-Body Humanoid-Object Interaction via Generative Modeling and Reinforcement Learning

    cs.RO 2025-08 conditional novelty 6.0 of 10

    SimGenHOI generates physically plausible humanoid-object interaction sequences by predicting sparse key actions with a diffusion transformer and tracking them with a contact-aware reinforcement learning policy in simulation.

  10. Extension of generalized KYP lemma: from LTI systems to LPV systems

    math.DS 2025-08 unverdicted novelty 6.0 of 10

    The abstract claims a gKYP lemma extension for LPV systems via frequency-range enlargement, but the submitted full text is an unrelated video generation paper.

  11. EPSilon: Efficient Point Sampling for Lightening of Hybrid-based 3D Avatar Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    EPSilon prunes empty rays and sampling intervals around the body mesh, cutting hybrid avatar rendering to 3.9% of the points and 20x faster inference with comparable quality.

  12. 3D Hand Mesh-Guided AI-Generated Malformed Hand Refinement with Hand Pose Transformation via Diffusion Model

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Using 3D hand meshes instead of depth maps to guide diffusion inpainting improves malformed hand refinement and enables training-free hand pose transformation.

  13. Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free Simulation

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Vid2Sim recovers 3D geometry, appearance, and elastic material parameters from multi-view videos using a feed-forward network plus a fast refinement, enabling mesh-free reduced-order simulation.

  14. EgoZero: Robot Learning from Smart Glasses

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Robot policies trained only on egocentric human videos from smart glasses transfer zero-shot to a Franka gripper, with 70% success across 7 manipulation tasks.

  15. MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MEgoHand generates egocentric hand-object interaction motions from an RGB image, a text instruction, and an initial MANO hand pose using VLM-based semantics, monocular depth, and flow matching.

  16. Hand-Shadow Poser

    cs.CG 2025-05 conditional novelty 6.0 of 10

    A three-stage AI pipeline estimates anatomically plausible bimanual hand poses whose cast shadows match a given target silhouette, outperforming differentiable-rendering and segmentation baselines on a 210-shape benchmark.

  17. TxP: Reciprocal Generation of Ground Pressure Dynamics and Activity Descriptions for Improving Human Activity Recognition

    cs.AI 2025-05 conditional novelty 6.0 of 10

    A bidirectional text-pressure model with a learned codebook generates synthetic pressure data from activity descriptions and classifies real pressure sequences via LLM-generated text, gaining up to 12.4 macro-F1 point...

  18. HOGSA: Bimanual Hand-Object Interaction Understanding with 3D Gaussian Splatting Based Data Augmentation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A 3D Gaussian Splatting based framework augments bimanual hand-object interaction datasets with diverse, realistic poses and views, improving baseline pose and contact estimation on Arctic and H2O.

  19. JADE: Joint-aware Latent Diffusion for 3D Human Generative Modeling

    cs.CV 2024-12 conditional novelty 6.0 of 10

    JADE learns a per-joint disentangled latent representation for 3D human bodies and generates new shapes with two cascaded diffusion models, one for skeleton structure and one for local surface geometry.

  20. A Powered Prosthetic Hand with Vision System for Enhancing the Anthropopathic Grasp

    cs.RO 2024-12 conditional novelty 6.0 of 10

    A vision-only prosthetic hand system maps hand-to-object distance to finger angles using per-object polynomial functions, and estimates grasp intent from wrist trajectory, achieving 95.43% grasp success and 94.35% int...

  21. DNF: Unconditional 4D Generation with Dictionary-based Neural Fields

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DNF generates novel 4D deforming shapes by diffusing over per-instance singular-value coefficients of a dictionary built from SVD of pretrained shape and motion neural fields.

  22. GigaHands: A Massive Annotated Dataset of Bimanual Hand Activities

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GigaHands provides 14,000 bimanual hand clips, 84,000 text annotations, and 183 million frames from 51 cameras, outperforming smaller datasets in text-to-motion and captioning tasks.

  23. FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 2D-keypoint-conditioned diffusion model trained on a new 10M-image hand dataset enables controllable hand reposing, appearance transfer, novel view synthesis, and zero-shot hand video generation.

  24. Bimanual Dexterity for Complex Tasks

    cs.RO 2024-11 conditional novelty 6.0 of 10

    BiDex combines Manus mocap gloves for finger tracking with GELLO-style teacher arms for wrist tracking to enable low-cost, portable, bimanual dexterous teleoperation that outperforms VR and SteamVR baselines on most t...

  25. S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A three-stage pipeline generates animatable 3D Gaussian head avatars from one image by diffusion-based splat synthesis, FLAME fitting, and inverse-distance binding with scale adaptation.

  26. Interaction-Aware 4D Gaussian Splatting for Dynamic Hand-Object Interaction Reconstruction

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A 4D Gaussian-splatting method with separate hand/object/background fields, learned importance and radius parameters, and hand-conditioned object deformation improves dynamic hand-object reconstruction from egocentric video.

  27. OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    OpenEgo is a 1,107-hour unified egocentric manipulation dataset with standardized 21-joint hand poses and timestamped action language, plus a small validation showing a language-conditioned policy learns short-horizon...

  28. Half-Physics: Enabling Kinematic 3D Human Model with Physical Interactions

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Half physics converts kinematic SMPL-X poses into velocities that drive a physics engine, preserving the original motion when contact-free and giving physically correct responses when collisions occur.

  29. Hand Shadow Art: A Differentiable Rendering Perspective

    cs.GR 2025-05 conditional novelty 5.0 of 10

    A differentiable rendering framework deforms MANO hand models so their cast shadows match target silhouette images, including interpolation between two targets.

  30. JGHand: Joint-Driven Animatable Hand Avater via 3D Gaussian Splatting

    cs.CV 2025-01 conditional novelty 5.0 of 10

    JGHand is a joint-driven 3D Gaussian Splatting hand model that renders photorealistic hand images in real time, using a zero-error skeleton transform and depth-based shadow simulation to beat prior state-of-the-art on...

  31. Human Grasp Generation for Rigid and Deformable Objects with Decomposed VQ-VAE

    cs.RO 2025-01 conditional novelty 5.0 of 10

    A decomposed VQ-VAE that generates diverse human grasps for rigid and deformable objects, with a Mesh UFormer backbone and normal-vector encoding for deformation simulation.

  32. MultiGO: Towards Multi-level Geometry Learning for Monocular 3D Textured Human Reconstruction

    cs.CV 2024-12 conditional novelty 5.0 of 10

    MultiGO combines skeleton, joint, and wrinkle level improvements on a Gaussian-based 3D human reconstruction model and reports SOTA results on CustomHuman and THuman3.0.

  33. FastGrasp: Efficient Grasp Synthesis with Diffusion

    cs.RO 2024-11 conditional novelty 5.0 of 10

    A one-stage latent diffusion model with an adaptation module generates MANO hand grasping poses from object point clouds faster and with lower penetration than two-stage optimization baselines.

  34. Towards motion from video diffusion models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    MotionDistill optimizes SMPL-X body poses with SDS gradients from video diffusion models, producing plausible animation for common actions but failing on rare ones.

  35. UniHands: Unifying Various Wild-Collected Keypoints for Personalized Hand Reconstruction

    cs.CV 2024-11 reject novelty 5.0 of 10

    A pipeline that reconstructs personalized MANO/NIMBLE hand meshes from arbitrary keypoints and derives a unified 25-joint set via a trained MLP.

  36. Learning Correlation-aware Aleatoric Uncertainty for 3D Hand Pose Estimation

    cs.CV 2025-09 reject novelty 4.0 of 10

    A single linear layer applied to per-joint Gaussian noise gives a correlation-aware covariance for 3D hand pose uncertainty, improving calibration metrics on two benchmarks.

  37. SAT: Supervisor Regularization and Animation Augmentation for Two-process Monocular Texture 3D Human Reconstruction

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A two-stage Gaussian-splatting framework with supervisor feature regularization and online animation augmentation improves monocular textured 3D human reconstruction on CustomHuman and THuman3.0.

  38. Grounding Intelligence in Movement

    cs.AI 2025-07 conditional novelty 4.0 of 10

    Movement should be treated as a first-class AI modeling modality, and a unified, biomechanically grounded movement foundation model built from aggregated data across species and sensors is the proposed path forward.

  39. BG-HOP: A Bimanual Generative Hand-Object Prior

    cs.CV 2025-06 conditional novelty 4.0 of 10

    BG-HOP is a transfer-learned diffusion prior for bimanual hand-object interaction, but its left-hand results are often implausible.

  40. Disentangled Human Body Representation Based on Unsupervised Semantic-Aware Learning

    cs.CV 2025-05 conditional novelty 4.0 of 10

    DHBR learns a disentangled 3D human body representation with a global shape code and per-bone-group pose codes, achieving lower mesh reconstruction error than prior unsupervised baselines.

  41. VM-BHINet:Vision Mamba Bimanual Hand Interaction Network for 3D Interacting Hand Mesh Recovery From a Single RGB Image

    cs.CV 2025-04 reject novelty 4.0 of 10

    VM-BHINet combines a Vision Mamba block with an interaction feature module to recover two interacting hand meshes from one RGB image, reporting 5.44 mm MPVPE and 5.09 mm MPJPE on InterHand2.6M.

  42. Shape Shifters: Does Body Shape Change the Perception of Small-Scale Crowd Motions?

    cs.HC 2024-12 conditional novelty 4.0 of 10

    In virtual crowds of twelve physics-based avatars, varying body shape did not help hide motion clones, but increasing motion variety reduced clone detection.

  43. Text to Image Generation and Editing: A Survey

    cs.CV 2025-05 conditional novelty 3.0 of 10

    A broad survey of text-to-image generation and editing research from 2021 to 2024, organized by architecture and comparison tables.

  44. Exploring Graph Mamba: A Comprehensive Survey on State-Space Models for Graph Learning

    cs.LG 2024-12 conditional novelty 2.0 of 10

    A survey of Graph Mamba, the adaptation of state-space models (Mamba, S4, S6) to graph learning, synthesizing roughly 30 recent papers into a taxonomy of architectures, applications, benchmarks, and open challenges.

Pith tools