REVIEW 40 cited by
Embodied Hands: Modeling and Capturing Hands and Bodies Together
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Embodied Hands: Modeling and Capturing Hands and Bodies Together
read the original abstract
Humans move their hands and bodies together to communicate and solve tasks. Capturing and replicating such coordinated activity is critical for virtual characters that behave realistically. Surprisingly, most methods treat the 3D modeling and tracking of bodies and hands separately. Here we formulate a model of hands and bodies interacting together and fit it to full-body 4D sequences. When scanning or capturing the full body in 3D, hands are small and often partially occluded, making their shape and pose hard to recover. To cope with low-resolution, occlusion, and noise, we develop a new model called MANO (hand Model with Articulated and Non-rigid defOrmations). MANO is learned from around 1000 high-resolution 3D scans of hands of 31 subjects in a wide variety of hand poses. The model is realistic, low-dimensional, captures non-rigid shape changes with pose, is compatible with standard graphics packages, and can fit any human hand. MANO provides a compact mapping from hand poses to pose blend shape corrections and a linear manifold of pose synergies. We attach MANO to a standard parameterized 3D body shape model (SMPL), resulting in a fully articulated body and hand model (SMPL+H). We illustrate SMPL+H by fitting complex, natural, activities of subjects captured with a 4D scanner. The fitting is fully automatic and results in full body models that move naturally with detailed hand motions and a realism not seen before in full body performance capture. The models and data are freely available for research purposes in our website (http://mano.is.tue.mpg.de).
Forward citations
Cited by 40 Pith papers
-
JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation
JointHOI jointly generates hand-object motion and distance-based contact maps in one diffusion stage to improve temporal stability and physical plausibility over prior multi-stage HOI methods.
-
Policy-as-Data: Learning Generalizable HOI Diffusion Models from Simulated Physics
A framework called Policy-as-Data generates task-oriented synthetic HOI data via RL policies in physics simulators, retargets it, and trains diffusion models that generalize to unseen objects and long horizons.
-
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
DreamDojo is a foundation world model pretrained on the largest human video dataset to date that uses continuous latent actions to transfer interaction knowledge and achieves controllable physics simulation after robo...
-
AGILE: Hand-Object Interaction Reconstruction from Video via Agentic Generation
AGILE generates complete object meshes via VLM-guided synthesis and tracks poses with anchor-and-track plus contact-aware optimization to achieve robust hand-object reconstruction from video.
-
CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos
CoMoVi co-generates 3D human motions and 2D videos synchronously in a single diffusion denoising loop using 3D-to-2D projection and dual-branch diffusion with 3D-2D cross attentions.
-
Large Video Planner Enables Generalizable Robot Control
A video foundation model trained on human demonstrations generates zero-shot plans that convert to executable robot actions on novel scenes and tasks.
-
SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting
SFHand presents the first streaming language-guided autoregressive framework for 3D hand forecasting, achieving up to 35.8% gains over prior methods and 13.4% better downstream embodied task performance.
-
Rodrigues Network for Learning Robot Actions
Proposes Rodrigues Network using a learnable Neural Rodrigues Operator to add kinematic inductive biases for improved robot action learning and prediction.
-
InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation
A first large multimodal 4D human–dog interaction dataset (6.8M frames) plus an autoregressive model that generates dog motion from human body/hand gestures and audio.
-
BODIESReg: An Open-Source Pipeline for Registering 3D Body Scans Using Pose-Aligned Initialization
BODIESReg automatically registers 3D body scans to SMPL-family models using pose-aligned initialization, achieving 82.9% success on CHI3D and 100% on MorphoMotion with mean surface-fit error below 10mm.
-
TactiDex: A Real-World Tactile-Guided Benchmark for Human-Like Dexterous Manipulation
A tactile-rich HOI dataset plus a tri-component force reward improves contact fidelity and success of human-to-robot dexterous transfer over kinematic imitation alone.
-
Human-Centric Transferable Tactile Pre-Training for Dexterous Robotic Manipulation
Introduces H-Tac human tactile-action dataset and TTP pre-training that unifies spaces and predicts future tactile signals to improve robotic dexterous manipulation transfer.
-
PoseShield: Neural Collision Fields for Human Self-Collision Resolution
PoseShield learns a neural collision field directly in SMPL pose space with Eikonal regularization to correct self-collisions post-hoc in human pose estimation and motion generation, achieving 95.8% success on a new b...
-
Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots
A relative wrist translation bridging action with a vision-language-action model using interleaved tokens and attention masking transfers human manipulation skills to robots more effectively than 6DoF actions.
-
Generative Relightable Avatars
GRA combines UV-space material optimization and physics rendering with feed-forward texture refinement and a fine-tuned video-to-video diffusion model to achieve controllable, high-detail relighting of full-body avatars.
-
Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
Qwen-RobotManip applies unified alignment across representation, motion, and behavior to enable large-scale training on heterogeneous manipulation data, yielding emergent generalization on out-of-distribution robotic ...
-
Hand-centric Human-to-Robot Trajectory Transfer from Video Demonstrations via Open-World Contact Localization
HOWTransfer recovers 3D hand motion from video, localizes contact intervals via hand-object cues, generates multi-modal grasp hypotheses, and edits trajectories to produce diverse robot-executable motions achieving 86...
-
Dexterous Point Policy: Learning Point-based Dexterous Hand Policies from Human Demonstrations
Dexterous Point Policy learns dexterous hand policies from human videos using 3D keypoints of hands and objects, achieving 75% success on real-robot tasks compared to 1% for a VLA baseline.
-
EgoPressDiff: Multimodal Video Diffusion for Egocentric UV-Domain Hand-Pressure Estimation
EgoPressDiff is a multimodal video diffusion framework that generates UV-pressure maps from egocentric visual inputs using PoseNet, Vertex Encoder, and a Distribution-Calibrated Spatial Layer, reporting over 34% relat...
-
HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model
HARP aligns human-robot visual and latent action representations via paired bridges and unpaired dynamics supervision to boost VLA policy performance on manipulation tasks.
-
UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation
UST-Hand is a self-supervised 3D hand pose estimation method using conditional normalizing flows for uncertainty-aware hypothesis sampling and probabilistic point cloud interactions to achieve up to 37.8% better MPVPE...
-
TouchMap-OR: Multi-View 3D Mapping of Hand-Surface Contacts
TouchMap-OR reconstructs multi-person 3D hand trajectories from RGB-D cameras, builds a semantic 3D model of the operating room, and infers contact events from proximity, achieving 0.75 binary contact F1 on three real...
-
GeoHand: Unlocking Prior Geometry Knowledge for Monocular 3D Hand Reconstruction
GeoHand adapts priors from a general-scene geometry estimator via a GeoAdapter, gated fusion, and keypoint-queried refiner to reach SOTA monocular 3D hand reconstruction on FreiHAND, DexYCB, and HO3Dv3 under heavy occlusion.
-
EgoExo-WM: Unlocking Exo Video for Ego World Models
Method converts exocentric videos to egocentric format via body-pose extraction and kinematics to improve egocentric world-model prediction and planning.
-
EgoExo-WM: Unlocking Exo Video for Ego World Models
Converting exocentric video to egocentric format via body-pose extraction and kinematics prior enables training of action-conditioned egocentric world models that improve prediction quality and goal-directed planning.
-
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation
MoT-HRA learns embodiment-agnostic human-intention priors from the HA-2.2M dataset of 2.2M human video episodes through a three-expert hierarchy to improve robotic motion plausibility and robustness under distribution shift.
-
Grasp in Gaussians: Fast Monocular Reconstruction of Dynamic Hand-Object Interactions
GraG reconstructs dynamic 3D hand-object interactions from monocular video 6.4x faster than prior work by using compact Sum-of-Gaussians tracking initialized from large models and refined with 2D losses.
-
HandDreamer: Zero-Shot Text to 3D Hand Model Generation using Corrective Hand Shape Guidance
HandDreamer is the first zero-shot text-to-3D method for hands that uses MANO initialization, skeleton-guided diffusion, and corrective shape guidance to produce view-consistent models.
-
Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints
Sparse 3D hand joints plus an occlusion-aware control module produce higher-fidelity, 3D-consistent egocentric hand-object videos than dense-2D or implicit-pose baselines.
-
Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints
A new occlusion-aware control module generates high-fidelity egocentric videos from sparse 3D hand joints, supported by a million-clip dataset and cross-embodiment benchmark.
-
Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views
Uni-Hand forecasts 2D/3D hand waypoints, head motion, and contact states in egocentric views using vision-language fusion and dual-branch diffusion, with new benchmarks for downstream robotics and action tasks.
-
Physically Plausible Human-Object Rendering from Sparse Views via 3D Gaussian Splatting
HOGS renders physically plausible human-object interactions from sparse views by optimizing dynamic 3D Gaussians with contact/separation losses guided by pre-trained pose refiner and contact predictor modules, claimin...
-
S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image
A three-stage pipeline generates animatable 3D Gaussian head avatars from one image by diffusion-based splat synthesis, FLAME fitting, and inverse-distance binding with scale adaptation.
-
PoseShield: Neural Collision Fields for Human Self-Collision Resolution
Neural collision field in SMPL pose space with Eikonal regularization resolves self-collisions at 95.8% success rate on a new benchmark.
-
Generative Relightable Avatars
GRA combines a microfacet-relit 3D avatar with a fine-tuned video diffusion model to generate photorealistic, freely viewable relit videos of a person under arbitrary environment lighting.
-
VEPHand: View-Efficient Photometric Hand Performance Capture at Scale
End-to-end neural pipeline extracts hand geometry from unmasked limited-view images and registers it to a personalized tetrahedral model via volumetric offsets, achieving SOTA on over 12,000 sequences.
-
Generator-Refiner-Examiner: A Tri-Module Data Augmentation Framework for 3D Human Avatar Learning from Monocular Videos
TrioMan is a tri-module data augmentation framework using a Generator for pose/camera perturbations, a Refiner with one-step diffusion, and an Examiner with dual-branch attention to improve 3D avatar learning from mon...
-
Uni-HOI:A Unified framework for Learning the Joint distribution of Text and Human-Object Interaction
Uni-HOI learns the joint distribution of text, human motion, and object motion using LLMs and VQ-VAEs in a two-stage training process for multiple HOI tasks.
-
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation
MoT-HRA learns embodiment-agnostic human-intention priors from a curated 2.2M-episode human video dataset via a three-expert hierarchical vision-language-action model to improve robotic manipulation under distribution shift.
-
EaDex: A Cross-Embodiment Dexterous Manipulation Framework from Low-Cost Demonstrations
EaDex combines single-camera RGB-D capture, MANO retargeting, and dynamic demonstration annealing to achieve 55.3% relative improvement over baseline on nine cross-embodiment dexterous object-opening tasks across three hands.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.