REVIEW 29 cited by
Generalizable Humanoid Manipulation with 3D Diffusion Policies
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Humanoid robots capable of autonomous operation in diverse environments have long been a goal for roboticists. However, autonomous manipulation by humanoid robots has largely been restricted to one specific scene, primarily due to the difficulty of acquiring generalizable skills and the expensiveness of in-the-wild humanoid robot data. In this work, we build a real-world robotic system to address this challenging problem. Our system is mainly an integration of 1) a whole-upper-body robotic teleoperation system to acquire human-like robot data, 2) a 25-DoF humanoid robot platform with a height-adjustable cart and a 3D LiDAR sensor, and 3) an improved 3D Diffusion Policy learning algorithm for humanoid robots to learn from noisy human data. We run more than 2000 episodes of policy rollouts on the real robot for rigorous policy evaluation. Empowered by this system, we show that using only data collected in one single scene and with only onboard computing, a full-sized humanoid robot can autonomously perform skills in diverse real-world scenarios. Videos are available at https://humanoid-manipulation.github.io .
Forward citations
Cited by 29 Pith papers
-
Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
A learned plug-in that denoises consumer depth cameras to simulation-like metric depth enables zero-shot sim-to-real transfer of depth-only manipulation policies trained on raw simulated depth.
-
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
A vision-language-action model that fuses stereo-derived geometric features with semantic features improves real-world grasping success and camera-pose robustness over single-view baselines.
-
Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation
A 10,300-demonstration, 260-task multimodal humanoid manipulation dataset with baseline policy evaluations and a cloud evaluation platform.
-
UMI-on-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies
Embodiment-Aware Diffusion Policy steers a UMI-trained diffusion policy with controller tracking-cost gradients at inference time, improving aerial manipulation success in simulation and real flights.
-
Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
Gaze-guided foveated patch tokenization reduces ViT tokens by 94%, accelerates training 7x and inference 3x, and improves robustness to distractors in bimanual manipulation policies.
-
TypeTele: Releasing Dexterity in Teleoperation by Dexterous Manipulation Types
A type-guided teleoperation system that selects predefined dexterous hand poses with a language model outperforms retargeting-based teleoperation on nine real-world tasks and improves imitation learning success.
-
Vision in Action: Learning Active Perception from Human Demonstrations
ViA trains bimanual manipulation policies from human demonstrations that include active head-camera movement, using a 6-DoF robot neck and a VR interface with point-cloud rendering, reporting large gains on three occl...
-
Versatile Loco-Manipulation through Flexible Interlimb Coordination
ReLIC lets a robot dog dynamically reassign its legs between walking and manipulating, achieving 78.9% average success across 12 real-world loco-manipulation tasks.
-
DemoSpeedup: Accelerating Visuomotor Policies via Entropy-Guided Demonstration Acceleration
DemoSpeedup accelerates visuomotor policies by downsampling high-entropy segments of demonstrations, achieving roughly 2x faster execution with maintained or improved success rates.
-
Object-Focus Actor for Data-efficient Robot Generalization Dexterous Manipulation
Object-Focus Actor makes dexterous manipulation policies generalize to new object positions and backgrounds by focusing on the hand-object region and using relative poses and actions, needing only 10 to 30 demonstrations.
-
Latent Theory of Mind: A Decentralized Diffusion Architecture for Cooperative Manipulation
A decentralized diffusion policy for two robot arms that aligns a learned consensus embedding across agents and uses theory-of-mind prediction to keep that embedding informative.
-
H$^3$DP: Triply-Hierarchical Diffusion Policy for Visuomotor Learning
H3DP couples depth-layered, multi-scale visual features to coarse-to-fine denoising stages, reporting a +27.5% relative success-rate improvement over DP3 across 44 simulation tasks.
-
Morphologically Symmetric Reinforcement Learning for Ambidextrous Bimanual Manipulation
SYMDEX decomposes bimanual tasks into per-hand equivariant policies and distills them into an ambidextrous policy, achieving strong results on six simulated tasks and two real-world deployments.
-
TWIST: Teleoperated Whole-Body Imitation System
Human MoCap drives a Unitree G1 humanoid in real time through a single teacher-student RL+BC controller that transfers zero-shot from simulation.
-
Physically Consistent Humanoid Loco-Manipulation using Latent Diffusion Models
A latent-diffusion-image-to-keyframe pipeline enables whole-body trajectory optimization to solve long-horizon humanoid loco-manipulation tasks in simulation, outperforming contact-only guidance.
-
CordViP: Correspondence-based Visuomotor Policy for Dexterous Manipulation in Real-World
CordViP achieves strong real-world dexterous manipulation by feeding a diffusion policy with pose-tracked 3D object models and hand point clouds, pretrained on contact maps and arm-hand coordination.
-
Habitizing Diffusion Planning for Efficient and Effective Decision Making
A variational-Bayes distillation framework (Habi) turns slow diffusion planners into fast feedforward policies that match their performance at orders-of-magnitude higher decision frequency on D4RL benchmarks.
-
Mimicking-Bench: A Benchmark for Generalizable Humanoid-Scene Interaction Learning via Human Mimicking
Mimicking-Bench provides six humanoid-scene interaction tasks with 23K human motion references and a retarget-track-imitate pipeline that beats data-free RL on average success.
-
Learning from Massive Human Videos for Universal Humanoid Pose Control
Humanoid-X contributes 163,800 text-annotated motion clips retargeted from human videos into humanoid robot poses, and UH-1 is an autoregressive transformer that maps text instructions to humanoid actions.
-
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
A robot policy generates its own language reasoning before acting and injects it into a diffusion action decoder, outperforming several VLA baselines on real-robot manipulation.
-
ARMOR: Egocentric Perception for Humanoid Robot Collision Avoidance and Motion Planning
Distributed arm-mounted ToF depth sensors plus a transformer imitation policy reduce collisions and improve success in humanoid collision avoidance compared to external cameras and cuRobo.
-
SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation
SkillMemo couples MoE-based skill discovery with episodic memory retrieval and reports consistent success-rate gains on diffusion and VLA policies for simulated and real manipulation tasks.
-
4D Visual Pre-training for Robot Learning
A next-frame point-cloud diffusion pre-training method (FVP) improves DP3 and RDT-1B manipulation success rates on the paper's own tasks.
-
Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach
The claimed result is that source-component-shift adaptation splits cleanly into offline component learning via EM and online mixing-weight updates, cutting cumulative test loss by up to 67.4%.
-
Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots
A humanoid-specific multimodal occupancy perception system with a new dataset, sensor layout, and a fusion network that claims state-of-the-art results on its own benchmark.
-
MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation
MoTo turns existing fixed-base manipulation models into mobile manipulators by using VLM-picked contact keypoints and trajectory optimization to find docking points, with no training of MoTo itself.
-
Leveraging OS-Level Primitives for Robotic Action Management
Applying OS-style exception handling, context caching, and replay to robotic action slices raises success rates 7x to 24x and cuts execution steps up to 74% for repetitive manipulation tasks, without retraining the VLA model.
-
A Survey: Learning Embodied Intelligence from Physical Simulators and World Models
Embodied intelligence learning is reviewed through the complementary lenses of physical simulators and world models, with a proposed IR-L0 to IR-L4 robot capability taxonomy.
-
A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI
A review of navigation and manipulation simulators, datasets, and methods, framed around the sim-to-real gap.
Discussion (0). Continue with ORCID to comment.