Pith. sign in

REVIEW 21 cited by

UMI on Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole-body Controllers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.10353 v1 pith:WC62WIAC submitted 2024-07-14 cs.RO

UMI on Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole-body Controllers

classification cs.RO
keywords manipulationdatarobotsimulationumi-on-legswhole-bodycontrollerdemonstrate
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce UMI-on-Legs, a new framework that combines real-world and simulation data for quadruped manipulation systems. We scale task-centric data collection in the real world using a hand-held gripper (UMI), providing a cheap way to demonstrate task-relevant manipulation skills without a robot. Simultaneously, we scale robot-centric data in simulation by training whole-body controller for task-tracking without task simulation setups. The interface between these two policies is end-effector trajectories in the task frame, inferred by the manipulation policy and passed to the whole-body controller for tracking. We evaluate UMI-on-Legs on prehensile, non-prehensile, and dynamic manipulation tasks, and report over 70% success rate on all tasks. Lastly, we demonstrate the zero-shot cross-embodiment deployment of a pre-trained manipulation policy checkpoint from prior work, originally intended for a fixed-base robot arm, on our quadruped system. We believe this framework provides a scalable path towards learning expressive manipulation skills on dynamic robot embodiments. Please checkout our website for robot videos, code, and data: https://umi-on-legs.github.io

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Being-H0.7: A Latent World-Action Model from Egocentric Videos

    cs.RO 2026-04 unverdicted novelty 7.0

    Being-H0.7 adds future-aware latent reasoning to direct VLA policies via dual-branch alignment on latent queries, matching world-model benefits at VLA efficiency.

  2. HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

    cs.RO 2026-07 conditional novelty 6.0

    Robot-free HiFi-UMI demonstrations can replace teleoperated real-robot data in post-training: three policy backbones matched in-domain teleoperation within 3.1 percentage points, including 85% success on a precision i...

  3. FT-WBC: Learning Fault-Tolerant Whole-Body Control for Legged Loco-Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0

    FT-WBC introduces a decoupled policy architecture with a Fault Estimator and Posture Adaptation Module that converts unstable arm-driven posture requests into safe base commands under actuator failures in legged manipulators.

  4. Grounding Generative Policies in Physics: Optimization-Guided Diffusion for Robot Control

    cs.RO 2026-06 unverdicted novelty 6.0

    Optimization-guided diffusion replaces sampling perturbations with constrained corrections to enforce physical feasibility in generative robot policies at inference time.

  5. EmbodiSteer: Steering Embodiment-Agnostic Visuomotor Policies with Joint-Space Guidance for Zero-Shot Cross-Embodiment Deployment

    cs.RO 2026-06 unverdicted novelty 6.0

    EmbodiSteer steers embodiment-agnostic Cartesian diffusion policies into joint space with Jacobian-based collision guidance after each denoising step for zero-shot cross-embodiment deployment.

  6. HCLM: A Hierarchical Framework for Cooperative Loco-Manipulation with Dual Quadrupeds

    cs.RO 2026-05 unverdicted novelty 6.0

    HCLM presents a hierarchical architecture that uses an SE(3)-invariant diffusion policy for coordination and a hybrid whole-body controller with MPC and admittance control for safe closed-chain loco-manipulation on du...

  7. BifrostUMI: Bridging Robot-Free Demonstrations and Humanoid Whole-Body Manipulation

    cs.RO 2026-05 unverdicted novelty 6.0

    BifrostUMI enables robot-free human demonstration capture via VR and wrist cameras to train visuomotor policies that predict keypoint trajectories for transfer to humanoid whole-body control through retargeting.

  8. Learning Tactile-Aware Quadrupedal Loco-Manipulation Policies

    cs.RO 2026-04 unverdicted novelty 6.0

    A hierarchical tactile-aware policy trained from human demos and sim RL improves real quadrupedal loco-manipulation by 28.54% on average over vision-only and visuotactile baselines.

  9. Learning Tactile-Aware Quadrupedal Loco-Manipulation Policies

    cs.RO 2026-04 unverdicted novelty 6.0

    A tactile-aware hierarchical policy for quadrupedal loco-manipulation improves real-world contact-rich task performance by 28.54% over vision-only and visuotactile baselines.

  10. UMI-3D: Extending Universal Manipulation Interface from Vision-Limited to 3D Spatial Perception

    cs.RO 2026-04 unverdicted novelty 6.0

    UMI-3D integrates LiDAR into the UMI hardware for robust multimodal 3D perception in manipulation demonstrations, yielding higher policy success rates and enabling previously infeasible tasks like deformable object handling.

  11. XRZero-G0: Pushing the Frontier of Dexterous Robotic Manipulation with Interfaces, Quality and Ratios

    cs.RO 2026-04 unverdicted novelty 6.0

    XRZero-G0 enables 2000-hour robot-free datasets that, when mixed 10:1 with real-robot data, match full real-robot performance at 1/20th the cost and support zero-shot transfer.

  12. TAMEn: Tactile-Aware Manipulation Engine for Closed-Loop Data Collection in Contact-Rich Tasks

    cs.RO 2026-04 unverdicted novelty 6.0

    TAMEn supplies a cross-morphology wearable interface and pyramid-structured visuo-tactile data regime that raises bimanual manipulation success rates from 34% to 75% via closed-loop collection.

  13. Humanoid Whole-Body Badminton via Multi-Stage Reinforcement Learning

    cs.RO 2025-11 unverdicted novelty 6.0

    A multi-stage RL curriculum produces a unified whole-body controller enabling humanoid robots to sustain badminton rallies in simulation and return shuttles at up to 19.1 m/s in real hardware, with both EKF-based and ...

  14. R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation

    cs.RO 2025-10 unverdicted novelty 6.0

    R2RGen introduces a simulator-free three-stage pipeline that parses, augments, and post-processes real pointcloud observation-action pairs to improve spatial generalization in robotic manipulation policies.

  15. HeLoM: Hierarchical Learning for Whole-Body Loco-Manipulation by a Hexapod Robot

    cs.RO 2025-09 conditional novelty 6.0

    A hexapod pushes boxes with unknown mass, size, and friction to target poses by coordinating front-leg contact with hind-leg balance via a hierarchical learned controller.

  16. DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

    cs.RO 2025-02 unverdicted novelty 6.0

    DexVLA combines a scaled diffusion action expert with embodiment curriculum learning to achieve better generalization and performance than prior VLA models on diverse robot hardware and long-horizon tasks.

  17. FT-WBC: Learning Fault-Tolerant Whole-Body Control for Legged Loco-Manipulation

    cs.RO 2026-06 unverdicted novelty 5.0

    FT-WBC is a decoupled-policy framework that uses fault estimation and posture adaptation to synthesize compensatory gaits and preserve arm workspace in legged manipulators under actuator failures.

  18. Learning Tactile-Aware Quadrupedal Loco-Manipulation Policies

    cs.RO 2026-04 unverdicted novelty 5.0

    A hierarchical tactile-aware policy combines human-demonstration training for contact cue prediction with sim-to-real reinforcement learning to improve quadrupedal loco-manipulation performance by 28.54% over vision b...

  19. Choose What to Manipulate: Revealing Data Scaling Laws in Bounding-Box Guided Policies for Semantic Manipulation

    cs.RO 2026-02 conditional novelty 5.0

    A bounding-box-conditioned diffusion policy shows a power-law improvement with the number of object classes in training data, reaching about 85% success on four semantic manipulation tasks.

  20. World Action Models: The Next Frontier in Embodied AI

    cs.RO 2026-05 unverdicted novelty 4.0

    The paper introduces World Action Models as a new paradigm unifying predictive world modeling with action generation in embodied foundation models and provides a taxonomy of existing approaches.

  21. Data Pyramid for Embodied Manipulation

    cs.RO 2026-07 conditional novelty 3.0

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.