Pith. sign in

REVIEW 2 cited by

TeleMoMa: A Modular and Versatile Teleoperation System for Mobile Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.07869 v2 pith:MGWJ7XOU submitted 2024-03-12 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords telemomamobilemanipulationdemonstrationsteleoperationdemonstrateinterfaceswhole-body
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A critical bottleneck limiting imitation learning in robotics is the lack of data. This problem is more severe in mobile manipulation, where collecting demonstrations is harder than in stationary manipulation due to the lack of available and easy-to-use teleoperation interfaces. In this work, we demonstrate TeleMoMa, a general and modular interface for whole-body teleoperation of mobile manipulators. TeleMoMa unifies multiple human interfaces including RGB and depth cameras, virtual reality controllers, keyboard, joysticks, etc., and any combination thereof. In its more accessible version, TeleMoMa works using simply vision (e.g., an RGB-D camera), lowering the entry bar for humans to provide mobile manipulation demonstrations. We demonstrate the versatility of TeleMoMa by teleoperating several existing mobile manipulators - PAL Tiago++, Toyota HSR, and Fetch - in simulation and the real world. We demonstrate the quality of the demonstrations collected with TeleMoMa by training imitation learning policies for mobile manipulation tasks involving synchronized whole-body motion. Finally, we also show that TeleMoMa's teleoperation channel enables teleoperation on site, looking at the robot, or remote, sending commands and observations through a computer network, and perform user studies to evaluate how easy it is for novice users to learn to collect demonstrations with different combinations of human interfaces enabled by our system. We hope TeleMoMa becomes a helpful tool for the community enabling researchers to collect whole-body mobile manipulation demonstrations. For more information and video results, https://robin-lab.cs.utexas.edu/telemoma-web.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Panorama-Aware VLA for Mobile Manipulation with Whole-Body Teleoperation

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Adding a panoramic camera feed to a vision-language-action policy raises end-to-end success on four real-world mobile two-arm tasks from 30% to 73%.

  2. ModPack: An Extensible Teleoperation Interface for Bimanual Mobile Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A modular backpack-based teleoperation interface enables bimanual mobile manipulation with haptic feedback and active perception across multiple robot platforms.

Pith tools