REVIEW 45 cited by
Open-TeleVision: Teleoperation with Immersive Active Visual Feedback
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Open-TeleVision: Teleoperation with Immersive Active Visual Feedback
read the original abstract
Teleoperation serves as a powerful method for collecting on-robot data essential for robot learning from demonstrations. The intuitiveness and ease of use of the teleoperation system are crucial for ensuring high-quality, diverse, and scalable data. To achieve this, we propose an immersive teleoperation system Open-TeleVision that allows operators to actively perceive the robot's surroundings in a stereoscopic manner. Additionally, the system mirrors the operator's arm and hand movements on the robot, creating an immersive experience as if the operator's mind is transmitted to a robot embodiment. We validate the effectiveness of our system by collecting data and training imitation learning policies on four long-horizon, precise tasks (Can Sorting, Can Insertion, Folding, and Unloading) for 2 different humanoid robots and deploy them in the real world. The system is open-sourced at: https://robot-tv.github.io/
Forward citations
Cited by 45 Pith papers
-
EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations
EgoEngine transforms egocentric human videos into high-fidelity robot data enabling zero-shot visuomotor dexterous policy learning without real-robot demonstrations.
-
Targeting World Models to Compromise Robot Learning Pipelines
World models introduce a stealthy poisoning vector into robot learning pipelines where malicious prompts or dynamics in teleoperated data activate only during synthetic trajectory generation, enabling backdoors in dow...
-
EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World
A 1,362-hour multi-source egocentric human dataset and multi-lab study show co-training improves robot manipulation, but only when aligned human–robot data anchors transfer from diverse human data.
-
DexDirect: Direct Kinesthetic Arm Guidance for Efficient Dexterous Demonstration Collection
A hybrid kinesthetic-arm-plus-webcam-hand teleoperation interface achieved 17x/3x higher demonstration throughput than vision baselines and trained a 90%-success pick-and-place policy in a ten-person study.
-
ModPack: An Extensible Teleoperation Interface for Bimanual Mobile Manipulation
A modular backpack-based teleoperation interface enables bimanual mobile manipulation with haptic feedback and active perception across multiple robot platforms.
-
AHEAD: Anticipatory Hand-Driven Teleoperation via Human Intent Prediction
A VR teleoperation system predicts the operator's intended grasp object and placement slot from recent hand and head motion, letting the robot start moving before the command is completed.
-
Worlds in One Demo: A Synthetic Data Engine for Learning Open-World Mobile Manipulation
From one real demonstration, WANDA synthesizes diverse mobile-manipulation trajectories, reaching 54.8% average real-world task progress and zero-shot deployment on a morphologically different robot.
-
Robot Trajectron V3: A Probabilistic Shared Control Framework for SE(3) Manipulation
RT-V3 learns a transformer-CVAE prior over multi-modal SE(3) trajectories conditioned on scene geometry and grasps, then continuously fuses it with noisy user twists via Bayesian posterior estimation for shared graspi...
-
TactiDex: A Real-World Tactile-Guided Benchmark for Human-Like Dexterous Manipulation
A tactile-rich HOI dataset plus a tri-component force reward improves contact fidelity and success of human-to-robot dexterous transfer over kinematic imitation alone.
-
Cross-Embodiment Robot Manipulation via a Unified Hand Action Space
UHAS maps hand actions to deformations of a shared unit sphere and recovers joint commands via cascade IK, enabling multi-hand RL, zero-shot transfer, and modest real-world cube reorientation on LEAP and Allegro.
-
Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?
Task-agnostic RL play pretraining on diverse objects yields a reusable dexterous prior that makes sparse-reward assembly learning ~33× more sample-efficient and enables zero-shot sim-to-real transfer on tight insertio...
-
RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning
A wearable interface with a shared dexterous hand module enables retargeting-free teleoperation and matched data collection, yielding policies with 88.75% average success across eight real-robot tasks that generalize ...
-
MonoDuo: Using One Robot Arm to Learn Bimanual Policies
MonoDuo generates synthetic bimanual demonstrations from single-arm teleoperation plus human collaboration to train policies achieving up to 70% zero-shot success on five manipulation tasks, with 65-70% gains from 25-...
-
DexTwist: Dexterous Hand Retargeting for Twist Motion via Mixed Reality-based Teleoperation
DexTwist detects tripod pinches, estimates the intended screw axis and twist magnitude, then applies real-time joint refinement to track turning progress while stabilizing the robot's tripod geometry.
-
DexSynRefine: Synthesizing and Refining Human-Object Interaction Motion for Physically Feasible Dexterous Robot Actions
DexSynRefine synthesizes HOI motions with an extended manifold method, refines them via task-space residual RL, and adapts for sim-to-real transfer, outperforming kinematic retargeting by 50-70 percentage points on fi...
-
DexSynRefine: Synthesizing and Refining Human-Object Interaction Motion for Physically Feasible Dexterous Robot Actions
DexSynRefine couples HOI motion manifold flow primitives with task-space residual RL and proprioceptive adaptation to convert human-object interaction data into executable dexterous robot motions, reporting 50-70 poin...
-
Lucid-XR: An Extended-Reality Data Engine for Robotic Manipulation
Lucid-XR uses XR-headset physics simulation and physics-guided video generation to create synthetic data that trains robot policies transferring zero-shot to unseen real-world manipulation tasks.
-
Learn Weightlessness: Imitate Non-Self-Stabilizing Motions on Humanoid Robot
The Weightlessness Mechanism lets humanoid robots imitate non-self-stabilizing motions by dynamically relaxing specific joints to exploit passive environmental contacts, generalizing from single demonstrations to vari...
-
Learn Weightlessness: Imitate Non-Self-Stabilizing Motions on Humanoid Robot
A weightlessness mechanism enables humanoid robots to dynamically relax joints for stable, contact-rich motions across diverse environments without task-specific tuning.
-
ActiveGlasses: Learning Manipulation with Active Vision from Ego-centric Human Demonstration
ActiveGlasses learns robot manipulation from ego-centric human demos captured with active vision via smart glasses, achieving zero-shot transfer using object-centric point-cloud policies.
-
EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World
EgoVerse releases 1,362 hours of standardized egocentric human data across 1,965 tasks and shows via multi-lab experiments that robot policy performance scales with human data volume when the data aligns with robot ob...
-
IGen: Scalable Data Generation for Robot Learning from Open-World Images
IGen generates realistic visuomotor training data including actions and temporally coherent visuals from unstructured open-world images via 3D reconstruction and VLM reasoning.
-
Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning
Isaac Lab is a unified GPU-native platform combining high-fidelity physics, photorealistic rendering, multi-frequency sensors, domain randomization, and learning pipelines for scalable multi-modal robot policy training.
-
Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation
A 10,300-demonstration, 260-task multimodal humanoid manipulation dataset with baseline policy evaluations and a cloud evaluation platform.
-
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
EgoVLA pretrains VLA models on egocentric human videos, retargets predicted actions to robots via IK, and fine-tunes on few robot demos to improve bimanual manipulation performance on a new simulation benchmark.
-
DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion
DreamPolicy integrates an autoregressive diffusion world model with policy learning to produce a single scalable policy that generalizes to unseen composite terrains for humanoid locomotion.
-
DexWild: Dexterous Human Interactions for In-the-Wild Robot Policies
DexWild co-trains dexterous robot policies on in-the-wild human hand interactions recorded with a low-cost system and limited robot data, achieving 68.5% success in unseen environments and 5.8x better cross-embodiment...
-
FAST: Efficient Action Tokenization for Vision-Language-Action Models
FAST applies discrete cosine transform to robot action sequences for efficient tokenization, enabling autoregressive VLAs to succeed on high-frequency dexterous tasks and scale to 10k hours of data while matching diff...
-
Towards Miniature Humanoid Tele-Loco-Manipulation Using Virtual Reality and Reinforcement Learning
A VR-teleoperated, reinforcement-learning-balanced control stack lets a miniature ROBOTIS OP3 humanoid walk and manipulate objects simultaneously.
-
MEVION: Low-Cost Open-Source Data Collection System for Powerful and High-Speed Dual-Arm Manipulation
MEVION is an open-source dual-arm teleoperation platform with 60 Nm joint torque, built from e-commerce parts for about $14,000 per four-arm system, enabling heavier, faster manipulation data collection.
-
DexTele: A Dual-Arm Dexterous Teleoperation System Based on Motion Retargeting and Adaptive Force Control
A dual-arm teleoperation system combines a graph-based motion retargeting network with VLM-informed MPC force control to achieve cross-platform motion mapping and adaptive grasping across multiple robots and objects.
-
HEFT: Heavy-Payload Full-size Humanoid Teleoperation with Privileged Motion Guidance and Windowed Payload Curriculum
HEFT enables tracking of human motions including locomotion and squats on a 175cm 65kg humanoid under up to 24kg payloads by combining Privileged Motion Guidance from noisy VR data with a Windowed Payload Curriculum.
-
WARP: Whole-Body Retargeting for Learning from Offline Human Demonstrations
WARP is an offline retargeting method using a SEW geometric solver to produce consistent whole-body robot trajectories from human demonstrations for zero-shot mobile manipulation.
-
Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?
Play2Perfect uses task-agnostic RL play pretraining on diverse objects to build reusable manipulation priors, then fine-tunes for assembly, yielding 33x sample efficiency gains and 60% success on 0.5mm-clearance inser...
-
ConTrack: Constrained Hand Motion Tracking with Adaptive Trade-off Control
ConTrack introduces a constrained RL method with online dual-variable adaptation and adaptive resets for improved long-horizon hand tracking in simulation and on real robots.
-
A study on a Real-Time VR-Based Teleoperation Framework for Manipulator in Dynamic Environment
A VR teleoperation framework integrates GPU-accelerated inverse kinematics and trajectory optimization to generate collision-aware joint commands for a 7-DoF manipulator in real time across obstacle-free, static, and ...
-
Switch: Learning Agile Skills Switching for Humanoid Robots
Switch enables humanoid robots to perform agile, seamless transitions between locomotion skills via a kinematic skill graph, DRL tracking policy, and real-time graph-search scheduler.
-
Learning Versatile Humanoid Manipulation with Touch Dreaming
HTD, a multimodal transformer policy trained with behavioral cloning and touch dreaming to predict future tactile latents, achieves a 90.9% relative success rate improvement over baselines on five real-world contact-r...
-
A Multi-View 3D Telepresence System for XR Robot Teleoperation
A multi-view point cloud VR system with wrist RGB detail outperforms RGB streams and stereo views in robot teleoperation tasks per a 31-participant user study.
-
Low-Cost Teleoperation Extension for Mobile Manipulators
An open-source teleoperation framework enables intuitive whole-body control of mobile manipulators using commodity smartphone, leader arms, and foot pedals instead of costly VR equipment.
-
Choose What to Manipulate: Revealing Data Scaling Laws in Bounding-Box Guided Policies for Semantic Manipulation
A bounding-box-conditioned diffusion policy shows a power-law improvement with the number of object classes in training data, reaching about 85% success on four semantic manipulation tasks.
-
A Multimodal Data Collection Framework for Dialogue-Driven Assistive Robotics to Clarify Ambiguities: A Wizard-of-Oz Pilot Study
A two-room Wizard-of-Oz pilot collected 53 multimodal trials from five users to capture dialogue ambiguities for training ambiguity-aware assistive robot controllers.
-
General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling
GAM framework uses arc-length parameterization for temporal invariance and schema-affine factorization for geometric invariance to build a covariant action manifold integrated into VLA models for improved generalizati...
-
Data Pyramid for Embodied Manipulation
Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.
-
Immersive Social Interaction with VR and LLM-Assisted Humanoids
Novice operators achieved 80% success on object manipulation and 70% on social cube-passing using a VR-and-LLM-assisted humanoid teleoperation framework on a Unitree H1.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.