Pith. sign in

REVIEW 6 cited by

SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.04147 v5 pith:PT3QB7WY submitted 2025-06-04 cs.RO cs.AIcs.LG

SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training

classification cs.RO cs.AIcs.LG
keywords slacactionlatentlearningreal-worldspacetasksautonomously
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Building capable household and industrial robots requires mastering the control of versatile, high-degree-of-freedom (DoF) systems such as mobile manipulators. While reinforcement learning (RL) holds promise for autonomously acquiring robot control policies, scaling it to high-DoF embodiments remains challenging. Direct RL in the real world demands both safe exploration and high sample efficiency, which are difficult to achieve in practice. Sim-to-real RL, on the other hand, is often brittle due to the reality gap. This paper introduces SLAC, a method that renders real-world RL feasible for complex embodiments by leveraging a low-fidelity simulator to pretrain a task-agnostic latent action space. SLAC trains this latent action space via a customized unsupervised skill discovery method designed to promote temporal abstraction, disentanglement, and safety, thereby facilitating efficient downstream learning. Once a latent action space is learned, SLAC uses it as the action interface for a novel off-policy RL algorithm to autonomously learn downstream tasks through real-world interactions. We evaluate SLAC against existing methods on a suite of bimanual mobile manipulation tasks, where it achieves state-of-the-art performance. Notably, SLAC learns contact-rich whole-body tasks in under an hour of real-world interactions, without relying on any demonstrations or hand-crafted behavior priors. More information and robot videos at robo-rl.github.io

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents

    cs.RO 2026-06 unverdicted novelty 6.0

    A systematic study of hierarchical VLA agents identifies design principles that improve robot manipulation performance over flat and naive hierarchical baselines in simulation and real-world experiments.

  2. VOFA: Visual Object Goal Pushing with Force-Adaptive Control for Humanoids

    cs.RO 2026-05 unverdicted novelty 6.0

    VOFA combines a high-level visuomotor policy with a low-level force-adaptive controller to let humanoids push objects up to 17 kg to arbitrary goals using only noisy onboard vision, achieving over 80% real-world success.

  3. VOFA: Visual Object Goal Pushing with Force-Adaptive Control for Humanoids

    cs.RO 2026-05 conditional novelty 6.0

    VOFA combines a depth-image visuomotor policy with a force-adaptive whole-body controller to push objects of unknown mass to arbitrary goals on a humanoid.

  4. Factored Latent Action World Models

    cs.LG 2026-02 conditional novelty 6.0

    FLAM splits a scene into separate factors, each with its own latent action, and reports better video prediction and downstream policy learning than monolithic latent-action models.

  5. Learning Agile Striker Skills for Humanoid Soccer Robots from Noisy Sensory Input

    cs.RO 2025-12 conditional novelty 5.0

    A four-stage RL system with teacher-student distillation and online constrained adaptation enables humanoid robots to achieve robust ball-kicking accuracy under noisy perception in simulation and on physical hardware.

  6. IDEA: Insensitive to Dynamics Mismatch via Effect Alignment for Sim-to-Real Transfer in Multi-Agent Control

    cs.RO 2026-06 unverdicted novelty 3.0

    IDEA elevates multi-agent policies to semantic actions with effect alignment and synchronization for improved sim-to-real robustness on navigation tasks.