Pith. sign in

REVIEW 1 cited by

Constrained Skill Discovery: Quadruped Locomotion with Unsupervised Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.07877 v1 pith:QK5IEWP5 submitted 2024-10-10 cs.RO

classification cs.RO
keywords discoveryskilllearningrobotunsupervisedbehaviorsconstrainedlatent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Representation learning and unsupervised skill discovery can allow robots to acquire diverse and reusable behaviors without the need for task-specific rewards. In this work, we use unsupervised reinforcement learning to learn a latent representation by maximizing the mutual information between skills and states subject to a distance constraint. Our method improves upon prior constrained skill discovery methods by replacing the latent transition maximization with a norm-matching objective. This not only results in a much a richer state space coverage compared to baseline methods, but allows the robot to learn more stable and easily controllable locomotive behaviors. We successfully deploy the learned policy on a real ANYmal quadruped robot and demonstrate that the robot can accurately reach arbitrary points of the Cartesian state space in a zero-shot manner, using only an intrinsic skill discovery and standard regularization rewards.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Divide, Discover, Deploy: Factorized Skill Learning with Symmetry and Style Priors

    cs.RO 2025-08 conditional novelty 6.0 of 10

    A factorized USD framework that mixes METRA and DIAYN per state factor, adds symmetry and style priors, and achieves sim-to-real transfer on a quadruped.

Pith tools