Pith. sign in

REVIEW 1 cited by

Learning Domain Randomization Distributions for Training Robust Locomotion Policies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.00410 v2 pith:GSOR43VP submitted 2019-06-02 cs.LG cs.ROstat.ML

classification cs.LGcs.ROstat.ML
keywords randomizationdomaindistributiontargettrainingpoliciessystemlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Domain randomization (DR) is a successful technique for learning robust policies for robot systems, when the dynamics of the target robot system are unknown. The success of policies trained with domain randomization however, is highly dependent on the correct selection of the randomization distribution. The majority of success stories typically use real world data in order to carefully select the DR distribution, or incorporate real world trajectories to better estimate appropriate randomization distributions. In this paper, we consider the problem of finding good domain randomization parameters for simulation, without prior access to data from the target system. We explore the use of gradient-based search methods to learn a domain randomization with the following properties: 1) The trained policy should be successful in environments sampled from the domain randomization distribution 2) The domain randomization distribution should be wide enough so that the experience similar to the target robot system is observed during training, while addressing the practicality of training finite capacity models. These two properties aim to ensure the trajectories encountered in the target system are close to those observed during training, as existing methods in machine learning are better suited for interpolation than extrapolation. We show how adapting the domain randomization distribution while training context-conditioned policies results in improvements on jump-start and asymptotic performance when transferring a learned policy to the target environment.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sampling-Based System Identification with Active Exploration for Legged Robot Sim2Real Learning

    cs.RO 2025-05 conditional novelty 7.0 of 10

    SPI-Active identifies legged-robot physical parameters via massive parallel sampling and uses Fisher-information-optimal command sequences to collect informative real-world data, improving sim-to-real transfer on quad...

Pith tools