Pith. sign in

REVIEW 3 cited by

How to pick the domain randomization parameters for sim-to-real transfer of reinforcement learning policies?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1903.11774 v1 pith:67BWXEGX submitted 2019-03-28 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords distributionrealalgorithmsdomainrandomizationlearningparametersperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, reinforcement learning (RL) algorithms have demonstrated remarkable success in learning complicated behaviors from minimally processed input. However, most of this success is limited to simulation. While there are promising successes in applying RL algorithms directly on real systems, their performance on more complex systems remains bottle-necked by the relative data inefficiency of RL algorithms. Domain randomization is a promising direction of research that has demonstrated impressive results using RL algorithms to control real robots. At a high level, domain randomization works by training a policy on a distribution of environmental conditions in simulation. If the environments are diverse enough, then the policy trained on this distribution will plausibly generalize to the real world. A human-specified design choice in domain randomization is the form and parameters of the distribution of simulated environments. It is unclear how to the best pick the form and parameters of this distribution and prior work uses hand-tuned distributions. This extended abstract demonstrates that the choice of the distribution plays a major role in the performance of the trained policies in the real world and that the parameter of this distribution can be optimized to maximize the performance of the trained policies in the real world

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Solving Rubik's Cube with a Robot Hand

    cs.LG 2019-10 accept novelty 7.0 of 10

    Reinforcement learning models trained only in simulation using automatic domain randomization solve Rubik's cube with a real robot hand.

  2. GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation

    cs.RO 2025-08 conditional novelty 5.0 of 10

    GBC unifies MoCap retargeting and imitation learning into one framework that trains whole-body humanoid policies across multiple robot morphologies in simulation.

  3. How Should a Simulation-to-Reality Transfer Budget Be Spent?

    cs.RO 2026-06 unverdicted novelty 4.0 of 10

    Controlled sim-to-sim experiments indicate that system identification with few rollouts outperforms widened domain randomization for closing sim-to-real gaps when dynamics are identifiable.

Pith tools