Pith. sign in

REVIEW 5 cited by

Understanding Domain Randomization for Sim-to-real Transfer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.03239 v2 pith:7HPQQIFW submitted 2021-10-07 cs.LG

classification cs.LG
keywords sim-to-realtransferdomainrandomizationrealworldalgorithmslearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement learning encounters many challenges when applied directly in the real world. Sim-to-real transfer is widely used to transfer the knowledge learned from simulation to the real world. Domain randomization -- one of the most popular algorithms for sim-to-real transfer -- has been demonstrated to be effective in various tasks in robotics and autonomous driving. Despite its empirical successes, theoretical understanding on why this simple algorithm works is limited. In this paper, we propose a theoretical framework for sim-to-real transfers, in which the simulator is modeled as a set of MDPs with tunable parameters (corresponding to unknown physical parameters such as friction). We provide sharp bounds on the sim-to-real gap -- the difference between the value of policy returned by domain randomization and the value of an optimal policy for the real world. We prove that sim-to-real transfer can succeed under mild conditions without any real-world training samples. Our theory also highlights the importance of using memory (i.e., history-dependent policies) in domain randomization. Our proof is based on novel techniques that reduce the problem of bounding the sim-to-real gap to the problem of designing efficient learning algorithms for infinite-horizon MDPs, which we believe are of independent interest.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 33 citations worldwide. Full citation record

  1. Hold My Beer: Learning Gentle Humanoid Locomotion and End-Effector Stabilization Control

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A slow-fast two-agent reinforcement learning architecture with separate upper- and lower-body policies reduces end-effector shaking during humanoid locomotion.

  2. Simulation Priors for Data-Efficient Deep Learning

    cs.LG 2025-09 conditional novelty 5.0 of 10

    SimPEL uses low-fidelity simulators plus a Gaussian-process gap as a functional prior for Bayesian neural networks, improving data efficiency in dynamics learning and model-based RL.

  3. Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A pessimism-based framework for zero-shot transfer RL builds conservative proxies from robust MDPs, yielding lower-bound performance guarantees and distributed algorithms that mitigate negative transfer.

  4. McARL:Morphology-Control-Aware Reinforcement Learning for Generalizable Quadrupedal Locomotion

    cs.RO 2025-05 conditional novelty 5.0 of 10

    A morphology-conditioned PPO policy trained only on the Unitree Go1 transfers zero-shot in simulation to Go2, A1 and Mini Cheetah, with the best variant reaching 3.5 m/s on the Go2.

  5. Foundation Model Driven Robotics: A Comprehensive Review

    cs.RO 2025-07 conditional novelty 2.0 of 10

    A review of foundation-model-driven robotics that synthesizes recent work across perception, planning, control, HRI, simulation, and sim-to-real transfer, and highlights open challenges.

Pith tools