Pith. sign in

REVIEW 3 cited by

Robust Visual Domain Randomization for Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.10537 v2 pith:AIOLDITX submitted 2019-10-23 cs.LG cs.AI

classification cs.LGcs.AI
keywords domainrandomizationlearningvisualacrossagentdomainsenvironment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Producing agents that can generalize to a wide range of visually different environments is a significant challenge in reinforcement learning. One method for overcoming this issue is visual domain randomization, whereby at the start of each training episode some visual aspects of the environment are randomized so that the agent is exposed to many possible variations. However, domain randomization is highly inefficient and may lead to policies with high variance across domains. Instead, we propose a regularization method whereby the agent is only trained on one variation of the environment, and its learned state representations are regularized during training to be invariant across domains. We conduct experiments that demonstrate that our technique leads to more efficient and robust learning than standard domain randomization, while achieving equal generalization scores.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Unified Causal-Origin Taxonomy of Distributional Shifts in Reinforcement Learning

    cs.LG 2026-06 conditional novelty 6.5 of 10

    Distributional shift in RL is classified by which POMDP generative component changes (internal agent vs external environment) and by whether the time boundary is explicit, implicit, or hybrid.

  2. DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    DADiff estimates cross-domain dynamics mismatch from diffusion-model latent-state trajectories and uses it for reward modification or data selection in policy adaptation.

  3. rQdia: Regularizing Q-Value Distributions With Image Augmentation

    cs.LG 2025-06 reject novelty 5.0 of 10

    rQdia regularizes Q-value distributions across augmented images, reporting improvements on continuous control and Atari benchmarks over DrQ, SAC, and Data-Efficient Rainbow.

Pith tools