REVIEW 3 cited by
Robust Visual Domain Randomization for Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Producing agents that can generalize to a wide range of visually different environments is a significant challenge in reinforcement learning. One method for overcoming this issue is visual domain randomization, whereby at the start of each training episode some visual aspects of the environment are randomized so that the agent is exposed to many possible variations. However, domain randomization is highly inefficient and may lead to policies with high variance across domains. Instead, we propose a regularization method whereby the agent is only trained on one variation of the environment, and its learned state representations are regularized during training to be invariant across domains. We conduct experiments that demonstrate that our technique leads to more efficient and robust learning than standard domain randomization, while achieving equal generalization scores.
Forward citations
Cited by 3 Pith papers
-
A Unified Causal-Origin Taxonomy of Distributional Shifts in Reinforcement Learning
Distributional shift in RL is classified by which POMDP generative component changes (internal agent vs external environment) and by whether the time boundary is explicit, implicit, or hybrid.
-
DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning
DADiff estimates cross-domain dynamics mismatch from diffusion-model latent-state trajectories and uses it for reward modification or data selection in policy adaptation.
-
rQdia: Regularizing Q-Value Distributions With Image Augmentation
rQdia regularizes Q-value distributions across augmented images, reporting improvements on continuous control and Atari benchmarks over DrQ, SAC, and Data-Efficient Rainbow.
Discussion (0). Sign in to comment.