Pith. sign in

REVIEW 2 cited by

Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.05396 v3 pith:LGVVVWON submitted 2019-10-11 cs.LG stat.ML

classification cs.LGstat.ML
keywords agentsdeeplearningtrainedacrossenvironmentsgeneralizationmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep reinforcement learning (RL) agents often fail to generalize to unseen environments (yet semantically similar to trained agents), particularly when they are trained on high-dimensional state spaces, such as images. In this paper, we propose a simple technique to improve a generalization ability of deep RL agents by introducing a randomized (convolutional) neural network that randomly perturbs input observations. It enables trained agents to adapt to new domains by learning robust features invariant across varied and randomized environments. Furthermore, we consider an inference method based on the Monte Carlo approximation to reduce the variance induced by this randomization. We demonstrate the superiority of our method across 2D CoinRun, 3D DeepMind Lab exploration and 3D robotics control tasks: it significantly outperforms various regularization and data augmentation methods for the same purpose.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-Predictive Dynamics for Generalization of Vision-based Reinforcement Learning

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A self-supervised auxiliary loss combining weak and strong augmentations, an adversarial discriminator, and inverse-then-forward latent dynamics improves both data efficiency and zero-shot generalization in vision-based RL.

  2. LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A preference-conditioned PPO routing policy with IRT-based model identity vectors selects cost-effective LLMs per query and generalizes to unseen models from a handful of evaluation prompts.

Pith tools