Pith. sign in

REVIEW 7 cited by

Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.13649 v4 pith:SEG5KOYF submitted 2020-04-28 cs.LG cs.CVeess.IVstat.ML

classification cs.LGcs.CVeess.IVstat.ML
keywords learningaugmentationmodel-freepixelsreinforcementapproachdeepenabling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a simple data augmentation technique that can be applied to standard model-free reinforcement learning algorithms, enabling robust learning directly from pixels without the need for auxiliary losses or pre-training. The approach leverages input perturbations commonly used in computer vision tasks to regularize the value function. Existing model-free approaches, such as Soft Actor-Critic (SAC), are not able to train deep networks effectively from image pixels. However, the addition of our augmentation method dramatically improves SAC's performance, enabling it to reach state-of-the-art performance on the DeepMind control suite, surpassing model-based (Dreamer, PlaNet, and SLAC) methods and recently proposed contrastive learning (CURL). Our approach can be combined with any model-free reinforcement learning algorithm, requiring only minor modifications. An implementation can be found at https://sites.google.com/view/data-regularized-q.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-Predictive Dynamics for Generalization of Vision-based Reinforcement Learning

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A self-supervised auxiliary loss combining weak and strong augmentations, an adversarial discriminator, and inverse-then-forward latent dynamics improves both data efficiency and zero-shot generalization in vision-based RL.

  2. DeGuV: Depth-Guided Visual Reinforcement Learning for Generalization and Interpretability in Manipulation

    cs.RO 2025-09 conditional novelty 6.0 of 10

    Depth-guided masking improves visual RL generalization, sample efficiency, and interpretability on manipulation tasks.

  3. Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A single multi-task RL agent using a large regularized critic, categorical value loss, and task embeddings achieves state-of-the-art results across 283 tasks and transfers efficiently to new tasks.

  4. SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Grid-parallel translational trajectory symmetries plus sample-time homographies cut on-robot RL wall-clock training 1.37–2.17× versus SERL on three real contact tasks.

  5. Trading Human Curation for Synthetic Augmentation in RLVR

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    Gated synthetic augmentations of a 10-task human base substitute for ~87 extra human RLVR tasks on aggregate held-out pass@1, with cost-adjusted trade rate ρ_cost in [1.4×, 11.6×].

  6. A Survey of State Representation Learning for Deep Reinforcement Learning

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A six-class taxonomy of state representation learning methods for model-free online deep reinforcement learning, with selection guidelines, evaluation metrics, and future directions.

  7. Dream to Generalize: Zero-Shot Model-Based Reinforcement Learning for Unseen Visual Distractions

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Dr. G combines dual contrastive learning and a recurrent inverse-dynamics objective in a Dreamer-style world model to improve zero-shot generalization to unseen visual distractions in control tasks.

Pith tools