Pith. sign in

REVIEW 3 cited by

Dynamics Generalization via Information Bottleneck in Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.00614 v1 pith:BBU64I5X submitted 2020-08-03 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords generalizationinformationtrainingagentslearningdeepdifferentdynamics
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Despite the significant progress of deep reinforcement learning (RL) in solving sequential decision making problems, RL agents often overfit to training environments and struggle to adapt to new, unseen environments. This prevents robust applications of RL in real world situations, where system dynamics may deviate wildly from the training settings. In this work, our primary contribution is to propose an information theoretic regularization objective and an annealing-based optimization method to achieve better generalization ability in RL agents. We demonstrate the extreme generalization benefits of our approach in different domains ranging from maze navigation to robotic tasks; for the first time, we show that agents can generalize to test parameters more than 10 standard deviations away from the training parameter distribution. This work provides a principled way to improve generalization in RL by gradually removing information that is redundant for task-solving; it opens doors for the systematic study of generalization from training to extremely different testing settings, focusing on the established connections between information theory and machine learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Maximum Total Correlation Reinforcement Learning

    cs.LG 2025-05 conditional novelty 7.0 of 10

    A reinforcement learning regularizer that maximizes trajectory-level total correlation produces simpler, more compressible policies that are more robust to perturbations.

  2. Generalization Capability for Imitation Learning

    cs.LG 2025-04 conditional novelty 4.0 of 10

    The paper argues that imitation learning generalization is governed by representation compression and encoder-data dependence, and that high conditional entropy in actions tightens the bound, but the key new bound is ...

  3. TEA: Trajectory Encoding Augmentation for Robust and Transferable Policies in Offline Reinforcement Learning

    cs.LG 2024-11 conditional novelty 4.0 of 10

    Appending autoencoder-derived trajectory encodings to the state space improves offline RL transfer to new CartPole dynamics compared with BCQ, though the effect is small in some environments.

Pith tools