REVIEW 9 cited by
Understanding and Preventing Capacity Loss in Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The reinforcement learning (RL) problem is rife with sources of non-stationarity, making it a notoriously difficult problem domain for the application of neural networks. We identify a mechanism by which non-stationary prediction targets can prevent learning progress in deep RL agents: \textit{capacity loss}, whereby networks trained on a sequence of target values lose their ability to quickly update their predictions over time. We demonstrate that capacity loss occurs in a range of RL agents and environments, and is particularly damaging to performance in sparse-reward tasks. We then present a simple regularizer, Initial Feature Regularization (InFeR), that mitigates this phenomenon by regressing a subspace of features towards its value at initialization, leading to significant performance improvements in sparse-reward environments such as Montezuma's Revenge. We conclude that preventing capacity loss is crucial to enable agents to maximally benefit from the learning signals they obtain throughout the entire training trajectory.
Forward citations
Cited by 9 Pith papers
-
Beyond Single-Model Optimization: Preserving Plasticity in Continual Reinforcement Learning
TeLAPA preserves behaviorally diverse policy neighborhoods in a shared latent space, improving MiniGrid continual RL transfer, revisit recovery, and retention over single-model preservation.
-
When Does Continual Learning Require Learning
Different patterns of environmental change (space vs time) require different LLM update behaviors; no single family of methods—prompts, distillation, RL, or compression—handles all regimes.
-
Memory Merge DQN: Sensitivity Weighted Target Updates for Stable Value Learning
Replacing the hard target-copy in DQN with a Q-value-sensitivity-weighted merge of the last K network copies yields competitive Atari performance, but only marginally beats architecture-matched baselines.
-
Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation
MPAIL2 demonstrates real-world manipulation learning from observation alone, without rewards or action labels, plus positive online transfer.
-
Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning
One-shot random pruning at initialization lets deep RL networks keep improving at model sizes where dense networks collapse in performance.
-
Parseval Regularization for Continual Reinforcement Learning
Parseval regularization, a cheap orthogonality-preserving penalty, improves continual RL agents' success on new tasks across gridworld, CARL and MetaWorld benchmarks.
-
Sustaining Plasticity via Learnable Wavelet Activations in Continual Learning
A learnable wavelet activation with dynamic capacity injection and slope regularization improves plasticity retention in continual learning.
-
Activation by Interval-wise Dropout: A Simple Way to Prevent Neural Networks from Plasticity Loss
AID, a stochastic activation applying different dropout rates to positive and negative preactivations, mitigates plasticity loss and improves continual learning, reinforcement learning, and standard supervised learnin...
-
Fisher-Guided Selective Forgetting: Mitigating The Primacy Bias in Deep Reinforcement Learning
FGSF, periodic FIM-scaled weight noise for SAC, improves Humanoid and Quadruped but underperforms plain SAC on five of ten DMC tasks.
Discussion (0). Continue with ORCID to comment.