Pith. sign in

REVIEW 9 cited by

Understanding and Preventing Capacity Loss in Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.09560 v2 pith:TWANQE5Q submitted 2022-04-20 cs.LG

classification cs.LG
keywords capacitylearninglossagentsenvironmentsnetworksperformancepreventing
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The reinforcement learning (RL) problem is rife with sources of non-stationarity, making it a notoriously difficult problem domain for the application of neural networks. We identify a mechanism by which non-stationary prediction targets can prevent learning progress in deep RL agents: \textit{capacity loss}, whereby networks trained on a sequence of target values lose their ability to quickly update their predictions over time. We demonstrate that capacity loss occurs in a range of RL agents and environments, and is particularly damaging to performance in sparse-reward tasks. We then present a simple regularizer, Initial Feature Regularization (InFeR), that mitigates this phenomenon by regressing a subspace of features towards its value at initialization, leading to significant performance improvements in sparse-reward environments such as Montezuma's Revenge. We conclude that preventing capacity loss is crucial to enable agents to maximally benefit from the learning signals they obtain throughout the entire training trajectory.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Single-Model Optimization: Preserving Plasticity in Continual Reinforcement Learning

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    TeLAPA preserves behaviorally diverse policy neighborhoods in a shared latent space, improving MiniGrid continual RL transfer, revisit recovery, and retention over single-model preservation.

  2. When Does Continual Learning Require Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Different patterns of environmental change (space vs time) require different LLM update behaviors; no single family of methods—prompts, distillation, RL, or compression—handles all regimes.

  3. Memory Merge DQN: Sensitivity Weighted Target Updates for Stable Value Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Replacing the hard target-copy in DQN with a Q-value-sensitivity-weighted merge of the last K network copies yields competitive Atari performance, but only marginally beats architecture-matched baselines.

  4. Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation

    cs.RO 2026-02 conditional novelty 6.0 of 10

    MPAIL2 demonstrates real-world manipulation learning from observation alone, without rewards or action labels, plus positive online transfer.

  5. Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    One-shot random pruning at initialization lets deep RL networks keep improving at model sizes where dense networks collapse in performance.

  6. Parseval Regularization for Continual Reinforcement Learning

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Parseval regularization, a cheap orthogonality-preserving penalty, improves continual RL agents' success on new tasks across gridworld, CARL and MetaWorld benchmarks.

  7. Sustaining Plasticity via Learnable Wavelet Activations in Continual Learning

    cs.LG 2026-08 conditional novelty 5.0 of 10

    A learnable wavelet activation with dynamic capacity injection and slope regularization improves plasticity retention in continual learning.

  8. Activation by Interval-wise Dropout: A Simple Way to Prevent Neural Networks from Plasticity Loss

    cs.LG 2025-02 conditional novelty 5.0 of 10

    AID, a stochastic activation applying different dropout rates to positive and negative preactivations, mitigates plasticity loss and improves continual learning, reinforcement learning, and standard supervised learnin...

  9. Fisher-Guided Selective Forgetting: Mitigating The Primacy Bias in Deep Reinforcement Learning

    cs.LG 2025-02 reject novelty 5.0 of 10

    FGSF, periodic FIM-scaled weight noise for SAC, improves Humanoid and Quadruped but underperforms plain SAC on five of ten DMC tasks.

Pith tools