Pith. sign in

REVIEW 15 cited by

Disentangling the Causes of Plasticity Loss in Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.18762 v1 pith:RGNXWHTE submitted 2024-02-29 cs.LG

classification cs.LG
keywords learningplasticitylossmechanismsnetworkalgorithmsassumptionhighly
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Underpinning the past decades of work on the design, initialization, and optimization of neural networks is a seemingly innocuous assumption: that the network is trained on a \textit{stationary} data distribution. In settings where this assumption is violated, e.g.\ deep reinforcement learning, learning algorithms become unstable and brittle with respect to hyperparameters and even random seeds. One factor driving this instability is the loss of plasticity, meaning that updating the network's predictions in response to new information becomes more difficult as training progresses. While many recent works provide analyses and partial solutions to this phenomenon, a fundamental question remains unanswered: to what extent do known mechanisms of plasticity loss overlap, and how can mitigation strategies be combined to best maintain the trainability of a network? This paper addresses these questions, showing that loss of plasticity can be decomposed into multiple independent mechanisms and that, while intervening on any single mechanism is insufficient to avoid the loss of plasticity in all cases, intervening on multiple mechanisms in conjunction results in highly robust learning algorithms. We show that a combination of layer normalization and weight decay is highly effective at maintaining plasticity in a variety of synthetic nonstationary learning tasks, and further demonstrate its effectiveness on naturally arising nonstationarities, including reinforcement learning in the Arcade Learning Environment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Single-Model Optimization: Preserving Plasticity in Continual Reinforcement Learning

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    TeLAPA preserves behaviorally diverse policy neighborhoods in a shared latent space, improving MiniGrid continual RL transfer, revisit recovery, and retention over single-model preservation.

  2. V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control

    cs.LG 2026-08 conditional novelty 6.0 of 10

    V-Simba, a visual RL architecture combining layer normalization, weight decay, and a distributional critic, matches or outperforms complex baselines on 29 continuous control tasks while using less compute.

  3. NeuMoSync: End-to-End Neuromodulatory Control for Plasticity and Adaptability in Continual Learning

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A global neuromodulatory controller with per-neuron weight, activation, and offset modulation preserves plasticity and improves forward and backward adaptation in continual learning.

  4. What Can Grokking Teach Us About Learning Under Nonstationarity?

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Periodically increasing the effective learning rate while constraining parameter norms induces feature-learning dynamics and mitigates primacy bias in grokking, warm-starting, and reinforcement learning.

  5. Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    One-shot random pruning at initialization lets deep RL networks keep improving at model sizes where dense networks collapse in performance.

  6. Continual Hyperbolic Learning of Instances and Classes

    cs.CV 2025-06 conditional novelty 6.0 of 10

    HyperCLIC embeds the instance-class hierarchy in hyperbolic space and uses hyperbolic classification and distillation losses to continuously learn both fine-grained instances and coarse-grained classes on EgoObjects, ...

  7. Optimistic critics can empower small actors

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Optimistic critic estimates (mean or max over two Q functions) counteract value underestimation and recover most of the performance lost by shrinking the actor in SAC and DrQ.

  8. Mitigating Plasticity Loss in Continual Reinforcement Learning by Reducing Churn

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Reducing churn in continual RL via C-CHAIN prevents NTK rank collapse and substantially improves learning across four benchmark suites.

  9. Silent Neuron Theory and Plasticity Preservation for Deep Reinforcement Learning in Adaptive Video Streaming

    cs.LG 2025-05 unverdicted novelty 6.0 of 10

    Silent Neuron theory provides a framework for plasticity degradation in deep RL, and ReSiN preserves it via forward-backward guided resets, yielding up to 168% higher bitrate and 108% better QoE in video streaming.

  10. Torque-Aware Momentum

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Torque-Aware Momentum damps momentum updates by the alignment between new gradients and previous momentum, giving small gains on some benchmarks but mixed results on large model fine-tuning.

  11. Calibrated Partial Resets: Preventing Policy Collapse in Continual Reinforcement Learning

    cs.LG 2026-07 conditional novelty 5.5 of 10

    Utility-scaled partial neuron resets prevent policy collapse in long-horizon continual RL while matching or beating binary-reset and uniform-decay baselines on several benchmarks.

  12. An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?

    cs.CL 2026-07 conditional novelty 5.5 of 10

    Emergent misalignment and realignment are brittle surface effects driven by dataset artifacts like response length rather than stable representational changes.

  13. Relative Value Learning

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A critic that learns antisymmetric value differences ∆(s_i,s_j)=V(s_i)−V(s_j) has a provably contracting Bellman operator and an unbiased advantage estimator, and PPO with this critic matches standard PPO on Atari.

  14. Preserving Plasticity in Continual Learning with Adaptive Linearity Injection

    cs.LG 2025-05 conditional novelty 5.0 of 10

    AdaLin, a per-neuron learnable linearity injection gated by activation saturation, preserves plasticity in continual learning and off-policy RL without task boundaries or extra method hyperparameters.

  15. Activation by Interval-wise Dropout: A Simple Way to Prevent Neural Networks from Plasticity Loss

    cs.LG 2025-02 conditional novelty 5.0 of 10

    AID, a stochastic activation applying different dropout rates to positive and negative preactivations, mitigates plasticity loss and improves continual learning, reinforcement learning, and standard supervised learnin...

Pith tools