REVIEW 3 cited by
Generalisation Guarantees for Continual Learning with Orthogonal Gradient Descent
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In Continual Learning settings, deep neural networks are prone to Catastrophic Forgetting. Orthogonal Gradient Descent was proposed to tackle the challenge. However, no theoretical guarantees have been proven yet. We present a theoretical framework to study Continual Learning algorithms in the Neural Tangent Kernel regime. This framework comprises closed form expression of the model through tasks and proxies for Transfer Learning, generalisation and tasks similarity. In this framework, we prove that OGD is robust to Catastrophic Forgetting then derive the first generalisation bound for SGD and OGD for Continual Learning. Finally, we study the limits of this framework in practice for OGD and highlight the importance of the Neural Tangent Kernel variation for Continual Learning with OGD.
Forward citations
Cited by 3 Pith papers
-
Interference and Retention in Continual Learning
In the frozen-feature regime, forgetting of task A after learning B equals exactly ½ΔᵀΣ_AΔ, and this geometry yields IGFA, a signed share-or-protect rule that is lossless on disjoint supports and relocates unavoidable...
-
Reactivation: Empirical NTK Dynamics Under Task Shifts
Task transitions in continual learning cause abrupt, width-persistent changes in the Neural Tangent Kernel of past data, a phenomenon the authors call reactivation, which is modulated by semantic novelty of the new classes.
-
Measuring Representational Shifts in Continual Learning: A Linear Transformation Perspective
Representation discrepancy, a new metric with theoretical bounds, shows continual learning forgets features faster in deeper layers and slower in wider networks.
Discussion (0). Sign in to comment.