Increasing network width reduces forgetting only in lazy training; the best continual-learning performance occurs at a small critical level of feature learning that transfers across model sizes.
Our theory characterizes the effect of increasing the degree of feature learningon CF, starting from the lazy training setting – for which CF is already known
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The Importance of Being Lazy: Scaling Limits of Continual Learning
Increasing network width reduces forgetting only in lazy training; the best continual-learning performance occurs at a small critical level of feature learning that transfers across model sizes.