Pith. sign in

REVIEW 1 cited by

On the Diminishing Returns of Width for Continual Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.06398 v3 pith:JTF3E5SR submitted 2024-03-11 cs.LG cs.AI

classification cs.LGcs.AI
keywords forgettingwidthcontinualdiminishinglearningreturnscatastrophicdemonstrated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While deep neural networks have demonstrated groundbreaking performance in various settings, these models often suffer from \emph{catastrophic forgetting} when trained on new tasks in sequence. Several works have empirically demonstrated that increasing the width of a neural network leads to a decrease in catastrophic forgetting but have yet to characterize the exact relationship between width and continual learning. We design one of the first frameworks to analyze Continual Learning Theory and prove that width is directly related to forgetting in Feed-Forward Networks (FFN). Specifically, we demonstrate that increasing network widths to reduce forgetting yields diminishing returns. We empirically verify our claims at widths hitherto unexplored in prior studies where the diminishing returns are clearly observed as predicted by our theory.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Measuring Representational Shifts in Continual Learning: A Linear Transformation Perspective

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Representation discrepancy, a new metric with theoretical bounds, shows continual learning forgets features faster in deeper layers and slower in wider networks.

Pith tools