REVIEW 1 cited by
Decoupled Weight Decay for Any $p$ Norm
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
With the success of deep neural networks (NNs) in a variety of domains, the computational and storage requirements for training and deploying large NNs have become a bottleneck for further improvements. Sparsification has consequently emerged as a leading approach to tackle these issues. In this work, we consider a simple yet effective approach to sparsification, based on the Bridge, or $L_p$ regularization during training. We introduce a novel weight decay scheme, which generalizes the standard $L_2$ weight decay to any $p$ norm. We show that this scheme is compatible with adaptive optimizers, and avoids the gradient divergence associated with $0<p<1$ norms. We empirically demonstrate that it leads to highly sparse networks, while maintaining generalization performance comparable to standard $L_2$ regularization.
Forward citations
Cited by 1 Pith paper
-
Deep Weight Factorization: Sparse Learning Through the Lens of Artificial Symmetries
Factorizing weights into D≥2 multiplicative factors and applying L2 weight decay induces a non-convex sparse L2/D penalty, and with tailored initialization and learning rates, achieves superior sparsity-accuracy tradeoffs.
Discussion (0). Continue with ORCID to comment.