REVIEW 2 cited by
Penalising the biases in norm regularisation enforces sparsity
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Controlling the parameters' norm often yields good generalisation when training neural networks. Beyond simple intuitions, the relation between regularising parameters' norm and obtained estimators remains theoretically misunderstood. For one hidden ReLU layer networks with unidimensional data, this work shows the parameters' norm required to represent a function is given by the total variation of its second derivative, weighted by a $\sqrt{1+x^2}$ factor. Notably, this weighting factor disappears when the norm of bias terms is not regularised. The presence of this additional weighting factor is of utmost significance as it is shown to enforce the uniqueness and sparsity (in the number of kinks) of the minimal norm interpolator. Conversely, omitting the bias' norm allows for non-sparse solutions. Penalising the bias terms in the regularisation, either explicitly or implicitly, thus leads to sparse estimators.
Forward citations
Cited by 2 Pith papers
-
The Barron-Lipschitz Energy Gap and Depth Separation Phenomena in Scientific Machine Learning
Barron functions can fail to reach the Lipschitz-class infimum of certain variational energies—including a thin-shell folding energy where circular folds beat straight-line folds—while compositions of two Barron funct...
-
Elliptic Regularity Theory in Barron Spaces and Applications to the Deep Ritz Method
Harmonic functions with Barron Dirichlet data fail to be Lipschitz or H², yet admit Barron approximants of norm ~|log ε| with error ~ε on half-spaces and 2D rectangles, giving Deep Ritz a priori rates.
Discussion (0). Sign in to comment.