REVIEW 2 cited by
Optimal bump functions for shallow ReLU networks: Weight decay, depth separation and the curse of dimensionality
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
In this note, we study how neural networks with a single hidden layer and ReLU activation interpolate data drawn from a radially symmetric distribution with target labels 1 at the origin and 0 outside the unit ball, if no labels are known inside the unit ball. With weight decay regularization and in the infinite neuron, infinite data limit, we prove that a unique radially symmetric minimizer exists, whose weight decay regularizer and Lipschitz constant grow as $d$ and $\sqrt{d}$ respectively. We furthermore show that the weight decay regularizer grows exponentially in $d$ if the label $1$ is imposed on a ball of radius $\varepsilon$ rather than just at the origin. By comparison, a neural networks with two hidden layers can approximate the target function without encountering the curse of dimensionality.
Forward citations
Cited by 2 Pith papers
-
The Barron-Lipschitz Energy Gap and Depth Separation Phenomena in Scientific Machine Learning
Barron functions can fail to reach the Lipschitz-class infimum of certain variational energies—including a thin-shell folding energy where circular folds beat straight-line folds—while compositions of two Barron funct...
-
Elliptic Regularity Theory in Barron Spaces and Applications to the Deep Ritz Method
Harmonic functions with Barron Dirichlet data fail to be Lipschitz or H², yet admit Barron approximants of norm ~|log ε| with error ~ε on half-spaces and 2D rectangles, giving Deep Ritz a priori rates.
Discussion (0). Sign in to comment.