REVIEW 2 cited by
A Note on Connectivity of Sublevel Sets in Deep Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
It is shown that for deep neural networks, a single wide layer of width $N+1$ ($N$ being the number of training samples) suffices to prove the connectivity of sublevel sets of the training loss function. In the two-layer setting, the same property may not hold even if one has just one neuron less (i.e. width $N$ can lead to disconnected sublevel sets).
Forward citations
Cited by 2 Pith papers
-
Understanding Mode Connectivity via Parameter Space Symmetry
Minima of linear networks have 2^{l-1} connected components, skip connections can reduce this count, and rescaling symmetries can make linear interpolation between minima fail.
-
Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization
Weight-decay-regularized two-layer ReLU networks need width exponential in the number of samples for a benign loss landscape, and small initialization can still converge to spurious minima.
Discussion (0). Continue with ORCID to comment.