REVIEW 1 cited by
How noise affects the Hessian spectrum in overparameterized neural networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Stochastic gradient descent (SGD) forms the core optimization method for deep neural networks. While some theoretical progress has been made, it still remains unclear why SGD leads the learning dynamics in overparameterized networks to solutions that generalize well. Here we show that for overparameterized networks with a degenerate valley in their loss landscape, SGD on average decreases the trace of the Hessian of the loss. We also generalize this result to other noise structures and show that isotropic noise in the non-degenerate subspace of the Hessian decreases its determinant. In addition to explaining SGDs role in sculpting the Hessian spectrum, this opens the door to new optimization approaches that may confer better generalization performance. We test our results with experiments on toy models and deep neural networks.
Forward citations
Cited by 1 Pith paper
-
Noise-Driven Exploration and Transient Freezing Select Flat Minima in Stochastic Gradient Descent
SGD's preference for flat minima is explained by a noise-controlled transient exploration phase that ends in a freezing transition; stronger noise delays freezing and biases selection toward flatter valleys.
Discussion (0). Continue with ORCID to comment.