Pith. sign in

REVIEW 4 cited by

Beyond the Quadratic Approximation: the Multiscale Structure of Neural Network Loss Landscapes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.11326 v3 pith:BSCGEISO submitted 2022-04-24 cs.LG

classification cs.LG
keywords lossnetworkneuralscalesstructureapproximationexplainmultiscale
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A quadratic approximation of neural network loss landscapes has been extensively used to study the optimization process of these networks. Though, it usually holds in a very small neighborhood of the minimum, it cannot explain many phenomena observed during the optimization process. In this work, we study the structure of neural network loss functions and its implication on optimization in a region beyond the reach of a good quadratic approximation. Numerically, we observe that neural network loss functions possesses a multiscale structure, manifested in two ways: (1) in a neighborhood of minima, the loss mixes a continuum of scales and grows subquadratically, and (2) in a larger region, the loss shows several separate scales clearly. Using the subquadratic growth, we are able to explain the Edge of Stability phenomenon [5] observed for the gradient descent (GD) method. Using the separate scales, we explain the working mechanism of learning rate decay by simple examples. Finally, we study the origin of the multiscale structure and propose that the non-convexity of the models and the non-uniformity of training data is one of the causes. By constructing a two-layer neural network problem we show that training data with different magnitudes give rise to different scales of the loss function, producing subquadratic growth and multiple separate scales.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Implicit Bias of SGD in Multivariate ReLU Networks: Effective Width Collapse

    cs.LG 2026-07 accept novelty 7.0 of 10

    Noisy SGD in the mean-field regime forces wide multivariate ReLU networks to an effective width of at most 2P-1, yielding a continuous piecewise-affine predictor whose hyperplanes are non-redundant with respect to the...

  2. On the Stability of Nonlinear Dynamics in GD and SGD: Beyond Quadratic Potentials

    cs.LG 2026-02 conditional novelty 7.0 of 10

    Stable oscillations of GD near sharp minima are characterized by a multivariate derivative condition, and SGD stability in expectation is governed by a worst-case batch.

  3. From Logistic Regression to the Perceptron Algorithm: Exploring Gradient Descent with Large Step Sizes

    cs.LG 2024-12 conditional novelty 7.0 of 10

    Logistic regression with gradient descent and infinite step size is the batch perceptron, and a normalized version achieves an n times better iteration complexity.

  4. Criteria and Bias of Parameterized Linear Regression under Edge of Stability Regime

    math.OC 2024-12 conditional novelty 6.0 of 10

    Under specific conditions, gradient descent converges in the unstable edge-of-stability regime for a quadratic loss on a depth-2 diagonal linear network, with a bias bound depending on step size and initialization.

Pith tools