Pith. sign in

REVIEW 1 cited by

Deep linear networks for regression are implicitly regularized towards flat minima

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.13456 v2 pith:SECUPJS2 submitted 2024-05-22 stat.ML cs.LG

classification stat.MLcs.LG
keywords initializationsharpnessgradientnetworksresidualarbitrarilyboundconstant
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The largest eigenvalue of the Hessian, or sharpness, of neural networks is a key quantity to understand their optimization dynamics. In this paper, we study the sharpness of deep linear networks for univariate regression. Minimizers can have arbitrarily large sharpness, but not an arbitrarily small one. Indeed, we show a lower bound on the sharpness of minimizers, which grows linearly with depth. We then study the properties of the minimizer found by gradient flow, which is the limit of gradient descent with vanishing learning rate. We show an implicit regularization towards flat minima: the sharpness of the minimizer is no more than a constant times the lower bound. The constant depends on the condition number of the data covariance matrix, but not on width or depth. This result is proven both for a small-scale initialization and a residual initialization. Results of independent interest are shown in both cases. For small-scale initialization, we show that the learned weight matrices are approximately rank-one and that their singular vectors align. For residual initialization, convergence of the gradient flow for a Gaussian initialization of the residual network is proven. Numerical experiments illustrate our results and connect them to gradient descent with non-vanishing learning rate.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Experimental data re-uploading with provable enhanced learning capabilities

    quant-ph 2025-07 conditional novelty 6.0 of 10

    Separating encoding from trainable gates in one-qubit data re-uploading yields finite VC dimension and provable generalization, demonstrated on a photonic chip.

Pith tools