Pith. sign in

REVIEW 2 cited by

Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.00293 v4 pith:22Q4326B submitted 2022-02-01 stat.ML cond-mat.dis-nncs.LG

classification stat.MLcond-mat.dis-nncs.LG
keywords descentgradienthigh-dimensionalnetworksconvergenceinvestigatestochasticable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the non-convex optimization landscape, over-parametrized shallow networks are able to achieve global convergence under gradient descent. The picture can be radically different for narrow networks, which tend to get stuck in badly-generalizing local minima. Here we investigate the cross-over between these two regimes in the high-dimensional setting, and in particular investigate the connection between the so-called mean-field/hydrodynamic regime and the seminal approach of Saad & Solla. Focusing on the case of Gaussian data, we study the interplay between the learning rate, the time scale, and the number of hidden units in the high-dimensional dynamics of stochastic gradient descent (SGD). Our work builds on a deterministic description of SGD in high-dimensions from statistical physics, which we extend and for which we provide rigorous convergence rates.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Joint Learning in the Gaussian Single Index Model

    cs.LG 2025-05 conditional novelty 6.0 of 10

    In Gaussian single-index models, joint gradient flow over direction and link function converges to the true regression function from either sign of initial alignment, with rate governed by the information exponent.

  2. Limit Theorems for Stochastic Gradient Descent in High-Dimensional Single-Layer Networks

    stat.ML 2025-11 unverdicted novelty 5.0 of 10

    At the critical step-size scaling for SGD in high-dimensional single-layer networks, effective dynamics gain a diffusive correction term that changes the phase diagram and reduces to an Ornstein-Uhlenbeck process near...

Pith tools