Pith. sign in

REVIEW 3 cited by

On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.13530 v1 pith:ZCL4F3SJ submitted 2020-05-27 math.AP cs.LGstat.ML

classification math.APcs.LGstat.ML
keywords convergenceconditiondescentdistributionfieldgradientmeanparameter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We describe a necessary and sufficient condition for the convergence to minimum Bayes risk when training two-layer ReLU-networks by gradient descent in the mean field regime with omni-directional initial parameter distribution. This article extends recent results of Chizat and Bach to ReLU-activated networks and to the situation in which there are no parameters which exactly achieve MBR. The condition does not depend on the initalization of parameters and concerns only the weak convergence of the realization of the neural network, not its parameter distribution.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 14 citations worldwide. Full citation record

  1. Elliptic Regularity Theory in Barron Spaces and Applications to the Deep Ritz Method

    math.AP 2026-07 accept novelty 7.0 of 10

    Harmonic functions with Barron Dirichlet data fail to be Lipschitz or H², yet admit Barron approximants of norm ~|log ε| with error ~ε on half-spaces and 2D rectangles, giving Deep Ritz a priori rates.

  2. Singular perturbations and hierarchical learning in two-layer neural networks

    cs.LG 2026-07 accept novelty 7.0 of 10

    Constant and linear Hermite components of a misspecified single-index target are recovered at the conjectured singular-perturbation timescales; quadratic learning remains coupled to them via an auxiliary constrained flow.

  3. Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Weight-decay-regularized two-layer ReLU networks need width exponential in the number of samples for a benign loss landscape, and small initialization can still converge to spurious minima.

Pith tools