Pith. sign in

REVIEW 3 cited by

Mean Field Limit of the Learning Dynamics of Multilayer Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1902.02880 v1 pith:AEFFHPXI submitted 2019-02-07 cs.LG cond-mat.dis-nncond-mat.stat-mechstat.ML

classification cs.LGcond-mat.dis-nncond-mat.stat-mechstat.ML
keywords networksbehaviorcomplexdynamicsmultilayerneuralneuronsnumber
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Can multilayer neural networks -- typically constructed as highly complex structures with many nonlinearly activated neurons across layers -- behave in a non-trivial way that yet simplifies away a major part of their complexities? In this work, we uncover a phenomenon in which the behavior of these complex networks -- under suitable scalings and stochastic gradient descent dynamics -- becomes independent of the number of neurons as this number grows sufficiently large. We develop a formalism in which this many-neurons limiting behavior is captured by a set of equations, thereby exposing a previously unknown operating regime of these networks. While the current pursuit is mathematically non-rigorous, it is complemented with several experiments that validate the existence of this behavior.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Information-theoretic reduction of deep neural networks to linear models in the overparametrized proportional regime

    math.ST 2025-05 conditional novelty 8.0 of 10

    Deep Bayesian neural networks in the proportional scaling regime are asymptotically equivalent, in mutual information and generalization error, to generalized linear models.

  2. The generalization error of random features regression: Precise asymptotics and double descent curve

    math.ST 2019-08 conditional novelty 8.0 of 10

    Mei and Montanari derive the exact asymptotic test error of random features ridge regression and show it reproduces the full double descent phenomenon without any misspecified structure.

  3. Convergence of Time-Averaged Mean Field Gradient Descent Dynamics for Continuous Multi-Player Zero-Sum Games

    math.OC 2025-05 conditional novelty 6.0 of 10

    An exponentially discounted time-averaged mean-field gradient descent flow reaches the entropy-regularized mixed Nash equilibrium exponentially fast, and its annealed version reaches the unregularized equilibrium at r...

Pith tools