Pith. sign in

REVIEW 2 cited by

Mean-Field Neural ODEs via Relaxed Optimal Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.05475 v3 pith:HL7SQKQ2 submitted 2019-12-11 math.PR math.OCstat.ML

classification math.PRmath.OCstat.ML
keywords controlgradientneuraldeepderiveerrorlearningmean-field
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We develop a framework for the analysis of deep neural networks and neural ODE models that are trained with stochastic gradient algorithms. We do that by identifying the connections between control theory, deep learning and theory of statistical sampling. We derive Pontryagin's optimality principle and study the corresponding gradient flow in the form of Mean-Field Langevin dynamics (MFLD) for solving relaxed data-driven control problems. Subsequently, we study uniform-in-time propagation of chaos of time-discretised MFLD. We derive explicit convergence rate in terms of the learning rate, the number of particles/model parameters and the number of iterations of the gradient algorithm. In addition, we study the error arising when using a finite training data set and thus provide quantitive bounds on the generalisation error. Crucially, the obtained rates are dimension-independent. This is possible by exploiting the regularity of the model with respect to the measure over the parameter space.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Genericity of Polyak-Lojasiewicz Inequalities for Entropic Mean-Field Neural ODEs

    math.OC 2025-07 conditional novelty 7.0 of 10

    For a continuum-depth ResNet model with entropic regularization, an open dense set of initial feature-label distributions admits a unique stable minimizer and a local Polyak-Lojasiewicz inequality near it.

  2. Dropout and Random Gradient Masking Are Asymptotically Equivalent in Large ResNets

    stat.ML 2026-07 conditional novelty 6.0 of 10

    As ResNets grow deep and wide with fixed dropout rate, dropout training and random-gradient-masking training converge to the same limiting dynamics, and the common masking variants collapse to one limit.

Pith tools