REVIEW 3 cited by
Convergence of stochastic gradient descent schemes for Lojasiewicz-landscapes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this article, we consider convergence of stochastic gradient descent schemes (SGD), including momentum stochastic gradient descent (MSGD), under weak assumptions on the underlying landscape. More explicitly, we show that on the event that the SGD stays bounded we have convergence of the SGD if there is only a countable number of critical points or if the objective function satisfies Lojasiewicz-inequalities around all critical levels as all analytic functions do. In particular, we show that for neural networks with analytic activation function such as softplus, sigmoid and the hyperbolic tangent, SGD converges on the event of staying bounded, if the random variables modelling the signal and response in the training are compactly supported.
Forward citations
Cited by 3 Pith papers
-
Unified convergence analysis for gradient descent optimization methods in the training of deep neural networks
Bounded trajectories of a broad class of GD optimizers (Adam, RMSprop, NAG, Adan, etc.) converge with polynomial rates to critical points of KL objectives with locally Lipschitz gradients, covering analytic-activation...
-
Momentum-based minimization of the Ginzburg-Landau functional on Euclidean spaces and graphs
The accelerated Allen-Cahn equation formally converges to the hyperbolic interface law ∂_t v = (1-v^2)(h-αv), and a large-step FISTA discretization empirically accelerates Ginzburg-Landau minimization.
-
Mathematical analysis of the gradients in deep learning
For deep feedforward networks with piecewise-smooth activations, the autodiff gradient is shown to be the unique limit of gradients of smoothed activations, a limiting Frechet subgradient, and equal to the true gradie...
Discussion (0). Continue with ORCID to comment.