Pith. sign in

REVIEW 3 cited by

Convergence of stochastic gradient descent schemes for Lojasiewicz-landscapes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.09385 v3 pith:3EZBNO2M submitted 2021-02-16 cs.LG math.PRmath.STstat.TH

classification cs.LGmath.PRmath.STstat.TH
keywords convergencedescentgradientstochasticanalyticboundedcriticalevent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this article, we consider convergence of stochastic gradient descent schemes (SGD), including momentum stochastic gradient descent (MSGD), under weak assumptions on the underlying landscape. More explicitly, we show that on the event that the SGD stays bounded we have convergence of the SGD if there is only a countable number of critical points or if the objective function satisfies Lojasiewicz-inequalities around all critical levels as all analytic functions do. In particular, we show that for neural networks with analytic activation function such as softplus, sigmoid and the hyperbolic tangent, SGD converges on the event of staying bounded, if the random variables modelling the signal and response in the training are compactly supported.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unified convergence analysis for gradient descent optimization methods in the training of deep neural networks

    math.OC 2026-07 accept novelty 7.0 of 10

    Bounded trajectories of a broad class of GD optimizers (Adam, RMSprop, NAG, Adan, etc.) converge with polynomial rates to critical points of KL objectives with locally Lipschitz gradients, covering analytic-activation...

  2. Momentum-based minimization of the Ginzburg-Landau functional on Euclidean spaces and graphs

    math.AP 2024-12 conditional novelty 7.0 of 10

    The accelerated Allen-Cahn equation formally converges to the hyperbolic interface law ∂_t v = (1-v^2)(h-αv), and a large-step FISTA discretization empirically accelerates Ginzburg-Landau minimization.

  3. Mathematical analysis of the gradients in deep learning

    cs.LG 2025-01 accept novelty 6.0 of 10

    For deep feedforward networks with piecewise-smooth activations, the autodiff gradient is shown to be the unique limit of gradients of smoothed activations, a limiting Frechet subgradient, and equal to the true gradie...

Pith tools