Pith. sign in

REVIEW 2 cited by

Non asymptotic analysis of Adaptive stochastic gradient algorithms and applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.01370 v1 pith:PEP65JEE submitted 2023-03-01 math.OC math.PRmath.STstat.MLstat.TH

classification math.OCmath.PRmath.STstat.MLstat.TH
keywords algorithmsstochasticgradientadaptiveadagradasymptoticlinearnewton
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In stochastic optimization, a common tool to deal sequentially with large sample is to consider the well-known stochastic gradient algorithm. Nevertheless, since the stepsequence is the same for each direction, this can lead to bad results in practice in case of ill-conditionned problem. To overcome this, adaptive gradient algorithms such that Adagrad or Stochastic Newton algorithms should be prefered. This paper is devoted to the non asymptotic analyis of these adaptive gradient algorithms for strongly convex objective. All the theoretical results will be adapted to linear regression and regularized generalized linear model for both Adagrad and Stochastic Newton algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On MUON optimization: From non-convergence to an error analysis with Polar Express and the Newton-Schulz polynomial from implementations

    math.OC 2026-08 accept novelty 8.0 of 10

    For a one-dimensional quadratic stochastic optimization problem, MUON with Newton-Schulz steps provably fails to converge to the minimizer for all sufficiently large mini-batch sizes when the data is skewed, while a n...

  2. Unified convergence analysis for gradient descent optimization methods in the training of deep neural networks

    math.OC 2026-07 accept novelty 7.0 of 10

    Bounded trajectories of a broad class of GD optimizers (Adam, RMSprop, NAG, Adan, etc.) converge with polynomial rates to critical points of KL objectives with locally Lipschitz gradients, covering analytic-activation...

Pith tools