Pith. sign in

REVIEW 2 cited by

AdaGrad stepsizes: Sharp convergence over nonconvex landscapes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.01811 v8 pith:KDL23JR7 submitted 2018-06-05 stat.ML cs.LG

classification stat.MLcs.LG
keywords adagradconvergencegradientstochasticadagrad-normguaranteesdescentexperiments
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Adaptive gradient methods such as AdaGrad and its variants update the stepsize in stochastic gradient descent on the fly according to the gradients received along the way; such methods have gained widespread use in large-scale optimization for their ability to converge robustly, without the need to fine-tune the stepsize schedule. Yet, the theoretical guarantees to date for AdaGrad are for online and convex optimization. We bridge this gap by providing theoretical guarantees for the convergence of AdaGrad for smooth, nonconvex functions. We show that the norm version of AdaGrad (AdaGrad-Norm) converges to a stationary point at the $\mathcal{O}(\log(N)/\sqrt{N})$ rate in the stochastic setting, and at the optimal $\mathcal{O}(1/N)$ rate in the batch (non-stochastic) setting -- in this sense, our convergence guarantees are 'sharp'. In particular, the convergence of AdaGrad-Norm is robust to the choice of all hyper-parameters of the algorithm, in contrast to stochastic gradient descent whose convergence depends crucially on tuning the step-size to the (generally unknown) Lipschitz smoothness constant and level of stochastic noise on the gradient. Extensive numerical experiments are provided to corroborate our theory; moreover, the experiments suggest that the robustness of AdaGrad-Norm extends to state-of-the-art models in deep learning, without sacrificing generalization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Linear Convergence of Adaptive Stochastic Gradient Descent

    stat.ML 2019-08 conditional novelty 8.0 of 10

    AdaGrad-Norm provably reaches ε error in O(log 1/ε) iterations for strongly convex and PL objectives from any initial step size, under new RUIG and zero-noise-at-optimum assumptions.

  2. Observability conditions for neural state-space models with eigenvalues and their roots of unity

    cs.LG 2025-04 reject novelty 5.0 of 10

    A set of sufficient conditions and training losses for enforcing observability in neural state-space models, with one clean Mamba condition and several unproven high-probability Fourier results.

Pith tools