Pith. sign in

REVIEW 1 cited by

A Unified Analysis of AdaGrad with Weighted Aggregation and Momentum Acceleration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1808.03408 v4 pith:USEYIGM3 submitted 2018-08-10 cs.LG cs.NAmath.NAmath.OCstat.ML

classification cs.LGcs.NAmath.NAmath.OCstat.ML
keywords momentumadagradadamlearningadaptiveadausmrmsproprate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Integrating adaptive learning rate and momentum techniques into SGD leads to a large class of efficiently accelerated adaptive stochastic algorithms, such as AdaGrad, RMSProp, Adam, AccAdaGrad, \textit{etc}. In spite of their effectiveness in practice, there is still a large gap in their theories of convergences, especially in the difficult non-convex stochastic setting. To fill this gap, we propose \emph{weighted AdaGrad with unified momentum}, dubbed AdaUSM, which has the main characteristics that (1) it incorporates a unified momentum scheme which covers both the heavy ball momentum and the Nesterov accelerated gradient momentum; (2) it adopts a novel weighted adaptive learning rate that can unify the learning rates of AdaGrad, AccAdaGrad, Adam, and RMSProp. Moreover, when we take polynomially growing weights in AdaUSM, we obtain its $\mathcal{O}(\log(T)/\sqrt{T})$ convergence rate in the non-convex stochastic setting. We also show that the adaptive learning rates of Adam and RMSProp correspond to taking exponentially growing weights in AdaUSM, thereby providing a new perspective for understanding Adam and RMSProp. Lastly, comparative experiments of AdaUSM against SGD with momentum, AdaGrad, AdaEMA, Adam, and AMSGrad on various deep learning models and datasets are also carried out.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Convergence of Momentum-Based Optimization Algorithms with Time-Varying Parameters

    math.OC 2025-06 conditional novelty 6.0 of 10

    A unified momentum-based optimization algorithm with time-varying parameters is shown to converge almost surely under generalized Robbins-Monro and Kiefer-Wolfowitz-Blum conditions, even with biased, unbounded-varianc...

Pith tools