Adam can achieve the accelerated momentum convergence rate locally on smooth strongly convex problems when its momentum and step size are tuned to the condition number, while RMSprop is shown to converge at the slower gradient descent rate.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
math.OC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Sharp higher order convergence rates for the Adam optimizer
Adam can achieve the accelerated momentum convergence rate locally on smooth strongly convex problems when its momentum and step size are tuned to the condition number, while RMSprop is shown to converge at the slower gradient descent rate.