AdaGrad-Norm provably reaches ε error in O(log 1/ε) iterations for strongly convex and PL objectives from any initial step size, under new RUIG and zero-noise-at-optimum assumptions.
However, after tuning η = 10000 in stochastic setting and η = 100 in batch setting, the convergence rate of AdaGrad-Norm is better again
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
stat.ML 1years
2019 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Linear Convergence of Adaptive Stochastic Gradient Descent
AdaGrad-Norm provably reaches ε error in O(log 1/ε) iterations for strongly convex and PL objectives from any initial step size, under new RUIG and zero-noise-at-optimum assumptions.