Pith. sign in

However, after tuning η = 10000 in stochastic setting and η = 100 in batch setting, the convergence rate of AdaGrad-Norm is better again

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

citation-role summary

background 1

citation-polarity summary

fields

stat.ML 1

years

2019 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Linear Convergence of Adaptive Stochastic Gradient Descent

stat.ML · 2019-08-28 · conditional · novelty 8.0

AdaGrad-Norm provably reaches ε error in O(log 1/ε) iterations for strongly convex and PL objectives from any initial step size, under new RUIG and zero-noise-at-optimum assumptions.

citing papers explorer

Showing 1 of 1 citing paper.

  • Linear Convergence of Adaptive Stochastic Gradient Descent stat.ML · 2019-08-28 · conditional · none · ref 2

    AdaGrad-Norm provably reaches ε error in O(log 1/ε) iterations for strongly convex and PL objectives from any initial step size, under new RUIG and zero-noise-at-optimum assumptions.