Pith. sign in

Dissipativity Theory for Accelerating Stochastic Variance Reduction: A Unified Analysis of SVRG and Katyusha Using Semidefinite Programs

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Techniques for reducing the variance of gradient estimates used in stochastic programming algorithms for convex finite-sum problems have received a great deal of attention in recent years. By leveraging dissipativity theory from control, we provide a new perspective on two important variance-reduction algorithms: SVRG and its direct accelerated variant Katyusha. Our perspective provides a physically intuitive understanding of the behavior of SVRG-like methods via a principle of energy conservation. The tools discussed here allow us to automate the convergence analysis of SVRG-like methods by capturing their essential properties in small semidefinite programs amenable to standard analysis and computational techniques. Our approach recovers existing convergence results for SVRG and Katyusha and generalizes the theory to alternative parameter choices. We also discuss how our approach complements the linear coupling technique. Our combination of perspectives leads to a better understanding of accelerated variance-reduced stochastic methods for finite-sum problems.

fields

cs.LG 1

years

2019 1

verdicts

CONDITIONAL 1

representative citing papers

Almost Tune-Free Variance Reduction

cs.LG · 2019-08-25 · conditional · novelty 6.0

SVRG and SARAH converge with a weighted averaging scheme driven by estimate sequences, and when combined with Barzilai-Borwein step sizes and an adaptive inner-loop rule they become almost tune-free in numerical tests.

citing papers explorer

Showing 1 of 1 citing paper.

  • Almost Tune-Free Variance Reduction cs.LG · 2019-08-25 · conditional · none · ref 2018 · internal anchor

    SVRG and SARAH converge with a weighted averaging scheme driven by estimate sequences, and when combined with Barzilai-Borwein step sizes and an adaptive inner-loop rule they become almost tune-free in numerical tests.