Pith. sign in

REVIEW 4 cited by

Two-Timescale Linear Stochastic Approximation: Constant Stepsizes Go a Long Way

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.13067 v1 pith:5MCHTVTL submitted 2024-10-16 eess.SY cs.LGcs.SYmath.OC

Two-Timescale Linear Stochastic Approximation: Constant Stepsizes Go a Long Way

classification eess.SY cs.LGcs.SYmath.OC
keywords betaalphaconstantstepsizesvarianceapproximationbiasiterate
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Previous studies on two-timescale stochastic approximation (SA) mainly focused on bounding mean-squared errors under diminishing stepsize schemes. In this work, we investigate {\it constant} stpesize schemes through the lens of Markov processes, proving that the iterates of both timescales converge to a unique joint stationary distribution in Wasserstein metric. We derive explicit geometric and non-asymptotic convergence rates, as well as the variance and bias introduced by constant stepsizes in the presence of Markovian noise. Specifically, with two constant stepsizes $\alpha < \beta$, we show that the biases scale linearly with both stepsizes as $\Theta(\alpha)+\Theta(\beta)$ up to higher-order terms, while the variance of the slower iterate (resp., faster iterate) scales only with its own stepsize as $O(\alpha)$ (resp., $O(\beta)$). Unlike previous work, our results require no additional assumptions such as $\beta^2 \ll \alpha$ nor extra dependence on dimensions. These fine-grained characterizations allow tail-averaging and extrapolation techniques to reduce variance and bias, improving mean-squared error bound to $O(\beta^4 + \frac{1}{t})$ for both iterates.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Shuffling the Data, Stretching the Step-size: Sharper Bias in constant step-size SGD

    math.OC 2026-04 unverdicted novelty 7.0

    Combining random reshuffling and Richardson-Romberg extrapolation yields cubic bias refinement and better MSE for constant-step SGD on structured non-monotone variational inequalities.

  2. SGD at the Edge of Stability: Stochastic Stabilization with Large Learning Rates

    stat.ML 2026-06 unverdicted novelty 6.0

    SGD on multiclass cross-entropy loss alternates between curvature-driven oscillations and stable regimes but self-stabilizes to enable best-iterate convergence with large learning rates for linear and two-layer models.

  3. Revisiting the Constant Stepsize Stochastic Approximation with Decision-Dependent Markovian Noise

    math.OC 2026-04 unverdicted novelty 6.0

    Constant stepsize SA with decision-dependent Markovian noise has stationary bias O(alpha) under Poisson-Gateaux differentiability, plus finite-time moment bounds and weak convergence.

  4. Finite-Time Decoupled Convergence in Nonlinear Two-Time-Scale Stochastic Approximation

    math.OC 2024-01 unverdicted novelty 6.0

    Under nested local linearity, nonlinear two-time-scale SA achieves finite-time decoupled convergence; nonlinearity in the slow update alone can destroy it.