Pith. sign in

REVIEW 1 cited by

Nonlinear Two-Time-Scale Stochastic Approximation: Convergence and Finite-Time Performance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.01868 v3 pith:AZSQBSD3 submitted 2020-11-03 math.OC cs.LGcs.SYeess.SY

classification math.OCcs.LGcs.SYeess.SY
keywords stochasticapproximationconvergencefinite-timenonlineartwo-time-scaleanalysiscontrol
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Two-time-scale stochastic approximation, a generalized version of the popular stochastic approximation, has found broad applications in many areas including stochastic control, optimization, and machine learning. Despite its popularity, theoretical guarantees of this method, especially its finite-time performance, are mostly achieved for the linear case while the results for the nonlinear counterpart are very sparse. Motivated by the classic control theory for singularly perturbed systems, we study in this paper the asymptotic convergence and finite-time analysis of the nonlinear two-time-scale stochastic approximation. Under some fairly standard assumptions, we provide a formula that characterizes the rate of convergence of the main iterates to the desired solutions. In particular, we show that the method achieves a convergence in expectation at a rate $\mathcal{O}(1/k^{2/3})$, where $k$ is the number of iterations. The key idea in our analysis is to properly choose the two step sizes to characterize the coupling between the fast and slow-time-scale iterates.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Linear Speedup of Personalized Federated Reinforcement Learning with Shared Representations

    cs.LG 2024-11 conditional novelty 7.0 of 10

    The paper proves that personalized federated temporal-difference learning with a shared linear representation converges at rate O(1/(N^{2/3} T^{2/3})), yielding linear speedup in the number of agents under Markovian noise.

Pith tools