REVIEW 2 cited by
Stochastic Recursive Momentum for Policy Gradient Methods
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
In this paper, we propose a novel algorithm named STOchastic Recursive Momentum for Policy Gradient (STORM-PG), which operates a SARAH-type stochastic recursive variance-reduced policy gradient in an exponential moving average fashion. STORM-PG enjoys a provably sharp $O(1/\epsilon^3)$ sample complexity bound for STORM-PG, matching the best-known convergence rate for policy gradient algorithm. In the mean time, STORM-PG avoids the alternations between large batches and small batches which persists in comparable variance-reduced policy gradient methods, allowing considerably simpler parameter tuning. Numerical experiments depicts the superiority of our algorithm over comparative policy gradient algorithms.
Forward citations
Cited by 2 Pith papers
-
Reusing Trajectories in Policy Gradients Enables Fast Convergence
Reusing past trajectories with a power-mean-corrected importance-weighting estimator gives policy-gradient methods a sample complexity of O~(epsilon^{-1}) in the full-reuse regime.
-
Variance-Reduced Conditional Gradient Methods under Markovian Sampling for Nonconvex Composite Optimization
MC-ALFCG attains eO((τmix^2 Gσ + τmix^{5/2} Gσ^2) ε^{-3} + τmix^5 ε^{-2}) expected sample complexity for the generalized Frank–Wolfe gap under a single ergodic Markovian stream.
Discussion (0). Continue with ORCID to comment.