Pith. sign in

REVIEW 2 cited by

Stochastic Recursive Momentum for Policy Gradient Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.04302 v1 pith:PNAEISQ4 submitted 2020-03-09 stat.ML cs.LG

classification stat.MLcs.LG
keywords gradientpolicystorm-pgalgorithmrecursivestochasticbatchesmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

In this paper, we propose a novel algorithm named STOchastic Recursive Momentum for Policy Gradient (STORM-PG), which operates a SARAH-type stochastic recursive variance-reduced policy gradient in an exponential moving average fashion. STORM-PG enjoys a provably sharp $O(1/\epsilon^3)$ sample complexity bound for STORM-PG, matching the best-known convergence rate for policy gradient algorithm. In the mean time, STORM-PG avoids the alternations between large batches and small batches which persists in comparable variance-reduced policy gradient methods, allowing considerably simpler parameter tuning. Numerical experiments depicts the superiority of our algorithm over comparative policy gradient algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reusing Trajectories in Policy Gradients Enables Fast Convergence

    cs.LG 2025-06 conditional novelty 7.0 of 10

    Reusing past trajectories with a power-mean-corrected importance-weighting estimator gives policy-gradient methods a sample complexity of O~(epsilon^{-1}) in the full-reuse regime.

  2. Variance-Reduced Conditional Gradient Methods under Markovian Sampling for Nonconvex Composite Optimization

    math.OC 2026-07 accept novelty 6.0 of 10

    MC-ALFCG attains eO((τmix^2 Gσ + τmix^{5/2} Gσ^2) ε^{-3} + τmix^5 ε^{-2}) expected sample complexity for the generalized Frank–Wolfe gap under a single ergodic Markovian stream.

Pith tools