Pith. sign in

REVIEW 1 cited by

Policy Optimization with Stochastic Mirror Descent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.10462 v5 pith:XKHVKM3H submitted 2019-06-25 cs.LG stat.ML

classification cs.LGstat.ML
keywords policysamplemathttvrmpogradientdescentefficiencyepsilon
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Improving sample efficiency has been a longstanding goal in reinforcement learning. This paper proposes $\mathtt{VRMPO}$ algorithm: a sample efficient policy gradient method with stochastic mirror descent. In $\mathtt{VRMPO}$, a novel variance-reduced policy gradient estimator is presented to improve sample efficiency. We prove that the proposed $\mathtt{VRMPO}$ needs only $\mathcal{O}(\epsilon^{-3})$ sample trajectories to achieve an $\epsilon$-approximate first-order stationary point, which matches the best sample complexity for policy optimization. The extensive experimental results demonstrate that $\mathtt{VRMPO}$ outperforms the state-of-the-art policy gradient methods in various settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods

    cs.LG 2019-08 conditional novelty 6.0 of 10

    A trajectory-wise control variate estimator removes both action-level and future-trajectory variance in policy gradients, and the natural time-ordering is proven optimal under exact critic assumptions.

Pith tools