Pith. sign in

REVIEW 1 cited by

On the Convergence Rates of Policy Gradient Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.07443 v2 pith:QKFWQBCL submitted 2022-01-19 math.OC cs.LG

classification math.OCcs.LG
keywords policyconvergencemethodgradientdescentmethodsmirrorprojected
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We consider infinite-horizon discounted Markov decision problems with finite state and action spaces and study the convergence rates of the projected policy gradient method and a general class of policy mirror descent methods, all with direct parametrization in the policy space. First, we develop a theory of weak gradient-mapping dominance and use it to prove sharper sublinear convergence rate of the projected policy gradient method. Then we show that with geometrically increasing step sizes, a general class of policy mirror descent methods, including the natural policy gradient method and a projected Q-descent method, all enjoy a linear rate of convergence without relying on entropy or other strongly convex regularization. Finally, we also analyze the convergence rate of an inexact policy mirror descent method and estimate its sample complexity under a simple generative model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 13 citations worldwide. Full citation record

  1. Safe and Efficient Online Convex Optimization with Linear Budget Constraints and Partial Feedback

    math.OC 2024-12 conditional novelty 6.0 of 10

    SELO combines Lyapunov-style virtual queues with pessimistic linear-bandit estimates to reach O(√T) regret and zero cumulative constraint violation under unknown linear budgets and partial constraint feedback.

Pith tools