Pith. sign in

REVIEW 1 cited by

Policy Gradient Method For Robust Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.07344 v1 pith:TA5EOQM2 submitted 2022-05-15 cs.LG

classification cs.LG
keywords policyrobustgradientmethodcomplexitygloballearningreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

This paper develops the first policy gradient method with global optimality guarantee and complexity analysis for robust reinforcement learning under model mismatch. Robust reinforcement learning is to learn a policy robust to model mismatch between simulator and real environment. We first develop the robust policy (sub-)gradient, which is applicable for any differentiable parametric policy class. We show that the proposed robust policy gradient method converges to the global optimum asymptotically under direct policy parameterization. We further develop a smoothed robust policy gradient method and show that to achieve an $\epsilon$-global optimum, the complexity is $\mathcal O(\epsilon^{-3})$. We then extend our methodology to the general model-free setting and design the robust actor-critic method with differentiable parametric policy class and value function. We further characterize its asymptotic convergence and sample complexity under the tabular setting. Finally, we provide simulation results to demonstrate the robustness of our methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Policy Gradient Learning for Distributionally Robust Markov Decision Processes under Wasserstein Ambiguity

    math.OC 2026-06 unverdicted novelty 7.0 of 10

    Wasserstein-robust finite-horizon MDP values admit exact directional-derivative policy-gradient recursions, reducing to a vector-valued gradient when the dual and transport optimizers are unique.

Pith tools