Pith. sign in

REVIEW 1 cited by

Policy Gradient in Robust MDPs with Global Convergence Guarantee

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.10439 v2 pith:GI7JMQ6K submitted 2022-12-20 cs.LG

classification cs.LG
keywords policyrmdpsrobustgradientconvergencealgorithmsdrpgerrors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Robust Markov decision processes (RMDPs) provide a promising framework for computing reliable policies in the face of model errors. Many successful reinforcement learning algorithms build on variations of policy-gradient methods, but adapting these methods to RMDPs has been challenging. As a result, the applicability of RMDPs to large, practical domains remains limited. This paper proposes a new Double-Loop Robust Policy Gradient (DRPG), the first generic policy gradient method for RMDPs. In contrast with prior robust policy gradient algorithms, DRPG monotonically reduces approximation errors to guarantee convergence to a globally optimal policy in tabular RMDPs. We introduce a novel parametric transition kernel and solve the inner loop robust policy via a gradient-based method. Finally, our numerical results demonstrate the utility of our new algorithm and confirm its global convergence properties.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hybrid Cross-domain Robust Reinforcement Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    HYDRO combines a small offline robust RL dataset with a mismatched online simulator, filtering simulator samples by uncertainty and gap to the worst-case model to improve robust policy performance.

Pith tools