Pith. sign in

REVIEW 1 cited by

Convergence of Policy Mirror Descent Beyond Compatible Function Approximation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.11033 v3 pith:RW5YQGKU submitted 2025-02-16 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords policyconvergenceclassesclosureconditionsdescentenvironmentsmirror
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern policy optimization methods roughly follow the policy mirror descent (PMD) algorithmic template, for which there are by now numerous theoretical convergence results. However, most of these either target tabular environments, or can be applied effectively only when the class of policies being optimized over satisfies strong closure conditions, which is typically not the case when working with parametric policy classes in large-scale environments. In this work, we develop a theoretical framework for PMD for general policy classes where we replace the closure conditions with a strictly weaker variational gradient dominance assumption, and obtain upper bounds on the rate of convergence to the best-in-class policy. Our main result leverages a novel notion of smoothness with respect to a local norm induced by the occupancy measure of the current policy, and casts PMD as a particular instance of smooth non-convex optimization in non-Euclidean space.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Convergence and Sample Complexity of First-Order Methods for Agnostic Reinforcement Learning

    cs.LG 2025-07 reject novelty 6.0 of 10

    Under variational gradient dominance, the paper derives state-space-independent sample complexity bounds for SDPO, CPI, DA-CPI, and PMD in agnostic policy learning.

Pith tools