Pith. sign in

REVIEW 4 minor 142 references

Unregularized policy mirror descent with any constant step size converges to an optimal policy for a wide class of mirror maps, even when many optima exist.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 04:16 UTC pith:TCMVQURW

load-bearing objection Clean unification of policy (not just value) convergence for constant-step unregularized PMD under general decomposable Legendre maps, with a useful four-case taxonomy.

arxiv 2607.11626 v1 pith:TCMVQURW submitted 2026-07-13 math.OC

On the Policy Convergence of Policy Mirror Descent Methods

classification math.OC MSC 90C4090C2568T05
keywords policy mirror descentpolicy convergenceLegendre mirror mapsconstant step sizefinite-time terminationsoftmax natural policy gradientprojected Q-ascentMarkov decision process
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Value convergence for policy mirror descent has been known for years, but value can approach the optimum without the policy itself settling when multiple optimal policies exist. This paper proves that, for any decomposable Legendre-type mirror map of the form sum of a scalar function psi, unregularized PMD with a fixed step size produces a policy sequence that actually converges to one limiting optimal policy. The same theory covers the squared Euclidean map behind projected Q-ascent, the negative entropy behind softmax natural policy gradient, Tsallis, Hellinger, and Fermi-Dirac maps. The qualitative behavior is completely determined by whether psi is differentiable at the endpoints 0 and 1: finite-time termination, pure asymptotic convergence, or an MDP-dependent split between the two. When the map also has positive finite curvature, the paper supplies matching local rates for both value error and policy distance.

Core claim

For every decomposable mirror map h(p)=sum psi(p(a)) obeying the standard Legendre conditions, the unregularized PMD iteration with arbitrary constant step size generates a policy sequence that converges in the policy domain to a limiting optimal policy, even when the optimal set is not a singleton. Differentiability of psi at 0 and 1 partitions the maps into four regimes that respectively yield finite-time termination, asymptotic convergence, an MDP-dependent dichotomy, or asymptotic convergence with half-rate local bounds.

What carries the argument

The summability criterion (Lemma 3.1): once the value errors are summable, Bregman distances to any optimal policy form a quasi-Fejer sequence whose only accumulation point is a single optimal policy; local dual-gap quantities G_k then supply the missing summability via the boundary behavior of psi'.

Load-bearing premise

The map must send gradients to infinity at every boundary point of its domain; without that essential-smoothness condition the dual-gap growth arguments and the four-case classification fail.

What would settle it

Construct a finite discounted MDP and a decomposable Legendre map satisfying the paper's assumptions for which the constant-step PMD sequence has two distinct accumulation points that are both optimal; any such counter-example would refute the global policy-convergence claim.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 4 minor

Summary. The paper establishes policy-domain convergence of unregularized policy mirror descent (PMD) with arbitrary constant step sizes for finite discounted MDPs, under decomposable Legendre-type mirror maps h(p)=\sum_a \psi(p(a)) satisfying Assumption 1.1. The central claim is that the generated policy sequence converges to a limiting optimal policy even when the optimal set is non-singleton. Convergence regimes are classified by the one-sided differentiability of \psi at 0 and 1 (finite-time termination, asymptotic convergence, or an MDP-dependent dichotomy), with local value-error bounds and policy rates under additional C^{2} curvature assumptions. The argument proceeds via a summability criterion (Lemma 3.1), dual-gap evolution (Lemma 3.2), and an integral test (Lemma 3.3), with complete proofs in Sections 4–5 and supporting grid-world experiments.

Significance. If correct, this supplies the first systematic, unified policy-convergence theory for unregularized constant-step PMD across the standard family of decomposable Legendre maps (squared Euclidean/PQA, Shannon entropy/softmax NPG, Tsallis, Hellinger, Fermi–Dirac). Prior guarantees were either instance-specific (relying on closed-form updates) or required regularization/homotopy. The four-case taxonomy, the reduction of policy convergence to value-error summability, and the local rates under positive finite curvature are technically substantial and immediately useful for understanding which mirror maps produce finite-time versus asymptotic behavior. Full proofs and matching numerical illustrations strengthen the contribution.

minor comments (4)
  1. Table 1 is dense; a short caption sentence clarifying that “/” means “not applicable / finite-time so no local rate” would help readers scan the four cases.
  2. In the numerical section the bisection accuracy 10^{-13} and fixed \eta=0.1 are stated, but a one-line remark that the qualitative regimes are insensitive to modest changes in these parameters would improve reproducibility claims.
  3. A few typographical slips remain (e.g., “Inthis paper”, missing spaces after periods in the introduction). A final proof-reading pass would clean them.
  4. The comparison with Lin & Zhang (2022) L-coercivity is clear in Remark 3.1; adding a one-sentence pointer in the abstract or introduction that the present result removes that assumption for the Tsallis q\in(1,2) case would make the novelty more immediately visible.

Circularity Check

0 steps flagged

No significant circularity: self-contained Legendre/Bregman analysis with comparison-only self-citations

full rationale

This is a pure mathematical optimization paper. The central claim (policy convergence of unregularized constant-step PMD for decomposable Legendre maps under Assumption 1.1, even when Π* is non-singleton) is derived from first-order KKT conditions of the PMD update (Lemma 2.2), the three-point Bregman identity (Lemma 2.3), dual-gap evolution (Lemma 3.2), a general summability criterion reducing policy convergence to ∑∥V*−Vπk∥∞<∞ (Lemma 3.1), and case-by-case local bounds obtained by tracking ψ′-gaps and applying the integral inverse criterion (Lemma 3.3). All proofs are written out in Sections 4–5; no parameters are fitted to data, no uniqueness theorem is imported as an external black box that forces the result, and no ansatz is smuggled in. Self-citations (Li et al. on HPMD and softmax NPG; Liu et al. on PQA) appear only for comparison of prior special cases or regularized variants; the unregularized constant-step theory for general decomposable maps is developed independently and does not rest on those results as load-bearing lemmas. The four-case taxonomy is a classification by boundary differentiability of ψ, not a renaming of a known empirical pattern. Score 0 is therefore appropriate.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

Pure mathematical theory resting on standard convex-analysis and MDP facts; no free parameters are fitted to data. The only non-standard ingredients are the Legendre-type conditions on ψ and the finite discounted MDP setting, both stated explicitly.

axioms (4)
  • domain assumption ψ satisfies Assumption 1.1 (proper closed convex, essentially smooth and strictly convex on int dom ψ, dom ψ ⊇ [0,1])
    Invoked from the first statement of the main theorems; essential for Bregman divergence properties and dual-gap monotonicity.
  • domain assumption Finite state-action discounted MDP with rewards in [0,1] and γ∈[0,1)
    Standard setting of Section 2; used for bounded values, existence of optimal policies, and the sub-optimality gap Δ.
  • standard math Three-point identity and first-order optimality conditions for Bregman projections onto the simplex (Lemmas 2.2–2.3)
    Classical convex-analysis facts (Rockafellar, Beck) used to derive the dual update and descent inequality.
  • standard math Global O(1/k) value convergence of PMD under Legendre maps (Xiao 2022, Lemma 2.4)
    Imported to guarantee that the local entrance time T(σ) is finite; the paper’s novelty begins after this global bound.

pith-pipeline@v1.1.0-grok45 · 51512 in / 2417 out tokens · 25131 ms · 2026-07-14T04:16:36.838289+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of On the Policy Convergence of Policy Mirror Descent Methods." pith.science (2026). https://pith.science/paper/TCMVQURW

@misc{pith2026260711626,
  author       = {Pith},
  title        = {Pith review of: On the Policy Convergence of Policy Mirror Descent Methods},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TCMVQURW}},
  note         = {Machine review of arXiv:2607.11626}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We study the policy convergence of unregularized policy mirror descent (PMD) with arbitrary constant step sizes for finite discounted Markov decision processes. We focus on decomposable mirror maps of the form $h(p)=\sum_a \psi(p(a))$, where $\psi$ satisfies standard Legendre-type assumptions. Under these conditions, we prove that the policy sequence generated by PMD converges in the policy domain to a limiting optimal policy, even when the optimal policy set is not a singleton. This result covers a broad class of commonly used mirror maps, including the squared Euclidean mirror map underlying projected Q-ascent, the negative Shannon entropy underlying softmax natural policy gradient, Tsallis entropy, the Hellinger mapping, and the Fermi-Dirac entropy. Although policy convergence has been established previously for specific PMD instances or for regularized variants such as homotopic PMD, to the best of our knowledge, this is the first systematic and unified policy convergence theory for unregularized PMD under general decomposable mirror maps and arbitrary constant step sizes. Our analysis further reveals that the convergence behavior is governed by the differentiability of $\psi$ at $0$ and $1$, leading to different behaviors, including finite-time termination, asymptotic convergence, and an MDP-dependent dichotomy. When $\psi$ is twice continuously differentiable with strictly positive finite curvature, we further establish local policy convergence rates for the asymptotic convergence cases, covering the standard mirror maps mentioned above.

Figures

Figures reproduced from arXiv: 2607.11626 by Ke Wei, Wenye Li.

Figure 1
Figure 1. Figure 1: Two Grid World MDPs that used in our experiment setting. [PITH_FULL_IMAGE:figures/full_fig_p017_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The plot of ∥πk+1 − πk∥2 for Case 1, tested on Grid World A. Tsallis entropy with q = 1/2, and use πlast := πT as an approximation to the limiting policy π∞. For the two representative states s1 = 3 ∈ S=1 and s2 = 7 ∈ S>1, we plot the curves of log [PITH_FULL_IMAGE:figures/full_fig_p019_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: The curves of log ∥πk(·|s) − πlast(·|s)∥2 for Case 2, tested on Grid World A (Figure 1a). The blue curve shows the result on state s1 = 3 ∈ S=1 and the orange one corresponds to state s2 = 7 ∈ S>1. Policy convergence dichotomy of Case 3. Since the policy convergence behavior in Case 3 differs depending on whether S>1 = S or S>1 ̸= S, we test PMD with the Hellinger mapping ψ(x) = − √ 1 − x 2 on both Grid Wo… view at source ↗
Figure 4
Figure 4. Figure 4: Policy convergence results for PMD with Hellinger mapping on Grid Worlds A (left) and B (right). [PITH_FULL_IMAGE:figures/full_fig_p020_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Policy convergence result for PMD with Fermi-Dirac entropy on Grid World A. [PITH_FULL_IMAGE:figures/full_fig_p020_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

142 extracted references · 15 linked inside Pith

  1. [1]

    Nature , volume=

    Human-level control through deep reinforcement learning , author=. Nature , volume=

  2. [2]

    Mastering the game of

    Silver, David and Huang, Aja and Maddison, Chris J and Guez, Arthur and Sifre, Laurent and Van Den Driessche, George and Schrittwieser, Julian and Antonoglou, Ioannis and Panneershelvam, Veda and Lanctot, Marc and Dieleman, Sander and Grewe, Dominik and Nham, John and Kalchbrenner, Nal and Sutskever, Ilya and Lillicrap, Timothy and Leach, Madeleine and Ka...

  3. [3]

    2019 , journal=

    Dota 2 with Large Scale Deep Reinforcement Learning , author=. 2019 , journal=

  4. [4]

    arxiv:2404.03372 , year =

    Liu, Jiacai and Li, Wenye and Wei, Ke , title =. arxiv:2404.03372 , year =

  5. [5]

    Starcraft

    Vinyals, Oriol and others , journal=. Starcraft

  6. [6]

    Science Robotics , year=

    Learning agile and dynamic motor skills for legged robots , author=. Science Robotics , year=

  7. [7]

    Science Robotics , volume =

    Lee, Joonho and Hwangbo, Jemin and Wellhausen, Lorenz and Koltun, Vladlen and Hutter, Marco , title =. Science Robotics , volume =

  8. [8]

    Learning robust perceptive locomotion for quadrupedal robots in the wild , volume =

    Miki, Takahiro and Lee, Joonho and Hwangbo, Jemin and Wellhausen, Lorenz and Koltun, Vladlen and Hutter, Marco , year =. Learning robust perceptive locomotion for quadrupedal robots in the wild , volume =

  9. [9]

    arXiv:1606.03966 , year=

    Making Contextual Decisions with Low Technical Debt , author=. arXiv:1606.03966 , year=

  10. [10]

    , booktitle=

    Chen, Minmin and Beutel, Alex and Covington, Paul and Jain, Sagar and Belletti, Francois and Chi, Ed H. , booktitle=. Top-k off-policy correction for a

  11. [11]

    Nature , volume=

    A graph placement methodology for fast chip design , author=. Nature , volume=

  12. [12]

    Advances in Neural Information Processing Systems , year=

    Policy Gradient Methods for Reinforcement Learning with Function Approximation , author=. Advances in Neural Information Processing Systems , year=

  13. [13]

    2019 , pages=

    A Theory of Regularized Markov Decision Processes , author=. 2019 , pages=

  14. [14]

    Adaptive Trust Region Policy Optimization:

    Shani, Lior and Efroni, Yonathan and Mannor, Shie , year=. Adaptive Trust Region Policy Optimization:. Proceedings of the AAAI Conference on Artificial Intelligence , pages=

  15. [15]

    On the Convergence Rates of Policy Gradient Methods , journal=

    Xiao, Lin , year=. On the Convergence Rates of Policy Gradient Methods , journal=

  16. [16]

    Mathematical Programming , author=

    Policy Mirror Descent for Reinforcement Learning:. Mathematical Programming , author=. 2021 , pages=

  17. [17]

    Escaping the Gravitational Pull of Softmax , booktitle=

    Mei, Jincheng and Xiao, Chenjun and Dai, Bo and Li, Lihong and Szepesvári, Csaba and Schuurmans, Dale , year=. Escaping the Gravitational Pull of Softmax , booktitle=

  18. [18]

    2023 , booktitle=

    Optimal Convergence Rate for Exact Policy Mirror Descent in Discounted Markov Decision Processes , author=. 2023 , booktitle=

  19. [19]

    Fast Global Convergence of Natural Policy Gradient Methods with Entropy Regularization , journal=

    Cen, Shicong and Cheng, Chen and Chen, Yuxin and Wei, Yuting and Chi, Yuejie , year=. Fast Global Convergence of Natural Policy Gradient Methods with Entropy Regularization , journal=

  20. [20]

    Mathematical Programming , author=

    Homotopic policy mirror descent:. Mathematical Programming , author=

  21. [21]

    SIAM Journal on Optimization , author=

    Policy Mirror Descent for Regularized Reinforcement Learning:. SIAM Journal on Optimization , author=. 2023 , volume =

  22. [22]

    Journal of Machine Learning Research , author=

    On the Theory of Policy Gradient Methods:. Journal of Machine Learning Research , author=. 2021 , volume=

  23. [23]

    On the Global Convergence Rates of Softmax Policy Gradient Methods , booktitle=

    Mei, Jincheng and Xiao, Chenjun and Szepesvári, Csaba and Schuurmans, Dale , year=. On the Global Convergence Rates of Softmax Policy Gradient Methods , booktitle=

  24. [24]

    2023 , booktitle=

    Linear Convergence of Natural Policy Gradient Methods with Log-Linear Policies , author=. 2023 , booktitle=

  25. [25]

    2022 , journal=

    Linear Convergence for Natural Policy Gradient with Log-linear Policy Parametrization , author=. 2022 , journal=

  26. [26]

    On the Linear Convergence of Natural Policy Gradient Algorithm , booktitle=

    Khodadadian, Sajad and Jhunjhunwala, Prakirt Raj and Varma, Sushil Mahavir and Maguluri, Siva Theja , year=. On the Linear Convergence of Natural Policy Gradient Algorithm , booktitle=

  27. [27]

    2020 , pages =

    Advances in Neural Information Processing Systems , author=. 2020 , pages =

  28. [28]

    International Conference on Artificial Intelligence and Statistics , author=

    On the Linear Convergence of Policy Gradient Methods for Finite. International Conference on Artificial Intelligence and Statistics , author=. 2021 , pages=

  29. [29]

    , year =

    Wang, Weiran and Carreira-Perpiñán, Miguel A. , year =. Projection onto the probability simplex:

  30. [30]

    International Conference on Machine Learning , pages=

    Approximately Optimal Approximate Reinforcement Learning , author=. International Conference on Machine Learning , pages=

  31. [31]

    On the Linear Convergence of Policy Gradient under

    Liu, Jiacai and Chen, Jinchi and Wei, Ke , journal=. On the Linear Convergence of Policy Gradient under

  32. [32]

    , title =

    Bertsekas, Dimitri P. , title =

  33. [33]

    Improved and Generalized Upper Bounds on the Complexity of Policy Iteration , volume =

    Scherrer, Bruno , booktitle =. Improved and Generalized Upper Bounds on the Complexity of Policy Iteration , volume =

  34. [34]

    Journal of Machine Learning Research , year=

    On the Convergence of Projected Policy Gradient for Any Constant Step Sizes , author=. Journal of Machine Learning Research , year=

  35. [35]

    arXiv:1206.2459 , year=

    Renyi Divergence and Kullback-Leibler Divergence , author=. arXiv:1206.2459 , year=

  36. [36]

    International Conference on Machine Learning , year =

    Mei, Jincheng and Gao, Yue and Dai, Bo and Szepesvári, Csaba and Schuurmans, Dale , title=. International Conference on Machine Learning , year =

  37. [37]

    Machine Learning , year =

    Simple statistical gradient-following algorithms for connectionist reinforcement learning , author =. Machine Learning , year =

  38. [38]

    Sutton and Andrew G

    Richard S. Sutton and Andrew G. Barto , title =

  39. [39]

    A natural policy gradient , year =

    Sham Kakade , booktitle =. A natural policy gradient , year =

  40. [40]

    International conference on machine learning , pages=

    Trust region policy optimization , author=. International conference on machine learning , pages=

  41. [41]

    arXiv:1707.06347 , year=

    Proximal policy optimization algorithms , author=. arXiv:1707.06347 , year=

  42. [42]

    Operations Research , year =

    Jalj Bhandari and Daniel Russo , title =. Operations Research , year =

  43. [43]

    AAAI Conference on Artifical Intelligence , year =

    Lior Shani and Yonathan Efroni and Shie Mannor , title =. AAAI Conference on Artifical Intelligence , year =

  44. [44]

    Advances in Neural Information Processing Systems , year =

    Carlo Alfano and Rui Yuan and Patrick Rebeschini , title =. Advances in Neural Information Processing Systems , year =

  45. [45]

    Mathematical Programming , year =

    Gen Li and Yuting Wei and Yuejie Chi and Yuxin Chen , title =. Mathematical Programming , year =

  46. [46]

    Puterman and Shelby L

    Martin L. Puterman and Shelby L. Brumelle , title =. Mathematics of Operations Research , year =

  47. [47]

    Advances in Neural Information Processing Systems , year =

    Ofir Nachum and Mohammad Norouzi and Kelvin Xu and Dale Schuurmans , title =. Advances in Neural Information Processing Systems , year =

  48. [48]

    Nature , author=

    Discovering faster matrix multiplication algorithms with reinforcement learning , volume=. Nature , author=. 2022 , pages=

  49. [49]

    Nature , year=

    Mastering the game of Go with deep neural networks and tree search , author=. Nature , year=

  50. [50]

    Nature , year=

    Highly accurate protein structure prediction with AlphaFold , volume=. Nature , year=

  51. [51]

    Puterman, Martin L , title=

  52. [52]

    Bellman, Richard , title =

  53. [53]

    Operations Research Letters , volume=

    Mirror descent and nonlinear projected subgradient methods for convex optimization , author=. Operations Research Letters , volume=. 2003 , publisher=

  54. [54]

    arXiv:1509.02971 , year=

    Continuous control with deep reinforcement learning , author=. arXiv:1509.02971 , year=

  55. [55]

    Advances in neural information processing systems , year=

    Policy gradient methods for reinforcement learning with function approximation , author=. Advances in neural information processing systems , year=

  56. [56]

    arxiv:2005.09814 , year =

    Manan Tomar and Lior Shani and Yonathan Efroni and Mohammad Ghavamzadeh , title =. arxiv:2005.09814 , year =

  57. [57]

    arxiv:1506.02438 , year =

    John Schulman and Philipp Moritz and Sergey Levine and Michael Jordan and Pieter Abbeel , title =. arxiv:1506.02438 , year =

  58. [58]

    International Conference on Machine Learning , pages=

    Stochastic gradient succeeds for bandits , author=. International Conference on Machine Learning , pages=. 2023 , organization=

  59. [59]

    arXiv preprint arXiv:2405.13136 , year=

    Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs , author=. arXiv preprint arXiv:2405.13136 , year=

  60. [60]

    2019 IEEE 58th Conference on Decision and Control (CDC) , pages=

    Convergence and iteration complexity of policy gradient method for infinite-horizon reinforcement learning , author=. 2019 IEEE 58th Conference on Decision and Control (CDC) , pages=. 2019 , organization=

  61. [61]

    SIAM Journal on Control and Optimization , volume=

    Global convergence of policy gradient methods to (almost) locally optimal policies , author=. SIAM Journal on Control and Optimization , volume=. 2020 , publisher=

  62. [62]

    arXiv preprint arXiv:2411.04913 , year=

    Structure Matters: Dynamic Policy Gradient , author=. arXiv preprint arXiv:2411.04913 , year=

  63. [63]

    Advances in neural information processing systems , volume=

    Policy search by dynamic programming , author=. Advances in neural information processing systems , volume=

  64. [64]

    2003 , publisher=

    On the sample complexity of reinforcement learning , author=. 2003 , publisher=

  65. [65]

    SIAM Journal on Optimization , volume=

    Convergence analysis of a proximal-like minimization algorithm using Bregman functions , author=. SIAM Journal on Optimization , volume=. 1993 , publisher=

  66. [66]

    International Conference on Machine Learning , pages=

    Safe reinforcement learning using advantage-based intervention , author=. International Conference on Machine Learning , pages=. 2021 , organization=

  67. [67]

    Advances in Neural Information Processing Systems , volume=

    Policy improvement via imitation of multiple oracles , author=. Advances in Neural Information Processing Systems , volume=

  68. [68]

    arXiv preprint arXiv:2108.05828 , year=

    A functional mirror ascent view of policy gradient methods with function approximation , author=. arXiv preprint arXiv:2108.05828 , year=

  69. [69]

    Advances in Neural Information Processing Systems , volume=

    An operator view of policy gradient methods , author=. Advances in Neural Information Processing Systems , volume=

  70. [70]

    International Conference on Learning Representations , year=

    Linear convergence of natural policy gradient methods with log-linear policies , author=. International Conference on Learning Representations , year=

  71. [71]

    International Conference on Learning Representations , year=

    Mirror descent policy optimization , author=. International Conference on Learning Representations , year=

  72. [72]

    International Conference on Artificial Intelligence and Statistics , pages=

    A general sample complexity analysis of vanilla policy gradient , author=. International Conference on Artificial Intelligence and Statistics , pages=. 2022 , organization=

  73. [73]

    arXiv preprint arXiv:2407.16602 , year=

    Functional Acceleration for Policy Mirror Descent , author=. arXiv preprint arXiv:2407.16602 , year=

  74. [74]

    1998 , publisher=

    Introduction to reinforcement learning , author=. 1998 , publisher=

  75. [75]

    International Conference on Machine Learning , pages=

    Stochastic policy gradient methods: Improved sample complexity for fisher-non-degenerate policies , author=. International Conference on Machine Learning , pages=. 2023 , organization=

  76. [76]

    International Conference on Artificial Intelligence and Statistics , pages=

    Improved sample complexity analysis of natural policy gradient algorithm with general parameterization for infinite horizon discounted reward markov decision processes , author=. International Conference on Artificial Intelligence and Statistics , pages=. 2024 , organization=

  77. [77]

    Journal of Scientific Computing , volume=

    Global Convergence of Natural Policy Gradient with Hessian-Aided Momentum Variance Reduction , author=. Journal of Scientific Computing , volume=. 2024 , publisher=

  78. [78]

    Advances in Neural Information Processing Systems , volume=

    An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods , author=. Advances in Neural Information Processing Systems , volume=

  79. [79]

    Mathematics of Operations Research , year=

    Stochastic first-order methods for average-reward markov decision processes , author=. Mathematics of Operations Research , year=

  80. [80]

    Advances in Neural Information Processing Systems , volume=

    Policy Mirror Descent with Lookahead , author=. Advances in Neural Information Processing Systems , volume=

Showing first 80 references.