Pith. sign in

REVIEW 3 cited by

Entropy annealing for policy mirror descent in continuous time and space

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.20250 v3 pith:FZOY3UOX submitted 2024-05-30 math.OC cs.LGmath.PR

classification math.OCcs.LGmath.PR
keywords entropypolicygradientregularizationdescentflowmirrorrate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Entropy regularization has been widely used in policy optimization algorithms to enhance exploration and the robustness of the optimal control; however it also introduces an additional regularization bias. This work quantifies the impact of entropy regularization on the convergence of policy gradient methods for stochastic exit time control problems. We analyze a continuous-time policy mirror descent dynamics, which updates the policy based on the gradient of an entropy-regularized value function and adjusts the strength of entropy regularization as the algorithm progresses. We prove that with a fixed entropy level, the mirror descent dynamics converges exponentially to the optimal solution of the regularized problem. We further show that when the entropy level decays at suitable polynomial rates, the annealed flow converges to the solution of the unregularized problem at a rate of $\mathcal O(1/S)$ for discrete action spaces and, under suitable conditions, at a rate of $\mathcal O(1/\sqrt{S})$ for general action spaces, with $S$ being the gradient flow running time. The technical challenge lies in analyzing the gradient flow in the infinite-dimensional space of Markov kernels for nonconvex objectives. This paper explains how entropy regularization improves policy optimization, even with the true gradient, from the perspective of convergence rate.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Policy Optimization for Continuous-time Linear-Quadratic Graphon Mean Field Games

    math.OC 2025-06 accept novelty 7.0 of 10

    A bilevel policy optimization algorithm for continuous-time linear-quadratic graphon mean field games converges linearly to best-response policies and globally to the Nash equilibrium.

  2. Mirror descent for constrained stochastic control problems

    math.OC 2025-06 conditional novelty 6.0 of 10

    Under uniform convexity of the Hamiltonian, continuous-time mirror descent converges linearly; under strong convexity relative to a Bregman divergence, it converges exponentially.

  3. Simulating Fokker-Planck equations via mean field control of score-based normalizing flows

    math.OC 2025-06 conditional novelty 4.0 of 10

    A mean field control formulation using score-based normalizing flows simulates Fokker-Planck equations deterministically, with a convergence theorem for Ornstein-Uhlenbeck processes and experiments on Langevin and cha...

Pith tools