Pith. sign in

REVIEW 2 cited by

A Policy Gradient Framework for Stochastic Optimal Control Problems with Global Convergence Guarantee

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.05816 v3 pith:SXJMJIEQ submitted 2023-02-11 math.OC cs.LGcs.SYeess.SY

classification math.OCcs.LGcs.SYeess.SY
keywords gradientcontrolconvergenceoptimalpolicycontinuousflowglobal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We consider policy gradient methods for stochastic optimal control problem in continuous time. In particular, we analyze the gradient flow for the control, viewed as a continuous time limit of the policy gradient method. We prove the global convergence of the gradient flow and establish a convergence rate under some regularity assumptions. The main novelty in the analysis is the notion of local optimal control function, which is introduced to characterize the local optimality of the iterate.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Convergence of Proximal Policy Gradient Method for Problems with Control Dependent Diffusion Coefficients

    math.OC 2025-05 conditional novelty 7.0 of 10

    For linear state dynamics with control-dependent diffusion, proximal policy gradient iterates converge linearly to a stationary control when the running or terminal cost is sufficiently strongly convex.

  2. Simulating Fokker-Planck equations via mean field control of score-based normalizing flows

    math.OC 2025-06 conditional novelty 4.0 of 10

    A mean field control formulation using score-based normalizing flows simulates Fokker-Planck equations deterministically, with a convergence theorem for Ornstein-Uhlenbeck processes and experiments on Langevin and cha...

Pith tools