Pith. sign in

REVIEW 2 cited by

Convergence Analysis for Entropy-Regularized Control Problems: A Probabilistic Approach

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.10959 v4 pith:G66VHO32 submitted 2024-06-16 math.OC cs.LG

classification math.OCcs.LG
keywords convergenceapproachcontrolalgorithmentropy-regularizedhorizonmodelpdes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper we investigate the convergence of the Policy Iteration Algorithm (PIA) for a class of general continuous-time entropy-regularized stochastic control problems. In particular, instead of employing sophisticated PDE estimates for the iterative PDEs involved in the algorithm (see, e.g., Huang-Wang-Zhou(2025)), we shall provide a simple proof from scratch for the convergence of the PIA. Our approach builds on probabilistic representation formulae for solutions of PDEs and their derivatives. Moreover, in the finite horizon model and in the infinite horizon model with large discount factor, the similar arguments lead to a super-exponential rate of convergence without tear. Finally, with some extra efforts we show that our approach can be extended to the diffusion control case in the one dimensional setting, also with a super-exponential rate of convergence.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond separability: convergence rate of vanishing viscosity approximations to mean field games via FBSDE stability

    math.OC 2025-05 conditional novelty 7.0 of 10

    The vanishing viscosity approximation to nonlocal, possibly non-separable mean field games converges at rate O(β) in L∞ on compact sets, matching the classical Hamilton-Jacobi rate.

  2. Simulating Fokker-Planck equations via mean field control of score-based normalizing flows

    math.OC 2025-06 conditional novelty 4.0 of 10

    A mean field control formulation using score-based normalizing flows simulates Fokker-Planck equations deterministically, with a convergence theorem for Ornstein-Uhlenbeck processes and experiments on Langevin and cha...

Pith tools