Pith. sign in

REVIEW 2 cited by

Finite-Time Analysis of Entropy-Regularized Neural Natural Actor-Critic Algorithm

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.00833 v1 pith:XUZOIXR3 submitted 2022-06-02 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords neuralcomplexitynetworkoptimizationregularizationachieveactoractor-critic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Natural actor-critic (NAC) and its variants, equipped with the representation power of neural networks, have demonstrated impressive empirical success in solving Markov decision problems with large state spaces. In this paper, we present a finite-time analysis of NAC with neural network approximation, and identify the roles of neural networks, regularization and optimization techniques (e.g., gradient clipping and averaging) to achieve provably good performance in terms of sample complexity, iteration complexity and overparametrization bounds for the actor and the critic. In particular, we prove that (i) entropy regularization and averaging ensure stability by providing sufficient exploration to avoid near-deterministic and strictly suboptimal policies and (ii) regularization leads to sharp sample complexity and network width bounds in the regularized MDPs, yielding a favorable bias-variance tradeoff in policy optimization. In the process, we identify the importance of uniform approximation power of the actor neural network to achieve global optimality in policy optimization due to distributional shift.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Wasserstein Adaptive Value Estimation for Actor-Critic Reinforcement Learning

    cs.LG 2025-01 reject novelty 4.0 of 10

    WAVE adds an adaptively weighted Sinkhorn approximation of the Wasserstein distance between successive Q-value distributions to the critic loss in actor-critic reinforcement learning.

  2. Partially Observed Optimal Stochastic Control: Regularity, Optimality, Approximations, and Learning

    math.OC 2024-12 conditional novelty 2.0 of 10

    A survey of regularity, approximation, and reinforcement learning guarantees for partially observed Markov decision processes, drawing mostly on the authors' earlier work.

Pith tools