Pith. sign in

REVIEW 2 cited by

Sample Efficient Reinforcement Learning with REINFORCE

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.11364 v2 pith:6NTN4ALP submitted 2020-10-22 cs.LG math.OC

classification cs.LGmath.OC
keywords gradientconvergenceglobalmethodsreinforceclassicalgradientslearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Policy gradient methods are among the most effective methods for large-scale reinforcement learning, and their empirical success has prompted several works that develop the foundation of their global convergence theory. However, prior works have either required exact gradients or state-action visitation measure based mini-batch stochastic gradients with a diverging batch size, which limit their applicability in practical scenarios. In this paper, we consider classical policy gradient methods that compute an approximate gradient with a single trajectory or a fixed size mini-batch of trajectories under soft-max parametrization and log-barrier regularization, along with the widely-used REINFORCE gradient estimation procedure. By controlling the number of "bad" episodes and resorting to the classical doubling trick, we establish an anytime sub-linear high probability regret bound as well as almost sure global convergence of the average regret with an asymptotically sub-linear rate. These provide the first set of global convergence and sample efficiency results for the well-known REINFORCE algorithm and contribute to a better understanding of its performance in practice.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AM-PPO: (Advantage) Alpha-Modulation with Proximal Policy Optimization

    cs.LG 2025-05 reject novelty 5.0 of 10

    AM-PPO modulates GAE advantages with a feedback-controlled tanh gate and reports improved reward trajectories on MuJoCo benchmarks, tested once per configuration.

  2. Model-free Reinforcement Learning for Model-based Control: Towards Safe, Interpretable and Sample-efficient Agents

    cs.LG 2025-07 conditional novelty 3.0 of 10

    A perspective paper argues that model predictive control can be used as a learned policy in model-free reinforcement learning and reviews the methods and open problems.

Pith tools