Pith. sign in

REVIEW

Augment-Reinforce-Merge Policy Gradient for Binary Stochastic Policy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1903.05284 v1 pith:GJCHRL7O submitted 2019-03-13 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords policygradientestimatoraugment-reinforce-mergebinaryvarianceachievesaction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Due to the high variance of policy gradients, on-policy optimization algorithms are plagued with low sample efficiency. In this work, we propose Augment-Reinforce-Merge (ARM) policy gradient estimator as an unbiased low-variance alternative to previous baseline estimators on tasks with binary action space, inspired by the recent ARM gradient estimator for discrete random variable models. We show that the ARM policy gradient estimator achieves variance reduction with theoretical guarantees, and leads to significantly more stable and faster convergence of policies parameterized by neural networks.

Discussion (0). Continue with ORCID to comment.

Pith tools