REVIEW 1 cited by
Efficient Wasserstein Natural Gradients for Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
A novel optimization approach is proposed for application to policy gradient methods and evolution strategies for reinforcement learning (RL). The procedure uses a computationally efficient Wasserstein natural gradient (WNG) descent that takes advantage of the geometry induced by a Wasserstein penalty to speed optimization. This method follows the recent theme in RL of including a divergence penalty in the objective to establish a trust region. Experiments on challenging tasks demonstrate improvements in both computational cost and performance over advanced baselines.
Forward citations
Cited by 1 Pith paper
-
Wasserstein Policy Optimization
WPO derives a closed-form policy update from Wasserstein gradient flows, which for Gaussian policies coincides with the standard policy gradient in expectation but with lower variance, and works for arbitrary stochast...
Discussion (0). Continue with ORCID to comment.