Pith. sign in

REVIEW 5 cited by

Wasserstein Policy Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.00663 v1 pith:P2XFTJ5K submitted 2025-05-01 cs.LG cs.AI

classification cs.LGcs.AI
keywords policygradientwassersteinactionalgorithmclassiccontinuouscontrol
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Wasserstein Policy Optimization (WPO), an actor-critic algorithm for reinforcement learning in continuous action spaces. WPO can be derived as an approximation to Wasserstein gradient flow over the space of all policies projected into a finite-dimensional parameter space (e.g., the weights of a neural network), leading to a simple and completely general closed-form update. The resulting algorithm combines many properties of deterministic and classic policy gradient methods. Like deterministic policy gradients, it exploits knowledge of the gradient of the action-value function with respect to the action. Like classic policy gradients, it can be applied to stochastic policies with arbitrary distributions over actions -- without using the reparameterization trick. We show results on the DeepMind Control Suite and a magnetic confinement fusion task which compare favorably with state-of-the-art continuous control methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Policy Gradient for Continuous-Time Robust Markov Decision Processes

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    Extends robust MDPs to continuous time with policy gradient derivations using differential equation methods and proposes optimizers achieving linear convergence and specific sample complexities.

  2. Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Wasserstein policy gradient converges globally in entropy-regularized RL via Bellman-induced distributional PL geometry and uniform LSI, yielding geometric contraction up to discretization bias.

  3. Ratio-Variance Regularized Policy Optimization

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    R²VPO uses ratio-variance regularization as a distributional soft brake on policy updates, claiming better performance than PPO on math reasoning and robotic control without hard clipping.

  4. A note on convergence of Wasserstein policy optimization

    cs.LG 2026-05 unverdicted novelty 4.0 of 10

    The note claims linear convergence of WPO in entropy-regularized MDPs by combining mean-field gradient flow analysis with a local log-Sobolev inequality under a regularity assumption.

  5. Challenges and opportunities for AI to help deliver fusion energy

    physics.plasm-ph 2026-03 unverdicted novelty 2.0 of 10

    AI offers opportunities to advance fusion energy R&D but requires responsible practices and expert collaborations to overcome its inherent challenges.

Pith tools