Pith. sign in

Paper Citation Record · LEDGER

Wasserstein Policy Optimization

As of 17 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 5 inbound Pith citation observations for arXiv:2505.00663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00663 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:47:05.795761Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:19:53.934535Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T06:56:44.503712Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact3
  • verified fuzzy22
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5b7dfde-d8ef-4053-a00f-f73925c912b2 · outbound

This paper cites write newline.

Wasserstein Policy Optimization write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.588081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.588081Z digest=sha256:31a5807d450a294e3ebe39c379c571808823316ae38967fe64ea559e6549d55e

Observation 2254d3da-e420-4c79-b6aa-9e78c980eec9 · outbound

This paper cites T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M.

Wasserstein Policy Optimization T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.331926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.593154Z digest=sha256:7abb0d6e2d8cd1978488c981727846bbc1e1074a914df1fe05e937516ee5e9b3

Observation 5ab9bc6b-cab4-4dbc-b48b-8645b4b1b909 · outbound

This paper cites Wasserstein Robust Reinforcement Learning.

Wasserstein Policy Optimization Wasserstein Robust Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.597189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.597189Z digest=sha256:40e2bd36f2a341da548c4d5491b0a5fcd877e6634e91157151927ab9a1f34ff5

Observation 733ba229-d450-4089-88ae-ba561c450cbc · outbound

This paper cites M., Lee, J.

Wasserstein Policy Optimization M., Lee, J

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.601296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.601296Z digest=sha256:e33c75c873c4099bcf4c0928328eb11f32bfa741bc9052571ec4fea277084fdb

Observation d34dd838-1579-4ca8-aba0-b935f4832c21 · outbound

This paper cites Gradient flows: in metric spaces and in the space of probability measures.

Wasserstein Policy Optimization Gradient flows: in metric spaces and in the space of probability measures

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.604877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.604877Z digest=sha256:2592f9e0b78e998c4ec66a45937c6fa441c5b817f1c7532351eab79601a10be2

Observation 56df3e18-0b0e-4184-b5d8-f649980ff82f · outbound

This paper cites W., Budden, D., Dabney, W., Horgan, D., Tb, D., Muldal, A., Heess, N., and Lillicrap, T.

Wasserstein Policy Optimization W., Budden, D., Dabney, W., Horgan, D., Tb, D., Muldal, A., Heess, N., and Lillicrap, T

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.306496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.608466Z digest=sha256:068a9d1e86610aada273bcb394135291febd79055e6e7b6db70954d80fb4ed1c

Observation 41dad690-926f-4135-b459-184cc07b59cb · outbound

This paper cites G., Sutton, R.

Wasserstein Policy Optimization G., Sutton, R

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.294957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.612357Z digest=sha256:28e485a5796e6d8a3e7989f4061e54bba7c3cab558c624bf1a78d9aec8acad6b

Observation c109af0b-e5a3-41d2-a497-fa0e246179c5 · outbound

This paper cites G., Dabney, W., and Munos, R.

Wasserstein Policy Optimization G., Dabney, W., and Munos, R

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.283686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.616220Z digest=sha256:2ac01afdd0f71ad76842db545917cd2a48e318450c701935eb56dd2c28a0004d

Observation 9662de86-3d7e-48cc-930a-e1e12a352561 · outbound

This paper cites and Brenier, Y.

Wasserstein Policy Optimization and Brenier, Y

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.619989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.619989Z digest=sha256:fbf64cfaa309176d700844a8d0dcecb20bb836484b9ba14dad5acdbfe5de40e9

Observation d8ac57cd-5b35-4274-82fb-516c7af50691 · outbound

This paper cites Development of free-boundary equilibrium and transport solvers for simulation and real-time interpretation of tokamak experiments.

Wasserstein Policy Optimization Development of free-boundary equilibrium and transport solvers for simulation and real-time interpretation of tokamak experiments

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.266647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.624192Z digest=sha256:78f6fb11f73b2f3717bf83f8d55baa43c4402d60b15edcfa3203093f705bc8b0

Observation 59307d2d-0cea-46e5-806f-be758d7ddaf0 · outbound

This paper cites MICo: Improved representations via sampling-based state similarity for Markov decision processes.

Wasserstein Policy Optimization MICo: Improved representations via sampling-based state similarity for Markov decision processes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.628603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.628603Z digest=sha256:1a8a471bcc8d01a2eeb4c756429b47d650a89834f0a8d56c8856a4635fd2e41a

Observation 5b4cddc1-b794-4c27-9d69-bffcd7d184c8 · outbound

This paper cites T., Rubanova, Y., Bettencourt, J., and Duvenaud, D.

Wasserstein Policy Optimization T., Rubanova, Y., Bettencourt, J., and Duvenaud, D

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.634035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.634035Z digest=sha256:c44f678c499b130932557cb65bbd971bcfe28cec98df05e7c092c90371ad787f

Observation c38ecf22-afac-4732-8aba-d3fcf077b551 · outbound

This paper cites Fast and accurate deep network learning by exponential linear units (elus).

Wasserstein Policy Optimization Fast and accurate deep network learning by exponential linear units (elus)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.248832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.638728Z digest=sha256:6c5061ede958ebb7e9433d1f89be4b73392201fd3c2e7e1d252e2b52ed5fbb82

Observation 6f464a34-6fdf-406c-87b2-0e9a0a2d4222 · outbound

This paper cites Magnetic control of tokamak plasmas through deep reinforcement learning.

Wasserstein Policy Optimization Magnetic control of tokamak plasmas through deep reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.642253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.642253Z digest=sha256:075cbb264fd4aa4cbd65577309863240a5e4528bb106ba51a16188d2c6b55544

Observation 0495d737-0898-4b96-bce4-c7eec5c1cbae · outbound

This paper cites Experimental research on the TCV tokamak.

Wasserstein Policy Optimization Experimental research on the TCV tokamak

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.645704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.645704Z digest=sha256:86a9edf1fbbb221bf55f49e92131592766907ddc6da2dae679e7d05a5fdddef9

Observation 3a3fc4cf-9920-41a2-b044-ee578f2224fb · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning.

Wasserstein Policy Optimization Sigmoid-weighted linear units for neural network function approximation in reinforcement learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.649683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.649683Z digest=sha256:0a8a8a51224101013f60914d98d020b445806b2dac0ac3ba630f658ea366c84c

Observation 2ec59336-bd68-4aaf-a304-ee1b46abba65 · outbound

This paper cites CALE: Continuous Arcade Learning Environment.

Wasserstein Policy Optimization CALE: Continuous Arcade Learning Environment

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:47:05.935740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.653202Z digest=sha256:64e8abfce955445df828625d338c802ee0bf54ccb9ad316a09571d70f36044ef

Observation b7d4ecf8-0e0e-4cbb-b112-33cb32583843 · outbound

This paper cites Metrics for finite markov decision processes.

Wasserstein Policy Optimization Metrics for finite markov decision processes

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.222698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.657534Z digest=sha256:6ca76eb0ffda524fb3032b1cdfcf2a45281540fe5bcaca8bd382ac9fbfe3ef8e

Observation 6335febe-a02a-473e-9ae8-dedb8a0a42fe · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Wasserstein Policy Optimization Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.661528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.661528Z digest=sha256:cd61898c496cf1963df534e2746e52023acf474a4d576e3a2300f9107143aae0

Observation 27536610-f19e-42b6-8a56-2f38f5d39127 · outbound

This paper cites H., Tirumala, D., Humplik, J., Wulfmeier, M., Tunyasuvunakool, S., Siegel, N.

Wasserstein Policy Optimization H., Tirumala, D., Humplik, J., Wulfmeier, M., Tunyasuvunakool, S., Siegel, N

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.202530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.666603Z digest=sha256:03ad3431425fe67b1026620ac6a7beee87177f3e266db57c6d69fb881183bab1

Observation 857f1670-9d4e-432b-99e7-5be13481fc37 · outbound

This paper cites Wasserstein Unsupervised Reinforcement Learning.

Wasserstein Policy Optimization Wasserstein Unsupervised Reinforcement Learning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:47:05.922391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.670497Z digest=sha256:b1a42f8938a6723e68756a736fd77430bbc7a5ee36e6ccfd3ecc6f2eeeefd1c3

Observation 744ec0f7-4171-4892-9043-dd086adef2b2 · outbound

This paper cites Learning continuous control policies by stochastic value gradients.

Wasserstein Policy Optimization Learning continuous control policies by stochastic value gradients

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.191434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.674549Z digest=sha256:cf977e2f1a19d8bc848e188e02962f66931c422bd9125b2fb36312a3abb6e3ca

Observation 2fd903b9-9f74-4b6a-b4c6-bf191014cea0 · outbound

This paper cites Acme: A Research Framework for Distributed Reinforcement Learning.

Wasserstein Policy Optimization Acme: A Research Framework for Distributed Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.678159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.678159Z digest=sha256:0aa69bb002d832b4569654f08c0e9fb81b723ef4f4c68ef5a6435c92b74cc93c

Observation 0395637c-0a5c-4c90-9a66-d09ac040974d · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.177293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.683628Z digest=sha256:6bb320867dcb877a7b2f6d282e5c8bbd245543ce507c091e587abebfcfcc4228

Observation 45aaa500-a603-4a85-a431-74aea021f4e5 · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.165323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.689320Z digest=sha256:61fc0548b10e348dee41ca93811eed0d8c27e7269bb554909e3cd8fb38d72d16

Observation b48467e3-ffd7-45c6-a333-91bda38df65b · outbound

This paper cites Categorical reparameterization with G umbel-softmax.

Wasserstein Policy Optimization Categorical reparameterization with G umbel-softmax

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.154538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.694063Z digest=sha256:0ee1c051d08da9af60ec4f3487418ab4626eb3e5df8c364b5b21173714902a81

Observation acc2e2a9-bed6-428e-bb47-9446929ffd68 · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.142634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.698562Z digest=sha256:686f62615685d914bbf2bd768b4917084cbf3154e326d090c88d5885b0a9e7de

Observation 51a4ad7c-dfc5-4b11-af34-2432dd3181ce · outbound

This paper cites Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions Control.

Wasserstein Policy Optimization Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions Control

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:47:05.898230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.702822Z digest=sha256:5b0aabb2822abf528c8b65b51feb239b90e5577c40f78994d76b4e96ecee073f

Observation 683d7e2a-d004-4fca-a499-b7fcbeebd1f1 · outbound

This paper cites Continuous control with deep reinforcement learning.

Wasserstein Policy Optimization Continuous control with deep reinforcement learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.706926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.706926Z digest=sha256:36c60ae5ec13ccf5d37e10c96c985be3c42a84b00b8b5767f94013488a4f8eac

Observation fa09afc5-6421-442b-89e4-cf543840fdda · outbound

This paper cites J., Mnih, A., and Teh, Y.

Wasserstein Policy Optimization J., Mnih, A., and Teh, Y

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.130226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.711300Z digest=sha256:66775c1b4e9c9627a28e33aceca60f5dbcb8797505e03afdb4bf268cf8b40fb3

Observation 9850b65c-b8ab-4482-8108-b8e4a356824a · outbound

This paper cites and Grosse, R.

Wasserstein Policy Optimization and Grosse, R

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.118288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.715064Z digest=sha256:4412b1e08862c7fa43dc6bbabc40303536f720195cf096e9ca4f7df56ef506c3

Observation ac03dda1-012e-4034-95eb-2c6f0a8b5bb0 · outbound

This paper cites M., Likmeta, A., and Restelli, M.

Wasserstein Policy Optimization M., Likmeta, A., and Restelli, M

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.105754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.718976Z digest=sha256:23b208f85f6617ceeb37964e0e7399072d1b790394af5ca170abe42298824bac

Observation e144be77-e467-494d-945b-4a65c3f7f94a · outbound

This paper cites Efficient Wasserstein Natural Gradients for Reinforcement Learning.

Wasserstein Policy Optimization Efficient Wasserstein Natural Gradients for Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.722747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.722747Z digest=sha256:2a57a2b68700fb4a9c535483eee68b6c300fbc3d1307d43e547bfa4ea67ba491

Observation afd51786-a8c3-48b1-b167-0d308880fc81 · outbound

This paper cites Wasserstein quantum M onte C arlo: a novel approach for solving the quantum many-body schr \"o dinger equation.

Wasserstein Policy Optimization Wasserstein quantum M onte C arlo: a novel approach for solving the quantum many-body schr \"o dinger equation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.094875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.727282Z digest=sha256:ca0a077451f0f11930db9ff85d333f895745bdc48be0c115c87884ac65e5f286

Observation f9fe4337-2e4f-4197-ae50-0dc6f03d9957 · outbound

This paper cites Learning to score behaviors for guided policy optimization.

Wasserstein Policy Optimization Learning to score behaviors for guided policy optimization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.083773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.731386Z digest=sha256:02fe984e46a3d4a2e56095aa46e9d1976276dcfe69436e719ad1bf727296740e

Observation 75eb8626-fb33-486c-b607-92acf35b5274 · outbound

This paper cites Revisiting Natural Gradient for Deep Networks.

Wasserstein Policy Optimization Revisiting Natural Gradient for Deep Networks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.734888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.734888Z digest=sha256:7c26519502d6bfd3ff21d8f42ba95f129a3bae0870755d4ecb60fe7c59fe5b03

Observation d6f13103-9698-4474-9d76-bc0048a81deb · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.073639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.739123Z digest=sha256:5ff5030d6f21a44bfefc4a3495b39dc0b2e819136a1db26f39fb32b015b2f6b8

Observation 800c788b-db9b-4b7d-a572-59a67930c74c · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:06.063889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.742610Z digest=sha256:da7fc2d79cb27f2ece0e655d88d70bcfea4d24874fdf5c239e32ee3b38ca1981

Observation 094c430b-f82f-4804-b88c-8209fbf8b329 · outbound

This paper cites Trust region policy optimization.

Wasserstein Policy Optimization Trust region policy optimization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.054524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.746388Z digest=sha256:8d0d9446c0458e5befe306163b9e6aeac2f8e3aa180f68e6d6789161c18d407e

Observation e8ff4f7c-2665-4a90-8f0b-8d98c02f2f19 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Wasserstein Policy Optimization High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.750096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.750096Z digest=sha256:442060622ac6b532a2725e60babf0557924c35e64d11fea9d3680da276e50dee

Observation 368ea6c0-f820-49db-9008-38cc22728977 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Wasserstein Policy Optimization Proximal Policy Optimization Algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.753537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.753537Z digest=sha256:95122a310b80b7080e0c7e803e71bc55b2bb71375908308a3c475dc271209df9

Observation 07849974-2fdd-4b13-8a43-d4366ea6a636 · outbound

This paper cites Deterministic policy gradient algorithms.

Wasserstein Policy Optimization Deterministic policy gradient algorithms

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.044687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.757112Z digest=sha256:0ec6852a3ccf8c9801f931bbee46f247fe0343d997263c66de1da1f4d59b4735

Observation 49344668-057b-44f7-aa5c-cc03726391f9 · outbound

This paper cites V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control.

Wasserstein Policy Optimization V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.760618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.760618Z digest=sha256:b24229f77b9121312bf94793ddbf30e90e4e306963eb91426daa4d3aa524e4f7

Observation 6dee8b22-00dd-4fd2-953b-4b7fa8930f26 · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.763904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.763904Z digest=sha256:7fda28edb4a3ac4217c01aa50cb4dfe85f55c4f6f6eba12d47b648d09cfb1f93

Observation e000d2e0-7711-4541-8ec6-49ec0d647327 · outbound

This paper cites S., McAllester, D., Singh, S., and Mansour, Y.

Wasserstein Policy Optimization S., McAllester, D., Singh, S., and Mansour, Y

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.767085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.767085Z digest=sha256:726b90db86d51ec652e9be6d867254e60d0ffd95afc8d19117266f64be42f810

Observation e5e35a59-79ac-48c1-99df-a0c872ebd9c8 · outbound

This paper cites DeepMind Control Suite.

Wasserstein Policy Optimization DeepMind Control Suite

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.770349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.770349Z digest=sha256:5c7499e3d439634c3ea3be9be9a196a326500f7c35174b9b9e07d594e7b09afb

Observation f4975490-ef45-488d-9b33-09be25f5bf14 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Wasserstein Policy Optimization Mujoco: A physics engine for model-based control

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.773879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.773879Z digest=sha256:c61925474e5058bbb8a39792ba2d803e6a81eab25ed4d2ea7af83f067144527e

Observation 1b4fa37c-7da9-453d-82ff-66061c19dcde · outbound

This paper cites D., Michi, A., Chervonyi, Y., Davies, I., Paduraru, C., Lazic, N., Felici, F., Ewalds, T., Donner, C., Galperti, C., et al.

Wasserstein Policy Optimization D., Michi, A., Chervonyi, Y., Davies, I., Paduraru, C., Lazic, N., Felici, F., Ewalds, T., Donner, C., Galperti, C., et al

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.776825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.776825Z digest=sha256:e88d88de5f7b2beff6b3b10ef8cf84535722b898df5085fa248451af46f2cebb

Observation 72b0774e-5686-4f6c-81fb-ee1b71650855 · outbound

This paper cites dm\_control: Software and tasks for continuous control.

Wasserstein Policy Optimization dm\_control: Software and tasks for continuous control

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.010833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.779820Z digest=sha256:6142494afc932fb6cdd158ecba27cc620d5c3c3c9645ccda1ae73b6c11b1402c

Observation 76ec6004-2123-427e-af7e-1944d25931b7 · outbound

This paper cites Double q-learning.

Wasserstein Policy Optimization Double q-learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:06.000229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.783069Z digest=sha256:b498bbf358a464e4b3844e5ba000cb149b1032a66e69e748dbe51bc0e7f7d994

Observation 3aa6ab08-bd3a-4050-a7d9-e48c39a41b19 · outbound

This paper cites and Wiering, M.

Wasserstein Policy Optimization and Wiering, M

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:05.989811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.786087Z digest=sha256:2c7951bb244b145432b5d5ee9107d38c79d7de9a8f533d16f7db1eb7ba20fdf4

Observation 2014b134-8925-4098-b8d1-a254e7e549b1 · outbound

This paper cites Beyond regression: N ew tools for prediction and analysis in the behavioral sciences.

Wasserstein Policy Optimization Beyond regression: N ew tools for prediction and analysis in the behavioral sciences

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:05.978635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.789226Z digest=sha256:70d4bc14e00f8fd3076a37398056086bed6c3e9a5d860cf9debd8756a62d731e

Observation 7ec0bc46-1d29-498e-8564-414468894dd6 · outbound

This paper cites an unresolved cited work.

Wasserstein Policy Optimization Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:05.792655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:47:05.792655Z digest=sha256:f37e9aa1ecd180dc11b3327b98059496e2c5f04992e4a11a9d35efa191be1287

Observation 8b443736-3676-4c3d-a32e-2b9135f53f77 · outbound

This paper cites Policy optimization as W asserstein gradient flows.

Wasserstein Policy Optimization Policy optimization as W asserstein gradient flows

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:05.962950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:47:05.795761Z digest=sha256:3506cfbb897dee3bbb134cb3ae986d432bc7a239e81d228466022cbf2b882625

Pith citing papers

Observation ac0cc622-9671-4b93-b2be-8695092a6a38 · inbound

Challenges and opportunities for AI to help deliver fusion energy cites this paper.

Challenges and opportunities for AI to help deliver fusion energy Wasserstein Policy Optimization

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:43:25.441685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T00:39:46.244035Z digest=sha256:246785e34f8088f24dcfd2abda5936656fdabfa1ff1dadf72666f1890972f720

Observation d42fbe4a-6cef-4066-bedb-7edae218c41c · inbound

A note on convergence of Wasserstein policy optimization cites this paper.

A note on convergence of Wasserstein policy optimization Wasserstein Policy Optimization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:41:10.858631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-22T06:36:57.345741Z digest=sha256:982ed71b4023ed6babe885193a415c15a3ab3e28a35718bada562ab6b180dfcb

Observation 5238f7d7-76f5-44de-b82b-131b8f3d9832 · inbound

Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning cites this paper.

Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning Wasserstein Policy Optimization

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:24:00.339861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T22:19:53.934535Z digest=sha256:28124b9575134a04cb02e29619c0cc4c04d7bb3d787590bcd5989a5bed5f59ec

Observation 42fb7b7b-e288-4bc2-a5d4-6af7262093e9 · inbound

Ratio-Variance Regularized Policy Optimization cites this paper.

Ratio-Variance Regularized Policy Optimization Wasserstein Policy Optimization

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:53:55.983387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:ed8581f09c3e873572f7c9c0e999d7d6eb02cae25febf911b60b65a88607ea32

Observation 2aca0c6c-ab04-45a0-9c52-c2ef9587a81b · inbound

Policy Gradient for Continuous-Time Robust Markov Decision Processes cites this paper.

Policy Gradient for Continuous-Time Robust Markov Decision Processes Wasserstein Policy Optimization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:56:44.505465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T07:18:48.673738Z digest=sha256:0e085909e209d6ba4a7c276422e66abb1cf9fcd0e17deb00f0a79d871c86a466