Pith. sign in

Paper Citation Record · LEDGER

On the Policy Convergence of Policy Mirror Descent Methods

As of 19 August 2026, this Paper Citation Record lists 100 of 142 outbound references and 0 inbound Pith citation observations for arXiv:2607.11626.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.11626 v1

Coverage vector

measured 100 of 142 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T04:17:12.534820Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 142 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b13e07c-8818-4835-acdc-1a5bc91c8e77 · outbound

This paper cites Nature , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Nature , volume=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:e370699df7497d277445d6a071f57180d64befebdecbd87b9689929e7d8973d3

Observation bd28bf77-378b-44cd-b4e5-650ded9bd911 · outbound

This paper cites Mastering the game of.

On the Policy Convergence of Policy Mirror Descent Methods Mastering the game of

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:7103129ca87dcec5fc08a4df06a35e3d6c04e19ebbb41d4ad166674d43a84947

Observation 9870bdbc-c5e3-4fa1-b43d-12c7135169d1 · outbound

This paper cites 2019 , journal=.

On the Policy Convergence of Policy Mirror Descent Methods 2019 , journal=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:5cef0d2ddce6fec3c8dad66842d8a3c1e8d62237e19a3450bcda55f135c8c046

Observation 6193d972-ea4e-4cff-b657-5a0c79578cce · outbound

This paper cites Elementary Analysis of Policy Gradient Methods.

On the Policy Convergence of Policy Mirror Descent Methods Elementary Analysis of Policy Gradient Methods

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:cf17fe425ed4bc29a0aeeb4cf9630944ed51b600864f7b394e27d0f403d02132

Observation d0dfd9fc-8476-45be-94a5-c784bf1c7a54 · outbound

This paper cites Starcraft.

On the Policy Convergence of Policy Mirror Descent Methods Starcraft

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:a0a89a229e600ceb3270ec8c22b3dcf6c8119453e342168ba4caa5738db6381f

Observation 4e9f8405-2753-46a4-9e5c-bd1895dbda7f · outbound

This paper cites Science Robotics , year=.

On the Policy Convergence of Policy Mirror Descent Methods Science Robotics , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:f065944348b9d8eb016c3ec7d40307a085436024f673fcb2679ef8fd53d6068b

Observation 09d04b13-9c04-40e1-a774-5085018c676f · outbound

This paper cites Science Robotics , volume =.

On the Policy Convergence of Policy Mirror Descent Methods Science Robotics , volume =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:7bee7859645736b85e0caf30ce16d208fd42983481c61075cdb63d587bf58b45

Observation d63c5b9f-acdc-4cfc-ae35-343b22f9f3dc · outbound

This paper cites Learning robust perceptive locomotion for quadrupedal robots in the wild , volume =.

On the Policy Convergence of Policy Mirror Descent Methods Learning robust perceptive locomotion for quadrupedal robots in the wild , volume =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:6d6ba4df04418476f2ad971d54387767520e19cda780987bfefdb78e2ebf5813

Observation 90cbff71-a513-4551-9c28-a964e005e6ca · outbound

This paper cites Making Contextual Decisions with Low Technical Debt.

On the Policy Convergence of Policy Mirror Descent Methods Making Contextual Decisions with Low Technical Debt

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:6dadc48115aace04b4b9da160d99ce84648e29032a76fe42b57c4e7ade6d1647

Observation 45d7be47-d62a-42af-895a-f02ae0c7fd91 · outbound

This paper cites , booktitle=.

On the Policy Convergence of Policy Mirror Descent Methods , booktitle=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:582be8d100313c5d86801f687bb72fb0b1866071935c070f2918b843641ed9e6

Observation afcc29dd-c109-444d-9af3-6fcb2ab98e78 · outbound

This paper cites Nature , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Nature , volume=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:589c84b301c35a12165475ec2a4cfbe9bfe7a082c2360d9fb0f262b6d596ac54

Observation 6963f7b2-8e41-4757-887c-98f08f9bdcdf · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

On the Policy Convergence of Policy Mirror Descent Methods Advances in Neural Information Processing Systems , year=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:660846eee79b78d2d7f69c0b10c05cf3dab5466501b8baeb3dae83e0f810cd01

Observation 4ed4838d-fba6-4376-be91-2fe1a0b00981 · outbound

This paper cites 2019 , pages=.

On the Policy Convergence of Policy Mirror Descent Methods 2019 , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:5818b3501048295f9720aae0ab1bdf91a8b28fbb88ff0a15ba306fa4838aa5aa

Observation c3fa16a0-8415-48f2-ac13-e8d422c37d8c · outbound

This paper cites Adaptive Trust Region Policy Optimization:.

On the Policy Convergence of Policy Mirror Descent Methods Adaptive Trust Region Policy Optimization:

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:23c11da55313a9f16bb9017378a4c0689829e6aea27f45db2c490cc98ed4d0b0

Observation 945439db-d11d-4d1c-8466-d9dcd2d8e959 · outbound

This paper cites On the Convergence Rates of Policy Gradient Methods , journal=.

On the Policy Convergence of Policy Mirror Descent Methods On the Convergence Rates of Policy Gradient Methods , journal=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:631af4cb84215542958eb8eb8360a2bd0eefaa1f48136186fe3634311eae17db

Observation 1d87765c-ac94-4cd7-b66e-b2c2e52b1107 · outbound

This paper cites Mathematical Programming , author=.

On the Policy Convergence of Policy Mirror Descent Methods Mathematical Programming , author=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:d71c1aec67f5b060788f8cade915cbc2d403ef1c8b9e0bc4388318f1d49c3247

Observation b99c6e00-d5f2-4f9c-8b1f-ea6c49544848 · outbound

This paper cites Escaping the Gravitational Pull of Softmax , booktitle=.

On the Policy Convergence of Policy Mirror Descent Methods Escaping the Gravitational Pull of Softmax , booktitle=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:94d8e971dd2ab198e206ada188c0487bfa06281e1e99d36f11d795f59458fe6e

Observation 135be19a-a105-4e10-9c84-35d3ea8d5be9 · outbound

This paper cites 2023 , booktitle=.

On the Policy Convergence of Policy Mirror Descent Methods 2023 , booktitle=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:f7fb9e72673cafe423687825c3af53d0b5922eceabdf64ff0e3233cba39b104b

Observation 18640ec0-9e5b-4089-85ea-26b0eea23c2b · outbound

This paper cites Fast Global Convergence of Natural Policy Gradient Methods with Entropy Regularization , journal=.

On the Policy Convergence of Policy Mirror Descent Methods Fast Global Convergence of Natural Policy Gradient Methods with Entropy Regularization , journal=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:d09af0976b6dd63fc2e43bf90861d58a77bacfd918f34a1026cfe32b18d6dacf

Observation 03b727e5-c8db-40cf-8866-3b17ae090699 · outbound

This paper cites Mathematical Programming , author=.

On the Policy Convergence of Policy Mirror Descent Methods Mathematical Programming , author=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:e73740c49322cd1999124b11342aafbdee52592888a6eba4440cbde3b790d1ab

Observation e102a90f-255c-47a6-84d6-e22f705e15a3 · outbound

This paper cites SIAM Journal on Optimization , author=.

On the Policy Convergence of Policy Mirror Descent Methods SIAM Journal on Optimization , author=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:32740d0b362c80e46fa56df9d888879d90dec6b11924cdd677d710038fa84e29

Observation a6335e65-410a-4856-808a-d3beda8c08a5 · outbound

This paper cites Journal of Machine Learning Research , author=.

On the Policy Convergence of Policy Mirror Descent Methods Journal of Machine Learning Research , author=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:438a67a6cfcee664def2278786a7c4e2559fba08fe8ce918ef00ca82ad3156dd

Observation 8347442e-f35f-4fe1-8958-87d9a20f1fe0 · outbound

This paper cites On the Global Convergence Rates of Softmax Policy Gradient Methods , booktitle=.

On the Policy Convergence of Policy Mirror Descent Methods On the Global Convergence Rates of Softmax Policy Gradient Methods , booktitle=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:a753e041329a23c4dc802035446eec8814fc963543c4641eedb80a3adbeeb46d

Observation 164c734b-210d-4e4c-bf80-074326ac9032 · outbound

This paper cites 2023 , booktitle=.

On the Policy Convergence of Policy Mirror Descent Methods 2023 , booktitle=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:8cd14a0fcaa508b560c3e2c8bb7d2686dd292ca3683d97e5ec7588cc82958b91

Observation 4ca8a9bc-7716-4b8b-8cef-63f82e6ff550 · outbound

This paper cites 2022 , journal=.

On the Policy Convergence of Policy Mirror Descent Methods 2022 , journal=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:bd10137a6460551f412ba7b2ff54afad1896827ce79252ebf5f539a891e1e53a

Observation a8f6de46-d4dd-4ea9-a1fa-a61d129806bc · outbound

This paper cites On the Linear Convergence of Natural Policy Gradient Algorithm , booktitle=.

On the Policy Convergence of Policy Mirror Descent Methods On the Linear Convergence of Natural Policy Gradient Algorithm , booktitle=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:bcf148b8df6d517dbd4350976e604f51f2ce8f008c184b6efa6109cd62d6cae1

Observation 0cf8e632-6f80-4343-9bb7-ddb9b27a3fc5 · outbound

This paper cites 2020 , pages =.

On the Policy Convergence of Policy Mirror Descent Methods 2020 , pages =

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:7f8b2217354243595f9623ce24bb353d5b0563ac796136fe22cc68e4ca1e7f41

Observation 33799fd1-79e0-49a1-912e-538909989665 · outbound

This paper cites International Conference on Artificial Intelligence and Statistics , author=.

On the Policy Convergence of Policy Mirror Descent Methods International Conference on Artificial Intelligence and Statistics , author=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:2a02b72fe016f152abf638d292a8556939c73e0286c068b70968a0accf6e46b9

Observation 18e5f31f-092b-4cc6-bebc-9aad12d58520 · outbound

This paper cites , year =.

On the Policy Convergence of Policy Mirror Descent Methods , year =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:6e3c626ed6c94ff6453106513137477bad2f3a02f948cb40782c045944d2a6ac

Observation 034bdb0b-dc20-424a-9b88-ff9bee76fa8c · outbound

This paper cites International Conference on Machine Learning , pages=.

On the Policy Convergence of Policy Mirror Descent Methods International Conference on Machine Learning , pages=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:833e5445fe7855f59a54695ca29ebbc755874e1fc584bc56d0e40c778fcc2b6a

Observation 62401023-607c-4265-8e37-ad4065ec6f3f · outbound

This paper cites On the Linear Convergence of Policy Gradient under.

On the Policy Convergence of Policy Mirror Descent Methods On the Linear Convergence of Policy Gradient under

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:76c449664958875294d72fb1fddf64950247fe457159a84bbe8efc4aca248113

Observation 4c54d006-81b5-4b42-84c9-871e9ddecb7b · outbound

This paper cites , title =.

On the Policy Convergence of Policy Mirror Descent Methods , title =

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:72d82c6d07abee29880ed02646c2fb82f9a7f05369f6e6162661d1632a13a5a8

Observation f31d2e45-5f33-4d19-a2b2-47cce4bc1abe · outbound

This paper cites Improved and Generalized Upper Bounds on the Complexity of Policy Iteration , volume =.

On the Policy Convergence of Policy Mirror Descent Methods Improved and Generalized Upper Bounds on the Complexity of Policy Iteration , volume =

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:a80d7860365a9f777f9db9903318dae87f7910401904e7899fc88300aab076e1

Observation 472c6665-bf13-4ba5-8bdb-52f2566c3bc4 · outbound

This paper cites Journal of Machine Learning Research , year=.

On the Policy Convergence of Policy Mirror Descent Methods Journal of Machine Learning Research , year=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:7ea1486f6f53207d201f0dff6cdcc45c6d8191e3f23c4d916f301725bbbd9e7f

Observation 70eb4c51-c10e-4e17-8f60-d8c49edc4223 · outbound

This paper cites R\'enyi Divergence and Kullback-Leibler Divergence.

On the Policy Convergence of Policy Mirror Descent Methods R\'enyi Divergence and Kullback-Leibler Divergence

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:b196009742d3da3cbb97d4b21a975bb577ada9ddb0dadc6b89268e236cc2bd97

Observation 5a0fa521-6daa-4166-aaef-dc8bf32902df · outbound

This paper cites International Conference on Machine Learning , year =.

On the Policy Convergence of Policy Mirror Descent Methods International Conference on Machine Learning , year =

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:084d9be526603338f6b7823ea8c164275a05632ddbf022d4dd3e66952e925f5d

Observation b46525f3-b451-4c6e-a227-3e11396e9428 · outbound

This paper cites Machine Learning , year =.

On the Policy Convergence of Policy Mirror Descent Methods Machine Learning , year =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:773a3f69aed3178311239b6aac4d46a3bf71155fee6744a6d866cf111e588aa0

Observation eb97c3f5-cd8b-43e1-8aee-3f59eec44e9d · outbound

This paper cites Sutton and Andrew G.

On the Policy Convergence of Policy Mirror Descent Methods Sutton and Andrew G

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:526ac186151ae6a00b897b1b53d996e7a76461643260bafbac54c13b0b5013a9

Observation 28cb6e32-828b-47d6-8163-24b0e92b2fb8 · outbound

This paper cites A natural policy gradient , year =.

On the Policy Convergence of Policy Mirror Descent Methods A natural policy gradient , year =

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:3a9262693c6ab2ad42e11942778cde5717279e273c09b7ebaea904e3ac7d9e93

Observation 7f894097-4641-435b-9a7c-8b0d412c70f9 · outbound

This paper cites International conference on machine learning , pages=.

On the Policy Convergence of Policy Mirror Descent Methods International conference on machine learning , pages=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:4f09e7a17494af93cd3f73def070dd7780a1b5f886d2543bdb7dc5e799a52676

Observation 4f44e2db-cee0-438b-bc35-88821db6cfb4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

On the Policy Convergence of Policy Mirror Descent Methods Proximal Policy Optimization Algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:1bc9f8fde71e03631dce6b225e1c56868c4650e30b4c5318590109866354c64e

Observation 46e7b2e8-efeb-4bd2-af3e-3e9f7736b45b · outbound

This paper cites Operations Research , year =.

On the Policy Convergence of Policy Mirror Descent Methods Operations Research , year =

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:7cbd4efdd1333a7caecc3aaaf77a6ba1074844eb069a4e390ec25a183f9b5783

Observation 2bdd53b0-c576-4375-ba53-208f77cb3836 · outbound

This paper cites AAAI Conference on Artifical Intelligence , year =.

On the Policy Convergence of Policy Mirror Descent Methods AAAI Conference on Artifical Intelligence , year =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:6995a50b624889f041cb7b6d6d6b243bf511ac8375f1cfe4005ad7781d1a9083

Observation 0b6af84a-3327-4e74-b3ef-a3c0e14c75a8 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

On the Policy Convergence of Policy Mirror Descent Methods Advances in Neural Information Processing Systems , year =

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:c0a37a1883c6535dabd2d10d7c2eb85eeac52fdcfa6c6485f4b5c615b3edd24c

Observation ed707d0c-562a-4181-9a4c-77f159d7cbd9 · outbound

This paper cites Mathematical Programming , year =.

On the Policy Convergence of Policy Mirror Descent Methods Mathematical Programming , year =

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:d01632f33abb53c28f7b293dcec519f4fd1e9873017358f0bc699c528399d9ab

Observation 19c708c3-cc84-416c-8a77-e492e231c0ad · outbound

This paper cites Puterman and Shelby L.

On the Policy Convergence of Policy Mirror Descent Methods Puterman and Shelby L

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:d78813425fc5e1bd453082c62bc63c3e2028dc4b141f76db6276c069a5813d1b

Observation aaae1373-0d08-451b-b0a2-ba1f26909209 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

On the Policy Convergence of Policy Mirror Descent Methods Advances in Neural Information Processing Systems , year =

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:74b665ff0a01fb696eb8caf9339310ec924d274cf554def27c509d32a74d421f

Observation e2203802-5d8c-4ad9-a83e-efa35c348cb7 · outbound

This paper cites Nature , author=.

On the Policy Convergence of Policy Mirror Descent Methods Nature , author=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:0988973c103d63bea35c0e9150cf849631dfb1d29432e2286b408f1ecf6de080

Observation 5d306a4b-267b-40f1-89be-8d1589d0eb5b · outbound

This paper cites Nature , year=.

On the Policy Convergence of Policy Mirror Descent Methods Nature , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:9923452f706cb31be82f79494ef92762447f9b6b9c72cf4a41a38026512322d3

Observation f9cc4840-8223-4c2c-bdc9-0f588d216654 · outbound

This paper cites Nature , year=.

On the Policy Convergence of Policy Mirror Descent Methods Nature , year=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:a2b28902286b35e01657ac03345ef46039a707180559e0e34d0a01d37d55b307

Observation 33c29a33-25c2-4ecc-a08e-fa74a89b9b73 · outbound

This paper cites an unresolved cited work.

On the Policy Convergence of Policy Mirror Descent Methods Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:6800f0db147683c86e7846f6f251d63b08aaabf4f634d2fa9ba8d01d2542d290

Observation 2b6985ed-aed0-4aa4-a835-b7dbeb7e6aff · outbound

This paper cites an unresolved cited work.

On the Policy Convergence of Policy Mirror Descent Methods Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:a7dff15d92fb096748b3db471c7335f0d50bf002882373456c453a14c1b58025

Observation 01b460f7-f249-47a6-94f7-e0d48acbf6b2 · outbound

This paper cites Operations Research Letters , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Operations Research Letters , volume=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:fef92f04f579932ca16f257ef8ede9c5ea895de55f4318545b9f7898bad56359

Observation ef7792f8-5f95-4745-be5b-15bfda397f15 · outbound

This paper cites Continuous control with deep reinforcement learning.

On the Policy Convergence of Policy Mirror Descent Methods Continuous control with deep reinforcement learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:ef8aaebe1e129b695ad0d8ca966f9000bb3e4b5b224344cfc668adaf122ed02e

Observation 7578b6b1-1ba4-4ccc-91fc-a378b30a19ce · outbound

This paper cites Advances in neural information processing systems , year=.

On the Policy Convergence of Policy Mirror Descent Methods Advances in neural information processing systems , year=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:e79f19af6bd87642ba34567952d4d8f2b56827f67b6d03af70610b0d75bd3c6c

Observation d3c3ed6e-a074-449b-92bc-82328076a10b · outbound

This paper cites Mirror Descent Policy Optimization.

On the Policy Convergence of Policy Mirror Descent Methods Mirror Descent Policy Optimization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:53ac331c74942a859b6a973091cf953a63c03dc461ac778c3cbab4d14c4bdf32

Observation 2d990d7e-bac3-41b0-b33f-164373c63ec6 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

On the Policy Convergence of Policy Mirror Descent Methods High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:71acad37c0e5d7a321f2c8e244432cc5eff57157c8f1af1bf1ecd24fc9e5e190

Observation 8a6cb66e-52a6-4468-a52d-72cb33f12bfb · outbound

This paper cites International Conference on Machine Learning , pages=.

On the Policy Convergence of Policy Mirror Descent Methods International Conference on Machine Learning , pages=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:7013a848b32038aadd6c28d72a8ed61029e53e79322ba6fd3247fe17f89c492b

Observation 3f3757d3-44f2-4fe6-8838-e40089b87ad8 · outbound

This paper cites Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs.

On the Policy Convergence of Policy Mirror Descent Methods Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:e05f08492b0a980e3ef0a9a39c1716d9ad75dc816e4c272608d15986787517ab

Observation 49e71c0e-1362-4504-abb5-95187d5d08f6 · outbound

This paper cites 2019 IEEE 58th Conference on Decision and Control (CDC) , pages=.

On the Policy Convergence of Policy Mirror Descent Methods 2019 IEEE 58th Conference on Decision and Control (CDC) , pages=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:42498ab6b5cb9d943ac2cf3bc342921192186be1a96f0ef2a41f3ee2ca2c3dc8

Observation 521377fe-2649-4037-ac7e-af8926112511 · outbound

This paper cites SIAM Journal on Control and Optimization , volume=.

On the Policy Convergence of Policy Mirror Descent Methods SIAM Journal on Control and Optimization , volume=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:c635a1af76e280b0c45c029805fcb4da6a27405f6088d962942eceb56f5545ea

Observation 342a263b-952b-469c-8926-c28e0ab8838a · outbound

This paper cites Structure Matters: Dynamic Policy Gradient.

On the Policy Convergence of Policy Mirror Descent Methods Structure Matters: Dynamic Policy Gradient

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:fb206d38eb6520479df800e275ba5d5958faf0180299a61f3febe3c1dd698884

Observation 7d9eaf87-8d9a-486e-897c-c9ce372d42b0 · outbound

This paper cites Advances in neural information processing systems , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Advances in neural information processing systems , volume=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:14a6989cfb013ab10375d5e2bc84fabd310d75b36f2f470129dd4963d610111f

Observation f830d578-4a73-4ad0-bca5-4e894cad8f3f · outbound

This paper cites 2003 , publisher=.

On the Policy Convergence of Policy Mirror Descent Methods 2003 , publisher=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:bab25977c5160f4e1effb34515a20adf6616cc1b9a1ffd922fc52d215e05c4e2

Observation f395aff0-9ea2-464c-aced-9610672671d3 · outbound

This paper cites SIAM Journal on Optimization , volume=.

On the Policy Convergence of Policy Mirror Descent Methods SIAM Journal on Optimization , volume=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:9d17a326bd698c7b0c5664c87e23ab7601d1d245da1273e3b2bdb37db69a7b4a

Observation 0f8ef427-73ed-4001-9118-11f252a778cb · outbound

This paper cites International Conference on Machine Learning , pages=.

On the Policy Convergence of Policy Mirror Descent Methods International Conference on Machine Learning , pages=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:777eb8d648a488b409edba21f902f1395b4f8d1eafa33a95b7f9498706c23616

Observation 512003c4-5fc3-484f-807c-0428c3b5dbe9 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Advances in Neural Information Processing Systems , volume=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:366759a71bcfaa1f5bb5ea6dd043e3d12ad8a3d6db73aa3130260594516d0a25

Observation 652097ef-63fc-4676-9b7f-82e914a7b81a · outbound

This paper cites A general class of surrogate functions for stable and efficient reinforcement learning.

On the Policy Convergence of Policy Mirror Descent Methods A general class of surrogate functions for stable and efficient reinforcement learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:ed2f70e28bc13e56b5b1977d09cec1d3f4262d8ece9cae2f3b7206f4c74d4a34

Observation ffbd5812-beb4-4135-861d-d07ddfa2cbe8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Advances in Neural Information Processing Systems , volume=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:2a5dd712b81c414b97a58b6aed0d31640d741c25b6dd04c6efe11b123e1200a5

Observation e086921c-112e-419c-bddf-fac9531fbfa2 · outbound

This paper cites International Conference on Learning Representations , year=.

On the Policy Convergence of Policy Mirror Descent Methods International Conference on Learning Representations , year=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:a0a428019513d9bdab467f6451766b38d8bf2ff8145cdf51c7d2c0b09b54d843

Observation 3cc7e455-32f6-4856-95db-640e86a392a6 · outbound

This paper cites International Conference on Learning Representations , year=.

On the Policy Convergence of Policy Mirror Descent Methods International Conference on Learning Representations , year=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:4ce06e7cae9bbd6795ba9be6374dc3f83450b2a8aca9597d2b28c60deb070a16

Observation 02e600da-1f06-43ed-8239-849df5a9fcba · outbound

This paper cites International Conference on Artificial Intelligence and Statistics , pages=.

On the Policy Convergence of Policy Mirror Descent Methods International Conference on Artificial Intelligence and Statistics , pages=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:25a13d7a29f5e57c9c14246b3db8cf210f49edebf202f08a4cc7a20111cdcd97

Observation 33863655-828d-40da-b7dc-1f73ec9125d7 · outbound

This paper cites Functional Acceleration for Policy Mirror Descent.

On the Policy Convergence of Policy Mirror Descent Methods Functional Acceleration for Policy Mirror Descent

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:74d9b4f897b8144c0cedb3c5ec00e862e6afe2c18a4d266e934f75a25db02c68

Observation feb4ff35-db7e-4997-8b25-7541486b7036 · outbound

This paper cites 1998 , publisher=.

On the Policy Convergence of Policy Mirror Descent Methods 1998 , publisher=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:293c6522fe40d749027d34d25c06c12a2073ce80438a4aa1fef7849d9dd69200

Observation 6c98bb8d-ad7b-48d3-a23a-ef0c059e9f9b · outbound

This paper cites International Conference on Machine Learning , pages=.

On the Policy Convergence of Policy Mirror Descent Methods International Conference on Machine Learning , pages=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:b45ab51b9a2165748f74196735e400db7386462815ef8ae6bea81113859140a2

Observation b462ec1f-2d39-4283-bb03-b81a208fd488 · outbound

This paper cites International Conference on Artificial Intelligence and Statistics , pages=.

On the Policy Convergence of Policy Mirror Descent Methods International Conference on Artificial Intelligence and Statistics , pages=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:b22c8b844e3e67a763ffa4bc521d98f8fc730cfd23ead0d67c4e4e98ae041d6c

Observation e8be1368-e58d-45fb-b4b8-0ce5bffdbfbd · outbound

This paper cites Journal of Scientific Computing , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Journal of Scientific Computing , volume=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:d44bd0ba16bd78220b4048efd0304fcf7119505a2ae900b443b2eef417484fdd

Observation 68d38a77-f8ef-4b9a-9c11-74da427bae41 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Advances in Neural Information Processing Systems , volume=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:8757074726acdfee49cf4f0669159b15b30a4b3d4ae82f719616754926de7095

Observation ef87b3a3-b0da-4125-b4ea-8d34e8fdc074 · outbound

This paper cites Mathematics of Operations Research , year=.

On the Policy Convergence of Policy Mirror Descent Methods Mathematics of Operations Research , year=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:b790e8e47c606d9f68aeb7cdb9c3d3799481302ccebf5365765a4e377c510e87

Observation aea7405f-b149-4395-8d27-432d03cc3129 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Advances in Neural Information Processing Systems , volume=

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:78228a93698ef5c905001674d9a8e45993af532421be047a09ba4f2318b9c44d

Observation 5b6b2522-4984-4d28-a942-a3ae84565144 · outbound

This paper cites Journal of Machine Learning Research , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Journal of Machine Learning Research , volume=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:2b2c6dcfab815b7a9022e65e119e529d71a3e9c3106fd846a0d6a2e6a89fce26

Observation 0a5d14ad-1d4b-4f08-8e00-15fdb3b7963b · outbound

This paper cites International Conference on Artificial Intelligence and Statistics , pages=.

On the Policy Convergence of Policy Mirror Descent Methods International Conference on Artificial Intelligence and Statistics , pages=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:c52a6302e80470353b99e18ebfcd13ae065dd72526683051840f34cd2ec748b4

Observation a32aafb6-2923-4a16-b100-2938738c4185 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Advances in Neural Information Processing Systems , volume=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:e7cb42fb16edd430383aa015df2bd6c4bb2789d297367bcf7af17d545eadef0a

Observation bafe69ab-b98e-476c-b10e-6e970efe7bab · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Advances in Neural Information Processing Systems , volume=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:7cbd4be0d24c2bf4508e7e31cf36b2854e6394b945968e3a5677de6782ec9a60

Observation 8b04eb64-53cf-4b12-81ba-ec1800dcb771 · outbound

This paper cites 2025 , booktitle=.

On the Policy Convergence of Policy Mirror Descent Methods 2025 , booktitle=

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:82a2006f25601c92c3983da107eccb821a7fe22a210b71dc7a6367e6cbc2a5eb

Observation 4eccd9a3-7b03-4589-9d17-12ca0fecf2b2 · outbound

This paper cites IEEE Transactions on Automatic Control , volume=.

On the Policy Convergence of Policy Mirror Descent Methods IEEE Transactions on Automatic Control , volume=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:efa5cfc63769c1d43d94042b4803ff1798e2a2de6d1689b87b0ea76bdf6fe28f

Observation a1cf2639-2dfc-4647-b8b8-e0fda75feb45 · outbound

This paper cites SIAM Journal on Control and Optimization , volume=.

On the Policy Convergence of Policy Mirror Descent Methods SIAM Journal on Control and Optimization , volume=

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:75d58e634a37e888b16fe6a87ec1acb7f84d0f4c090a045e31cec1f3e1157536

Observation e7f02beb-4ebe-429b-8808-f4dd854d7106 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Advances in Neural Information Processing Systems , volume=

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:2f78e17ec2876e995b0396d8649cfea531763551a5dc458c4de4fc62eade087a

Observation 7b168b0a-d3fe-4286-acb4-8801f14da9b3 · outbound

This paper cites Advances in neural information processing systems , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Advances in neural information processing systems , volume=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:898f1af5fe8ecd872fdb9228f82f2986b6c14d01bbacf9ee56cd7ba66ffccf0a

Observation 470aba16-58b9-416b-8cbe-ac56273bec8d · outbound

This paper cites SIAM Journal on Optimization , volume=.

On the Policy Convergence of Policy Mirror Descent Methods SIAM Journal on Optimization , volume=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:b6fe33909945757ee1a84ba3d3867c9a396cc73ea14576a847052811976057c1

Observation 33fd2a0c-302f-47be-955e-13268ecbd8d6 · outbound

This paper cites International Conference on Machine Learning , pages=.

On the Policy Convergence of Policy Mirror Descent Methods International Conference on Machine Learning , pages=

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:62f74cff595aa772060d2aeaae04afc4e6b4b60b98ef8bddcbd9f45ec71b869e

Observation f4ca55f0-41a9-48fa-b4b4-36ee7f30c021 · outbound

This paper cites International Conference on Learning Representations , year=.

On the Policy Convergence of Policy Mirror Descent Methods International Conference on Learning Representations , year=

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:a928b95789d2c8b07fc4094c99cff9abbbdd9b89df77302950e8aa9c963f73e4

Observation f95e448a-d11d-43dd-852d-e375044c556e · outbound

This paper cites Non-asymptotic Convergence Analysis of Two Time-scale (Natural) Actor-Critic Algorithms.

On the Policy Convergence of Policy Mirror Descent Methods Non-asymptotic Convergence Analysis of Two Time-scale (Natural) Actor-Critic Algorithms

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:64af2cd2b9d732d95f891b3f9a74d942dfd87c8fb17d13a4c32c22ded68096dc

Observation 867cac06-44df-4423-9ebe-c6aff6e4d351 · outbound

This paper cites International Conference on Machine Learning , year=.

On the Policy Convergence of Policy Mirror Descent Methods International Conference on Machine Learning , year=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:2d398f8a8b6f2a78efd26c67817df9f9a92b7adfe6cf5f34d1a532988acd83f9

Observation a78ac04f-5113-46ae-a60b-c08b2d947af8 · outbound

This paper cites Neurocomputing , volume=.

On the Policy Convergence of Policy Mirror Descent Methods Neurocomputing , volume=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:72d59604a49d5769e0451170957cba3fedbf574390b95e518cabb8f7c05b5dcf

Observation 50288f46-9d08-4bfc-b8ff-907c79171fce · outbound

This paper cites Transactions on Machine Learning Research , year=.

On the Policy Convergence of Policy Mirror Descent Methods Transactions on Machine Learning Research , year=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:b2042175c1d362460df1efd9ca92455fdf5b6b8bcb12bb2a05a8e790af119a82

Observation 4dd62a56-6969-42fd-8f10-7c9486b19b07 · outbound

This paper cites Finite-sample analysis for.

On the Policy Convergence of Policy Mirror Descent Methods Finite-sample analysis for

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:1a1cd9bee68eda5d5f6fd84ec6733c44f45bddc1d80845e28c47f241482d96b0

Observation 69f3772f-5ee6-40cd-8adb-5b6f2330c69f · outbound

This paper cites 2021 , journal=.

On the Policy Convergence of Policy Mirror Descent Methods 2021 , journal=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:c94ae9a398f504d76a9e3f65acda25ed96ef153968a5298846af824fdee62f06

Observation adcf5910-d8bb-41e0-900b-cada2093544c · outbound

This paper cites Bernstein-type inequalities for.

On the Policy Convergence of Policy Mirror Descent Methods Bernstein-type inequalities for

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:4c8b0a1d2f818bf8cbfb3b98843a3979b223e5cad55cc72bcdaad302a68b7227

Observation 533d3da6-0b4d-4914-97ea-d3ba052ba6d4 · outbound

This paper cites Conference on Learning Theory , year=.

On the Policy Convergence of Policy Mirror Descent Methods Conference on Learning Theory , year=

Reference 100

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:db446db0cbbc249cea4e48c924f9c00d437b9de839d17c9e229467aa257fcae7

Pith citing papers

No inbound Pith citation observations are available.