Pith. sign in

Paper Citation Record · LEDGER

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying

As of 5 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 1 inbound Pith citation observation for arXiv:2606.00151.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.00151 v1

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T23:51:47.193267Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:41:12.548554Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

90 of 90 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved89
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f396977-e942-4c10-a86e-1184be9f322a · outbound

This paper cites and Barto, Andrew G.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying and Barto, Andrew G

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:9f2d3587cfd8a5f6066ed750ee557c33849a87dd9be8705d6779d3c552656c03

Observation 66aa3f75-3006-47aa-a333-3a36e9217675 · outbound

This paper cites Why generalization in.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Why generalization in

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:5bf79a0666c1cadc054572a65c76ee6ae046bc7c32266911551e86a2421be432

Observation 04384a00-28f4-467f-97f1-13ca01c50d49 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:807ecff06ac60d895645e8c45ea36c303910898722a3c47d783bc6228f2b6b80

Observation 53c2b838-1da1-4cbe-b1b7-698de14a6aff · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:ee9de2f16e443e8d7dfb2b1902236d291df5d824e47d2635a709572d22d0af85

Observation 88d319e0-6419-46b4-9c04-9e97da1d74ef · outbound

This paper cites Strehl and Michael L.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Strehl and Michael L

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:743947a5be08c0be4bf7765c1205497478e5e2e280611280a1547d68c647c382

Observation d1a75296-f3b7-4d3c-969a-4a1c8f41af11 · outbound

This paper cites and Li, Lihong and Wiewiora, Eric and Langford, John and Littman, Michael L.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying and Li, Lihong and Wiewiora, Eric and Langford, John and Littman, Michael L

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:4ae630d52c00d462c6e2469068fe479763e5a90993032d010b81b28453b37a92

Observation cac9b50f-16f8-4f48-b2e8-79931f21a1b3 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:242469b727804e316c2fde300b166f840412222a09f0be55b0e337c502dd1091

Observation a07c8584-9b9f-4232-bdd2-edf75d4e10bb · outbound

This paper cites and Littman, Michael L.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying and Littman, Michael L

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:5ee20c84ec3c9503bbc99818976f209a39e6f878e98f63658b4a702b298d78fc

Observation 59aff438-2a21-48df-9140-6e78e23b3113 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:26ff8cd381ee4b46abf915bd06b60b04284f7c29aa12801e6c6ad36c49581e4b

Observation a1898857-5169-4006-8d7f-f43f85ffee79 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:fc6a86f0fef0bc66ca819228102273f53271446797d9051c75e004a23f4b5801

Observation c6d2efc7-f9e1-4c26-8b9a-95bb7e50b997 · outbound

This paper cites Stadie and Sergey Levine and Pieter Abbeel , year = 2015, journal =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Stadie and Sergey Levine and Pieter Abbeel , year = 2015, journal =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:15c20d18bac4f42cc6e0dd08fc23be9bd05e5a803b6331e0ab037ba97fd72bac

Observation f742a8ed-324e-4f92-991b-4621a0d5ba31 · outbound

This paper cites and Darrell, Trevor , year = 2017, booktitle =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying and Darrell, Trevor , year = 2017, booktitle =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:7be31b05b4e18f85e143cc7a8f4412e3d5e1409063a60b5d45056cbcca74ff4b

Observation aadd9723-2509-4f22-b7c2-0a3a556546d5 · outbound

This paper cites Thrun and Knut Möller , year = 1991, institution =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Thrun and Knut Möller , year = 1991, institution =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:aa0a507000f089bf8b01b223ab6c3e75cacb3b5159b4650e139a8cade59796af

Observation ee5ab933-a037-4513-9aa9-7314366988a5 · outbound

This paper cites Gomez and J.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Gomez and J

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:c3cf52dedd550337e311b9fda9846115280752467026b5eb4dd35454245b8f30

Observation d9174b7f-74b8-497c-ba17-464e9e84c7ac · outbound

This paper cites IEEE Transactions on Autonomous Mental Development , volume = 2, number = 3, pages =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying IEEE Transactions on Autonomous Mental Development , volume = 2, number = 3, pages =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:55b7e5af3d65e4585bfe5363689b459115f3d909a99483f346ff7d14124c89d8

Observation 260a3bf3-24f5-4659-a5f2-adadb7e503f8 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:10306fc27b29ac825887f82b31c0631501cacc9af7b6c95541a9b23139cbc460

Observation 15f6e43b-1e0c-437a-aa02-274ca91c8b7d · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:bd3f439bea4ef4c8ea8fa0f139798d578d6faf7e4447e8d7cb80f53bb402a0fd

Observation f2a274f8-109e-4918-b5e7-e00223366ff1 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:207f3784472275b75eec34c4c8123c6feea517d109e57fb383170fe6cafccfa5

Observation 2b18caf4-3be9-4d06-82e7-872744a98fad · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:d82f5f1ef71b7e28aa2d8d215b703673d0ce2015f9f1b75ac2485d88883ce8d5

Observation a89c7a21-e6ef-468d-bb3a-d4a49f9f35e4 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:9ea3b1ad03d3f9f6c135bfa77043dc5f2e364fe453c327f1a0cf2c273606cbde

Observation 1bce3a46-e364-462e-9080-4b1e623d6d3b · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:155df175e31cbc4bfe48c8e966c9525f28a502c8f65a9ec994fff2b3d8e093e7

Observation 7b4f4011-ef96-42be-9425-d3cf164cb876 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:dc4bbabd2bd2a99f46fb87337385829d1638878dd91e8e1882879dd87ecf21c2

Observation 86e05baa-62e5-4ccc-baf5-95732d665c2f · outbound

This paper cites Nature , volume=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Nature , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:c26a4abdd773ae64c221aab027ae87af08557c4a04bd7b8aae2d17d8b9144739

Observation a93fc4a4-ba43-46d4-90d1-3b61193c6348 · outbound

This paper cites Williams and Jing Peng , year = 1991, journal =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Williams and Jing Peng , year = 1991, journal =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:3a0ac8ded7bc60a42da8dd594412bcb0dbc04b3505886831296753e18ada1b67

Observation a8c38882-5a73-40eb-8231-ec2e54277c87 · outbound

This paper cites Machine Learning , volume=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Machine Learning , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:1a187ff58568f920bb4753525cc53cab94dd5c2758b6037ce58c6fa17110677d

Observation d8f6ec25-bb09-4f70-9763-9eee5a33c29a · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:619abe25dee849455d4d4f5d1667768136a214d8a2d3c6b808659d49607b53b5

Observation 9b423f32-d9db-4b41-81a5-6244e31b96ba · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:0c991670de56ca8c7035cf90038433eab5cacdeac86674b82a00b9fc4630a88f

Observation 945c1c00-efde-4466-b982-b7ca67106a54 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:0758e0c21cbfdd9865f773279ccbb64dacc532c8795300b97dd8c163783aa596

Observation 12e31e08-fe0f-4649-999e-a1df50a6c2e0 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:8f7ac6f01e2884d033b53936cdaef4f682dee7b4d6ec9c1b56e2bb4218adbfce

Observation 204384a4-640b-4a96-8e57-f60de5cc534d · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:5c96efdb57615a31d247fb29fa0b1ae99aa6c71acdd246b8bc22b0119ca89f08

Observation 0d490338-b480-4b78-8606-e7bf3ac08c10 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:34033303b43232868a72e0108eb881711e8503bf2e0b2ffd21c240c8d7d2be31

Observation cb4fbd68-dfe6-4475-9fe9-c34b3080ec08 · outbound

This paper cites and Maas, Andrew and Bagnell, J.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying and Maas, Andrew and Bagnell, J

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:70278d2db20b610e06b36143f5860ca008cb72ae1f48d1d8938cad36b34f3854

Observation 41f463b7-0ea3-4376-8435-6a0976766924 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:7b9b813f6004a33746ded8124bd09cebb04ef72569d8df09dd169e5a918cfe5b

Observation 11c9b177-90b0-415f-a706-9c1ea5596376 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:6eac5644f9d49d1c2a66608095ef4ca39ab81cff63c244617e1d84eec204ea5a

Observation dc57d2cd-0c27-4159-8e10-8f6b87a00c12 · outbound

This paper cites International Conference on Machine Learning , pages=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying International Conference on Machine Learning , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:2fd61857151768c56a559a74f6c9ee698b6a875d48b431271907c19c0805deb3

Observation 8037c2df-0873-45ba-88e9-13d5aa173e85 · outbound

This paper cites Deep exploration via bootstrapped.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Deep exploration via bootstrapped

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:a5dfb22ca0ce5bffe990b21e4c67c0794a66b0102a5060e7b0fddd910a24cd42

Observation bfe5d583-bb38-4acb-8b69-969becd2d622 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:fcc698324cad6c4919ff3048cac8fe7c599ff2aec3f0ed0a5c64772042c82291

Observation 790c50b6-f7f8-459f-ab4e-83884c727cf3 · outbound

This paper cites Journal of Machine Learning Research , volume = 20, number = 124, pages =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Journal of Machine Learning Research , volume = 20, number = 124, pages =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:6b93d94532385aeadb2d234b05e9e8a9b8c7e7a6901f24f90ea32602431e2089

Observation 6e8f8a2e-9272-462f-be9f-6d8a5cff5e00 · outbound

This paper cites Efficient exploration through bayesian deep.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Efficient exploration through bayesian deep

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:2009e2932e26ac7699b468c0688fcb54877e660c34d9617372c68a946289d36c

Observation 438fc4de-93f1-490d-bf94-d21145c9f9de · outbound

This paper cites International Conference on Learning Representations , volume=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying International Conference on Learning Representations , volume=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:9cd7423d3f690d24dc961a68f3d74de667508939f091dd46d18df31c5f07acec

Observation b134b419-7dad-40cb-bf66-5df1a5d7ad46 · outbound

This paper cites Uncertainty in Artificial Intelligence , pages =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Uncertainty in Artificial Intelligence , pages =

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:303bb389c25c74862cca0c206dea120b6ec33a0819ace52d6598d4c66b8aa8b1

Observation e84a7a12-3701-4ae3-8653-c9a0fb5224da · outbound

This paper cites International Conference on Machine Learning , pages =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying International Conference on Machine Learning , pages =

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:d1794cd183189188439e4955f8ad04b482aac631ad4f7573864b4f9d17ce3091

Observation 76bf6ead-c15f-4fdc-a12b-5564daebf609 · outbound

This paper cites International Conference on Learning Representations , volume=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying International Conference on Learning Representations , volume=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:3f479047d482c13724e53a6a38e3c783eccaf237daf70fd41b316ea35588fe8a

Observation a9ca94e6-ffd6-42ee-a980-4a2842f250ca · outbound

This paper cites International Conference on Machine Learning , pages =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying International Conference on Machine Learning , pages =

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:4ab98ea2bd6f7b891d8f64660f921a8cacda0953de7309b625eaf6373a1ca950

Observation 1eea2419-29fd-4511-bceb-91dbba4b73fe · outbound

This paper cites International Conference on Machine Learning , pages =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying International Conference on Machine Learning , pages =

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:905832bcce6710f307985a5b1dfdca50cb64742662bdfeae3a16f59b1f2235f5

Observation 44906199-4c35-4447-8ba6-2653663bb128 · outbound

This paper cites International Conference on Learning Representations , year=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying International Conference on Learning Representations , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:bdc51fbc312ce320492fcbce73fcb81444ff166a88930ff51f00bca7667a101f

Observation 6feed398-e68a-4141-bcee-40437d1841de · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:144245dc9c6d81bffbb1a9b4eeb756fcc2109c701799adc8951ccd37381b06c7

Observation a8971fd8-68e1-40f2-aac3-7acd49148511 · outbound

This paper cites Advances in Neural Information Processing Systems , volume = 36, pages =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Advances in Neural Information Processing Systems , volume = 36, pages =

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:599717bc3f3618bf294a352277cfa542a604960a3e56c3e2722d0ee3243cd59e

Observation a3c6aa87-14af-4f49-a216-95e89642a080 · outbound

This paper cites International Conference on Learning Representations , volume=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying International Conference on Learning Representations , volume=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:7984dd20d77c1bf3b7d82c50cd1186daef267e6d25645e4a37dac1c9cd402f7f

Observation 4770aa04-f0dc-4cc7-addc-5827a65ec306 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:3f1f3ff494a71478bb7d1de7de19e3fb10f73b2241458762a8a2eca35b960678

Observation ebeb77cc-a45a-4db5-bd7d-08656538a854 · outbound

This paper cites Journal of Artificial Intelligence Research , volume = 47, pages =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Journal of Artificial Intelligence Research , volume = 47, pages =

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:77c6d667eb8c26d66533992f2ed79814b156f16af64190aa56b00e68d89fe6e1

Observation f35cec72-c1ac-4dc0-8058-8c0cb3f419c1 · outbound

This paper cites Advances in Neural Information Processing Systems , volume = 35, pages =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Advances in Neural Information Processing Systems , volume = 35, pages =

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:e3ebca44d649b427cb6ed9758eb5f712b9bd4a3e39b4f39cca40bb79930fac53

Observation 9376ef5c-bbed-4566-9521-66e8dd21aa8f · outbound

This paper cites Advances in Neural Information Processing Systems , volume = 34, pages =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Advances in Neural Information Processing Systems , volume = 34, pages =

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:813c5314323fa1cf5ace1c60ea99e05ccb432251114639003c1645ba94c82944

Observation 2ed91220-3dcf-472c-8d1e-47123858796a · outbound

This paper cites Journal of the ACM (JACM) , publisher =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Journal of the ACM (JACM) , publisher =

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:cc5c48ce5f2f67c561415d6ffdf487a403eddfb006070946ad2dcdcd840b8cd1

Observation 9ef9e9fe-2029-4303-ad05-3fce038d6680 · outbound

This paper cites Artificial Intelligence and Statistics , pages =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Artificial Intelligence and Statistics , pages =

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:015b4de38bbe4ffa18dad7731d685e7255603e36717ad92680dece388270274a

Observation 99f1bc58-71f8-4716-be71-3f0adaf9deec · outbound

This paper cites Biometrika , publisher =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Biometrika , publisher =

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:fc25a0bd2816fbb751e11e7d4f16297808244f8e95eec3f1aec7f2f74a71a065

Observation 693e94f0-acb1-405d-90bf-c6a51886ff81 · outbound

This paper cites Machine Learning , publisher =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Machine Learning , publisher =

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:b667cdd6130b52773fdd8a6a1d7be6978864794475312fc50e3e6c678c485946

Observation 2c5ec34c-6783-45e4-8061-4233cc7c5622 · outbound

This paper cites Journal of Global Optimization , publisher =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Journal of Global Optimization , publisher =

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:a2d0a7e123aef7fd6bfae8e64cc7e889bff4bb36e29e9e2e81006f5aff67ffbd

Observation 185c9c26-f80a-4d18-88a8-7f0be66c25ff · outbound

This paper cites Proceedings of the 19th international conference on autonomous agents and multiagent systems , pages=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Proceedings of the 19th international conference on autonomous agents and multiagent systems , pages=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:7cfd1759fbca6c823ab775d948ba129393b601e7ec74066acb6d78552e9311a1

Observation 9f4c9f12-2010-46ba-ab38-3c65f90785ea · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:dea5c402d20853d1fe39d3e1614979669b2b9c1d2ab9bed2a0b27db62b36b65a

Observation 9caadd62-01d0-47a6-a79d-b5c3e8e573e6 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:eaa1dc1fd32eeb37e704facee4664e684aeb198c24fc684dd8964a3e8f31f1ab

Observation b25580dd-a8a4-4149-8471-00528ee90289 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:879233477c681e52dc63065d021e612224b217cf04b3c10a5a409dc908e1bf0d

Observation d75a9694-8abf-4779-b6c0-734acef5c68f · outbound

This paper cites Uncertainty-based offline reinforcement learning with diversified.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Uncertainty-based offline reinforcement learning with diversified

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:d46e09695ccdafe1be887d375f13528846ed2da08a7d54b4feaa16eda92c2ebb

Observation 3df2e2c7-6eae-4595-a26f-20e7a3ecd917 · outbound

This paper cites Maximum entropy.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Maximum entropy

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:9b1fc7488e4896db29633c111d9ecbf7527f1dcc2453ba48256f9d6f424e6a8c

Observation 572f28a8-d80b-4738-8848-00c15138f6b6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume = 33, pages =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Advances in Neural Information Processing Systems , volume = 33, pages =

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:43ba1e084224c73e053be65aa204bf732b8f7b798f57c09967c2e8c609698233

Observation e9697fb8-6d9b-40df-a1c1-fa2ec88aeba8 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:9d1a3e4bca557aee4b44863f9d3e73d5d2fad8252836759caecc0ff2d722df4b

Observation 9ff0a4c3-26f9-40c1-945c-0ec364e656ed · outbound

This paper cites International Conference on Learning Representations , year=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying International Conference on Learning Representations , year=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:3de301590139131bb503f0b7fdd1de4faf06b7f983b17ff0a74abdf9ed8ae9c9

Observation c7633eec-1e77-4b2a-82a6-a345f558e068 · outbound

This paper cites Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T23:52:49.347617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:5992da9d7b0b95e05258ba93b8a2d28f2f6c18935863909c93418e78be38eff6

Observation a57cbaf3-586b-4095-9bc6-b8f2ce585988 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Advances in Neural Information Processing Systems , volume=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:0102b0ca17a0a2c18a034860ff96d6e8464d0807a0842890a02f572c9f254e98

Observation 1f91ad50-d475-4ac4-903e-ea3e0412b2c3 · outbound

This paper cites No representation, no trust: connecting representation, collapse, and trust issues in.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying No representation, no trust: connecting representation, collapse, and trust issues in

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:c79f3d7933213e32d9f6506f0b86de4db671d84e25575c89546423849762e507

Observation 270b1b7d-77b5-41d1-9369-acd4106d393d · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Proceedings of the 41st International Conference on Machine Learning , pages=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:15dbf25d61d06f980ecaba2af57b3a2ae977ee82d7d42daa5408c45f29228e90

Observation c2eab59b-02fc-49f6-a0bd-595c52963911 · outbound

This paper cites International Conference on Learning Representations , year=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying International Conference on Learning Representations , year=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:8e667ee547b2623d1f466232f19879950d049e444a266ce0f7692ff27d227082

Observation e852d4ca-744b-4c80-99e0-1079980139a5 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:be7f4e82eb2bdf2bf2de1df47b1837556647ed64ebd4bf2cdb35f66aab012f66

Observation cea13e32-4892-4b0b-8283-d59626da5d5e · outbound

This paper cites Araújo , title =.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Araújo , title =

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:80896c32bf43365e5c13b520630f5a5cc980fc63e768c9edcedb8f9b78fab6a1

Observation 97b41722-d851-4be0-ba92-03f2985d1941 · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:b0960e7f034f96eba4539ef9a1c159d8b23e5ce70c1e3b7b3c722f71f08f67c4

Observation fc9cc110-affc-4a1b-b501-d8ca5638eac7 · outbound

This paper cites International Conference on Learning Representations , volume=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying International Conference on Learning Representations , volume=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:236254f6edc5f13f9749afb682fadc4049f35a1c24e94d2da1d2a890299f49e9

Observation c95f9b1a-c6e5-4947-8b2e-788a638d29eb · outbound

This paper cites International Conference on Machine Learning , pages=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying International Conference on Machine Learning , pages=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:f2b8c7d7c63341eebdc06fad5abf8537eeab02e098718a7b2ab4b3530c139e5f

Observation fcecc30f-967a-46a8-a6b5-0c3ed007f485 · outbound

This paper cites 2022 , url=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying 2022 , url=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:cd860db3ca8fdddbfcbadbb3f9fe27bdd5c491266c87581911b9fd1877c804db

Observation e70d6bb2-5768-4c06-8a2a-962d293d7981 · outbound

This paper cites Reinforcement Learning Conference , year=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Reinforcement Learning Conference , year=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:6c4dcb0894c6723f32719b7a1df9600f5b631557d7678c6dc793492fcb3fc7ac

Observation c3cdab19-2a70-4cc9-a565-aba07326e981 · outbound

This paper cites International Conference on Learning Representations , year=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying International Conference on Learning Representations , year=

Reference 80

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:2f7d8e5294a3b7c139b3ab69ec2146a963af7feb639be51e05d9f2f3ea5200b2

Observation 3848e43d-653c-47a7-b0d2-db3e19e7c5eb · outbound

This paper cites Beyond Softmax and entropy: convergence rates of policy gradients with.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Beyond Softmax and entropy: convergence rates of policy gradients with

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:88ae60300f325aa2ae5e4678eee3ba93e68586c6b3b36cc30f9666953c62f94e

Observation f42f263e-72b7-456d-b310-839281a78f17 · outbound

This paper cites 2024 , organization=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying 2024 , organization=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:8abc3d879815496335a15d85ba90b4a8bd8187a4feb8eafdb2de39f41f905921

Observation 9360dd9c-942a-4c88-bd0d-325c619d7363 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Advances in Neural Information Processing Systems , volume=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:641fba34680e206dc2816d7ec74ab6153a18fce8bd23ff691b7c6ed3a1e97bdc

Observation 644b7d82-c7aa-4d19-a5a0-708845348a40 · outbound

This paper cites Management science , volume=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Management science , volume=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:9eefc08c0aeef5c93f3b9fdb1a42e860f7bf3221049211f6cb71c16c3ac61319

Observation fc88741a-c101-4a7c-bbf9-652670b5d889 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Advances in Neural Information Processing Systems , volume=

Reference 85

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:d7850ad706cfe7c0181c51372c0d9071bb25ee94da6ffd410a12163dc034dbc0

Observation 3e8c564d-0386-4053-aff1-a869081b1abf · outbound

This paper cites Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence,.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:e85afbbd66f0bcc75bfa7c72db4a489a0f7a55d76f2d18c20280b43a5725a55b

Observation 4e76756f-41eb-48fa-ac7e-06ae6f89fd51 · outbound

This paper cites Journal of Risk , volume=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Journal of Risk , volume=

Reference 87

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:048dad70ffc12261ec5df77602c442353d279c65f1a50bcaa3b334ca86f3818b

Observation 0265642c-0e7b-44d2-a0da-edae386085ec · outbound

This paper cites an unresolved cited work.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:63f852ebeb383a5dced75536e3cd48d4985db945f6c3b9aded0044d3132a8753

Observation b0ded8b7-c7ab-4216-b153-b5a2126a6979 · outbound

This paper cites 2023 , publisher=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying 2023 , publisher=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:13922a948d253b5de2f272ae541442b1316174eb25004da89753e2d427104d67

Observation 0ba91134-924e-42c0-8610-5886b24dc0e0 · outbound

This paper cites International Conference on Machine Learning , pages=.

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying International Conference on Machine Learning , pages=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-06-28T23:51:47.193267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:51:47.193267Z digest=sha256:50fe74dd7329fc6141983bd07dea3b538d3a09454a0a780d5cf29cd7d09d93ea

Pith citing papers

Observation 3f06b6f3-348d-4473-9602-e638dd50a9ab · inbound

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation cites this paper.

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T00:41:12.548554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:41:12.548554Z digest=sha256:d01d06889acca4a416cb57d644f808a6a9b3da7900f0132d39c0ca7565982452