Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:53:15.489700Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2506.07054.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:53:15.489700Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-16T08:12:57.430291Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-16T08:17:36.523927Z
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 56c0e748-0fab-45bb-bd14-b08996d7944f · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the theory of policy gradient methods: Optimality, approximation, and distribution shift.The Journal of Machine Learning Research, 22(1):4431–4506, 2021
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3422e75-af8d-4b42-81e8-c13b283b4395 · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Natural actor- critic algorithms.Automatica, 45:2471–2482, 11 2009
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2a85ad48-e844-4c5d-932c-ce636f67b2a4 · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Finite-time analysis of single-timescale actor-critic.Advances in Neural Information Processing Systems, 36, 2024
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 43763294-3e93-43a4-b32b-352057071b5e · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Bellemare
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 81b3254c-48b1-4fee-9ddf-acd97723552c · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Beyond the One Step Greedy Approach in Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 51ece1e1-2ea5-4ff7-9766-b6dc50795aea · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Kakade, Karan Singh, and Abby Van Soest
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 57b80b14-6b61-46d9-83f0-f94a58620f0d · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Actor-critic algorithms
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation af50b33e-cec3-45cf-a19d-9e94afbfea7e · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy gradient with tree search (PGTS) in reinforcement learning evades local maxima
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b7909ac8-de03-407f-80da-d6f6a105ef17 · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a08650c2-ebf2-4ecf-9276-7e15532c7613 · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Elementary analysis of policy gradient methods, 2024
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc90bb24-e7bf-416f-8c44-9edb6cbf28f3 · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the global conver- gence rates of softmax policy gradient methods
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1e918b11-ad93-439d-8d9c-730f3fba57df · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Rusu, Joel Veness, Marc G
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7eb16917-a3a7-4ac4-8ac6-0ec7260376bf · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy mirror descent with lookahead, 2024
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a2d3ff54-86d7-456c-9820-2ef2c66ea579 · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead John Wiley & Sons, 2014
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec476048-53ed-420f-823f-3c0793c92f18 · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Trust region policy optimization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34db7301-cfab-47c5-b2da-56ab6f3284c1 · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Jordan, and Pieter Abbeel
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 68dbb265-dee3-47f6-984a-c97c6eda7e5c · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362(6419):1140–1144, 2018
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3dcbdbcf-1631-4f47-aaa6-d304346c9057 · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Sutton and Andrew G
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c23fe72e-c3e4-43d5-a58f-114fe4743eb8 · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy gradient methods for reinforcement learning with function approximation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 95b5e946-8bc8-42ee-9e6e-91def601df2d · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy gradient methods for reinforcement learning with function approximation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a92f577c-8c28-4350-b9fc-696876e7c04e · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the convergence rates of policy gradient methods, 2022
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 86e9b9f1-d77b-4dec-a355-26195e177e7e · outbound
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the convergence rates of policy gradient methods.Journal of Machine Learning Research, 23(282):1–36, 2022
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ada1f7a4-89f3-4e80-98c2-0ba288ee1c01 · inbound
Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.