Pith. sign in

Paper Citation Record · LEDGER

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead

As of 8 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2506.07054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07054 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:53:15.489700Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T08:12:57.430291Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T08:17:36.523927Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56c0e748-0fab-45bb-bd14-b08996d7944f · outbound

This paper cites On the theory of policy gradient methods: Optimality, approximation, and distribution shift.The Journal of Machine Learning Research, 22(1):4431–4506, 2021.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the theory of policy gradient methods: Optimality, approximation, and distribution shift.The Journal of Machine Learning Research, 22(1):4431–4506, 2021

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:15.402520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:53:15.402520Z digest=sha256:c688634f69d20c7bb9c731afbfb045ee86fdbddcdd1003c300029d3b8eab21d6

Observation b3422e75-af8d-4b42-81e8-c13b283b4395 · outbound

This paper cites Natural actor- critic algorithms.Automatica, 45:2471–2482, 11 2009.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Natural actor- critic algorithms.Automatica, 45:2471–2482, 11 2009

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.739448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.407155Z digest=sha256:a2a11fe3353c1bc665aa300e02aed40b0d0a0005e12dd88a0d82404a267bbe31

Observation 2a85ad48-e844-4c5d-932c-ce636f67b2a4 · outbound

This paper cites Finite-time analysis of single-timescale actor-critic.Advances in Neural Information Processing Systems, 36, 2024.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Finite-time analysis of single-timescale actor-critic.Advances in Neural Information Processing Systems, 36, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.728469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.411205Z digest=sha256:b91c3fdd5025af0f1504eaf413fde698360901db80d223db6eae89c8e0dbd99c

Observation 43763294-3e93-43a4-b32b-352057071b5e · outbound

This paper cites Bellemare.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Bellemare

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.717309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.415171Z digest=sha256:a3f2140417703ec85af63fbb24ce6b2b11c776408db2f45f3f180824aed44604

Observation 81b3254c-48b1-4fee-9ddf-acd97723552c · outbound

This paper cites Beyond the One Step Greedy Approach in Reinforcement Learning.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Beyond the One Step Greedy Approach in Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:53:15.526463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.419078Z digest=sha256:9fcb12bbe118b03140b81045e77b393e7a5a7271dffa7d4dfa2b952e7b02b389

Observation 51ece1e1-2ea5-4ff7-9766-b6dc50795aea · outbound

This paper cites Kakade, Karan Singh, and Abby Van Soest.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Kakade, Karan Singh, and Abby Van Soest

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.706317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.423850Z digest=sha256:8e0ef538dbb5403e2fda23b63af95fed730f8e985043441b1ff54b8fc0fe99e3

Observation 57b80b14-6b61-46d9-83f0-f94a58620f0d · outbound

This paper cites Actor-critic algorithms.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Actor-critic algorithms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.695828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.428386Z digest=sha256:d47d148a1eeaef8fd098906f3e3622c4e34be4de322528d2367f60184318fc2f

Observation af50b33e-cec3-45cf-a19d-9e94afbfea7e · outbound

This paper cites Policy gradient with tree search (PGTS) in reinforcement learning evades local maxima.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy gradient with tree search (PGTS) in reinforcement learning evades local maxima

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.684281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.435726Z digest=sha256:87f9afbc3f69942abe0dcdba492b8d13a4ed160d64a166b6219fdc65da8c22e4

Observation b7909ac8-de03-407f-80da-d6f6a105ef17 · outbound

This paper cites an unresolved cited work.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:53:15.672487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.439433Z digest=sha256:8890e4a7df7fdc5b975815f634a49d65aaa88a92119558f3992dcee378b503c2

Observation a08650c2-ebf2-4ecf-9276-7e15532c7613 · outbound

This paper cites Elementary analysis of policy gradient methods, 2024.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Elementary analysis of policy gradient methods, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.662105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.442956Z digest=sha256:af3fb66df09d7d003ae8dcf37dae80483ffaa8d787cced907e2ef1210c08e677

Observation bc90bb24-e7bf-416f-8c44-9edb6cbf28f3 · outbound

This paper cites On the global conver- gence rates of softmax policy gradient methods.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the global conver- gence rates of softmax policy gradient methods

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.651383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.446740Z digest=sha256:113f4b7b3c7734ffe156b53124452bd2e7fa204d595157ee6fd40a414d0014e3

Observation 1e918b11-ad93-439d-8d9c-730f3fba57df · outbound

This paper cites Rusu, Joel Veness, Marc G.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Rusu, Joel Veness, Marc G

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.640881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.450997Z digest=sha256:72b7c5441c572ebc21c90ba5d7c2030487f6a6199557a2628c18827037199aba

Observation 7eb16917-a3a7-4ac4-8ac6-0ec7260376bf · outbound

This paper cites Policy mirror descent with lookahead, 2024.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy mirror descent with lookahead, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.629318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.454437Z digest=sha256:143faf7dbf205bf3ec951b123984c7699c05b9c307c20281ae415c3c41f06cc6

Observation a2d3ff54-86d7-456c-9820-2ef2c66ea579 · outbound

This paper cites John Wiley & Sons, 2014.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead John Wiley & Sons, 2014

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:15.457879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:53:15.457879Z digest=sha256:2eb6b475afe62d6e30474ab3d929f625630d6d728104e074c56612df6c706fc0

Observation ec476048-53ed-420f-823f-3c0793c92f18 · outbound

This paper cites Trust region policy optimization.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Trust region policy optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:15.461456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:53:15.461456Z digest=sha256:a6b2a421c7cf9d7104ddf3a16887a5aa2a74eed76a640dc442bd134ac2de9e8f

Observation 34db7301-cfab-47c5-b2da-56ab6f3284c1 · outbound

This paper cites Jordan, and Pieter Abbeel.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Jordan, and Pieter Abbeel

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.604094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.464987Z digest=sha256:393163295aec2653f70cf6cd2a6f203b201f8ee5d13e5249a61238ce413801ad

Observation 68dbb265-dee3-47f6-984a-c97c6eda7e5c · outbound

This paper cites A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362(6419):1140–1144, 2018.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362(6419):1140–1144, 2018

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.593271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.469607Z digest=sha256:1c372873fb20a68c09fc17d2a83b51904f06a41b39363b3666785ec8c1bc6c8c

Observation 3dcbdbcf-1631-4f47-aaa6-d304346c9057 · outbound

This paper cites Sutton and Andrew G.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Sutton and Andrew G

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:15.473939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:53:15.473939Z digest=sha256:06984d3614ed76410292899e47261a343d9d70809c79a6a9f7131b02937200c2

Observation c23fe72e-c3e4-43d5-a58f-114fe4743eb8 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy gradient methods for reinforcement learning with function approximation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.572872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.478052Z digest=sha256:95cbdb67e9b42be4cfecc73ff673b10500c2f20c8bfabe9debd79258f8de4516

Observation 95b5e946-8bc8-42ee-9e6e-91def601df2d · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy gradient methods for reinforcement learning with function approximation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.561490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.482171Z digest=sha256:a8135d9cf8da9611622e050c27bf437d269a4a4edf2fe9a60ae3e138640622a3

Observation a92f577c-8c28-4350-b9fc-696876e7c04e · outbound

This paper cites On the convergence rates of policy gradient methods, 2022.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the convergence rates of policy gradient methods, 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.549983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.485818Z digest=sha256:c7ff761f57aebdd4332f7c82fec40c94122273b4e523d2b24918a14896c01962

Observation 86e9b9f1-d77b-4dec-a355-26195e177e7e · outbound

This paper cites On the convergence rates of policy gradient methods.Journal of Machine Learning Research, 23(282):1–36, 2022.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the convergence rates of policy gradient methods.Journal of Machine Learning Research, 23(282):1–36, 2022

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.539115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:53:15.489700Z digest=sha256:2d6f7788d61b34bf4ee51cfd312a12fde886a16273e12e073182c6d825a4519b

Pith citing papers

Observation ada1f7a4-89f3-4e80-98c2-0ba288ee1c01 · inbound

Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum cites this paper.

Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:17:36.525446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:12:57.430291Z digest=sha256:21c62f168c51162a660ab642a8e9a512855cee4a4129ac8ee86ad709ca8482a7