Pith. sign in

Paper Citation Record · LEDGER

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead

As of 10 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2506.07054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07054 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:53:15.489700Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T08:12:57.430291Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T08:17:36.523927Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56c0e748-0fab-45bb-bd14-b08996d7944f · outbound

This paper cites On the theory of policy gradient methods: Optimality, approximation, and distribution shift.The Journal of Machine Learning Research, 22(1):4431–4506, 2021.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the theory of policy gradient methods: Optimality, approximation, and distribution shift.The Journal of Machine Learning Research, 22(1):4431–4506, 2021

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:15.402520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:53:15.402520Z digest=sha256:c688634f69d20c7bb9c731afbfb045ee86fdbddcdd1003c300029d3b8eab21d6

Observation b3422e75-af8d-4b42-81e8-c13b283b4395 · outbound

This paper cites Natural actor- critic algorithms.Automatica, 45:2471–2482, 11 2009.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Natural actor- critic algorithms.Automatica, 45:2471–2482, 11 2009

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.739448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.407155Z digest=sha256:e778e00209453a41498f46a3ab3dd2bf230f85333eace817c3013c5bf5e59e49

Observation 2a85ad48-e844-4c5d-932c-ce636f67b2a4 · outbound

This paper cites Finite-time analysis of single-timescale actor-critic.Advances in Neural Information Processing Systems, 36, 2024.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Finite-time analysis of single-timescale actor-critic.Advances in Neural Information Processing Systems, 36, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.728469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.411205Z digest=sha256:a62dd025276f369427f2ee6da357da86c51a98dd6a34bc707d55f0c6e97576cb

Observation 43763294-3e93-43a4-b32b-352057071b5e · outbound

This paper cites Bellemare.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Bellemare

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.717309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.415171Z digest=sha256:2d0b3392e2e3bfcb404546896918d05c9543c90c2560391adee2b3f6bf77aec2

Observation 81b3254c-48b1-4fee-9ddf-acd97723552c · outbound

This paper cites Beyond the One Step Greedy Approach in Reinforcement Learning.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Beyond the One Step Greedy Approach in Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:53:15.526463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.419078Z digest=sha256:d6102602ab75f31fbcbafcfe606013e709133aaff70e1090dd3b090bbcfb6d95

Observation 51ece1e1-2ea5-4ff7-9766-b6dc50795aea · outbound

This paper cites Kakade, Karan Singh, and Abby Van Soest.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Kakade, Karan Singh, and Abby Van Soest

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.706317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.423850Z digest=sha256:4a16dca67ac0742f795fe34981de651bcbed2ea9886ab2fa83d1df5a3109316f

Observation 57b80b14-6b61-46d9-83f0-f94a58620f0d · outbound

This paper cites Actor-critic algorithms.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Actor-critic algorithms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.695828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.428386Z digest=sha256:dd9b6404119b1958ad1566c993c5b0355d055c686fd8c603d85460f29e2ce233

Observation af50b33e-cec3-45cf-a19d-9e94afbfea7e · outbound

This paper cites Policy gradient with tree search (PGTS) in reinforcement learning evades local maxima.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy gradient with tree search (PGTS) in reinforcement learning evades local maxima

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.684281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.435726Z digest=sha256:e78ee3e96be6699931e64757f88bd62f49ea699fc4a8772be17ae6a7d45da082

Observation b7909ac8-de03-407f-80da-d6f6a105ef17 · outbound

This paper cites an unresolved cited work.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:53:15.672487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.439433Z digest=sha256:5b34e4dbf055491d44f001a4bc97c091fd007a4f13f460c58375b79c4ddc06a1

Observation a08650c2-ebf2-4ecf-9276-7e15532c7613 · outbound

This paper cites Elementary analysis of policy gradient methods, 2024.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Elementary analysis of policy gradient methods, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.662105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.442956Z digest=sha256:9fc9c7343caf6d61f590a35a29bb3d6c4e2e1946df5995f0412b69beaf6e7296

Observation bc90bb24-e7bf-416f-8c44-9edb6cbf28f3 · outbound

This paper cites On the global conver- gence rates of softmax policy gradient methods.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the global conver- gence rates of softmax policy gradient methods

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.651383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.446740Z digest=sha256:e7a5e1ed33f28db95f2811c3a454a2721217398070312333920e6cbd0972a07a

Observation 1e918b11-ad93-439d-8d9c-730f3fba57df · outbound

This paper cites Rusu, Joel Veness, Marc G.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Rusu, Joel Veness, Marc G

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.640881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.450997Z digest=sha256:2c29dd88c4a6cee53c43305011e06df7bbb2e0899852329fc4aeba817f13408d

Observation 7eb16917-a3a7-4ac4-8ac6-0ec7260376bf · outbound

This paper cites Policy mirror descent with lookahead, 2024.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy mirror descent with lookahead, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.629318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.454437Z digest=sha256:ffb7ec7b86000fa3802ec1f1a16cbc74c193e0040986142a0e20a7c7e334a76e

Observation a2d3ff54-86d7-456c-9820-2ef2c66ea579 · outbound

This paper cites John Wiley & Sons, 2014.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead John Wiley & Sons, 2014

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:15.457879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:53:15.457879Z digest=sha256:2eb6b475afe62d6e30474ab3d929f625630d6d728104e074c56612df6c706fc0

Observation ec476048-53ed-420f-823f-3c0793c92f18 · outbound

This paper cites Trust region policy optimization.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Trust region policy optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:15.461456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:53:15.461456Z digest=sha256:a6b2a421c7cf9d7104ddf3a16887a5aa2a74eed76a640dc442bd134ac2de9e8f

Observation 34db7301-cfab-47c5-b2da-56ab6f3284c1 · outbound

This paper cites Jordan, and Pieter Abbeel.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Jordan, and Pieter Abbeel

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.604094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.464987Z digest=sha256:20c4317e703882fca7dcb7684f8e544ee0eaddb6d9fc7f4a2dd4d851698adce9

Observation 68dbb265-dee3-47f6-984a-c97c6eda7e5c · outbound

This paper cites A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362(6419):1140–1144, 2018.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362(6419):1140–1144, 2018

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.593271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.469607Z digest=sha256:74e58e6c12544dd182e6397c972f5366ba319527c114bbb4d725a8ec3e7b986d

Observation 3dcbdbcf-1631-4f47-aaa6-d304346c9057 · outbound

This paper cites Sutton and Andrew G.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Sutton and Andrew G

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:15.473939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:53:15.473939Z digest=sha256:06984d3614ed76410292899e47261a343d9d70809c79a6a9f7131b02937200c2

Observation c23fe72e-c3e4-43d5-a58f-114fe4743eb8 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy gradient methods for reinforcement learning with function approximation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.572872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.478052Z digest=sha256:56a12c8f2caa65df77db54a4603f55cc615e483a1d77b4b81af0b0fa715c50f3

Observation 95b5e946-8bc8-42ee-9e6e-91def601df2d · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy gradient methods for reinforcement learning with function approximation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.561490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.482171Z digest=sha256:9306b48631c3eb1c880d566765ef25a915e275de31ee6ed8313c5d052dd7f78b

Observation a92f577c-8c28-4350-b9fc-696876e7c04e · outbound

This paper cites On the convergence rates of policy gradient methods, 2022.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the convergence rates of policy gradient methods, 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.549983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.485818Z digest=sha256:e7825ea0b681c0e2f9a79a76a50eb0ac140cb20a583b38b9038cb241ff2866da

Observation 86e9b9f1-d77b-4dec-a355-26195e177e7e · outbound

This paper cites On the convergence rates of policy gradient methods.Journal of Machine Learning Research, 23(282):1–36, 2022.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the convergence rates of policy gradient methods.Journal of Machine Learning Research, 23(282):1–36, 2022

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.539115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:53:15.489700Z digest=sha256:fa5630fa3a402e320f0a6f57e0d8310c1d2a892ce0888089039eafd27abfc9f2

Pith citing papers

Observation ada1f7a4-89f3-4e80-98c2-0ba288ee1c01 · inbound

Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum cites this paper.

Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:17:36.525446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:12:57.430291Z digest=sha256:7bcbdf60939fc29892bff99d72319e20ed088505aab28b7ebbb166e0a602ab86