Pith. sign in

Paper Citation Record · LEDGER

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead

As of 8 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2506.07054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07054 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:53:15.489700Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T08:12:57.430291Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T08:17:36.523927Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56c0e748-0fab-45bb-bd14-b08996d7944f · outbound

This paper cites On the theory of policy gradient methods: Optimality, approximation, and distribution shift.The Journal of Machine Learning Research, 22(1):4431–4506, 2021.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the theory of policy gradient methods: Optimality, approximation, and distribution shift.The Journal of Machine Learning Research, 22(1):4431–4506, 2021

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:15.402520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:53:15.402520Z digest=sha256:c688634f69d20c7bb9c731afbfb045ee86fdbddcdd1003c300029d3b8eab21d6

Observation b3422e75-af8d-4b42-81e8-c13b283b4395 · outbound

This paper cites Natural actor- critic algorithms.Automatica, 45:2471–2482, 11 2009.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Natural actor- critic algorithms.Automatica, 45:2471–2482, 11 2009

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.739448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.407155Z digest=sha256:22186b7af112c962679b69f26579a4a8100455c509fb4f943b50440ef11f6160

Observation 2a85ad48-e844-4c5d-932c-ce636f67b2a4 · outbound

This paper cites Finite-time analysis of single-timescale actor-critic.Advances in Neural Information Processing Systems, 36, 2024.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Finite-time analysis of single-timescale actor-critic.Advances in Neural Information Processing Systems, 36, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.728469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.411205Z digest=sha256:6701a287b0dd07d7809e4fc282f383cacef1d7e568d06b5a004f0699c1dd8f74

Observation 43763294-3e93-43a4-b32b-352057071b5e · outbound

This paper cites Bellemare.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Bellemare

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.717309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.415171Z digest=sha256:254496d5f5542e92f4ca06d2d4a22d002af782fd89193e9a13f7598c0284eaa1

Observation 81b3254c-48b1-4fee-9ddf-acd97723552c · outbound

This paper cites Beyond the One Step Greedy Approach in Reinforcement Learning.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Beyond the One Step Greedy Approach in Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:53:15.526463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.419078Z digest=sha256:888b880262268d102f9a2b150733f4e517fa7f23e74ee37d0e5d53b790e96a84

Observation 51ece1e1-2ea5-4ff7-9766-b6dc50795aea · outbound

This paper cites Kakade, Karan Singh, and Abby Van Soest.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Kakade, Karan Singh, and Abby Van Soest

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.706317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.423850Z digest=sha256:7d6932785853417b88deaff7512948bd37d77850f4d606b4206491940ebd8da1

Observation 57b80b14-6b61-46d9-83f0-f94a58620f0d · outbound

This paper cites Actor-critic algorithms.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Actor-critic algorithms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.695828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.428386Z digest=sha256:9a4ed2cd8ce4b7c9e50f13a9d66d9f97d2347c16d6b3b2d2af8626905abb7527

Observation af50b33e-cec3-45cf-a19d-9e94afbfea7e · outbound

This paper cites Policy gradient with tree search (PGTS) in reinforcement learning evades local maxima.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy gradient with tree search (PGTS) in reinforcement learning evades local maxima

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.684281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.435726Z digest=sha256:4fd095014ad5bc2cd8c6b773794b9049aa9a24dca830f713eb5a18e8bc1ce683

Observation b7909ac8-de03-407f-80da-d6f6a105ef17 · outbound

This paper cites an unresolved cited work.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:53:15.672487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.439433Z digest=sha256:66a4162e9822e5c6c8378ba144fefb31f2773c97fb0321ccbe9a5aff6ff12209

Observation a08650c2-ebf2-4ecf-9276-7e15532c7613 · outbound

This paper cites Elementary analysis of policy gradient methods, 2024.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Elementary analysis of policy gradient methods, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.662105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.442956Z digest=sha256:07f43c97c84fa143d8b5436c72367306d2536a7c1f4303fb076b63b3962c9ba8

Observation bc90bb24-e7bf-416f-8c44-9edb6cbf28f3 · outbound

This paper cites On the global conver- gence rates of softmax policy gradient methods.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the global conver- gence rates of softmax policy gradient methods

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.651383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.446740Z digest=sha256:b7471a50ee273a395afd985284641fb0f68460a108457f4dbb321b3316d88689

Observation 1e918b11-ad93-439d-8d9c-730f3fba57df · outbound

This paper cites Rusu, Joel Veness, Marc G.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Rusu, Joel Veness, Marc G

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.640881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.450997Z digest=sha256:dd81b2aca4bf9ea637174938d2d537ad05884aca014d3c6ec38ece3d11ad3844

Observation 7eb16917-a3a7-4ac4-8ac6-0ec7260376bf · outbound

This paper cites Policy mirror descent with lookahead, 2024.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy mirror descent with lookahead, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.629318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.454437Z digest=sha256:b26a5626b3a8759e57c2b7fb85dce01d4f3db076130b97c229386cdfd1e2eb67

Observation a2d3ff54-86d7-456c-9820-2ef2c66ea579 · outbound

This paper cites John Wiley & Sons, 2014.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead John Wiley & Sons, 2014

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:15.457879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:53:15.457879Z digest=sha256:2eb6b475afe62d6e30474ab3d929f625630d6d728104e074c56612df6c706fc0

Observation ec476048-53ed-420f-823f-3c0793c92f18 · outbound

This paper cites Trust region policy optimization.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Trust region policy optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:15.461456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:53:15.461456Z digest=sha256:a6b2a421c7cf9d7104ddf3a16887a5aa2a74eed76a640dc442bd134ac2de9e8f

Observation 34db7301-cfab-47c5-b2da-56ab6f3284c1 · outbound

This paper cites Jordan, and Pieter Abbeel.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Jordan, and Pieter Abbeel

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.604094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.464987Z digest=sha256:6462e6683fde2e7dcc9edb36cacd41ef96f2c630ed27e0c5a69f25c8e9038b1f

Observation 68dbb265-dee3-47f6-984a-c97c6eda7e5c · outbound

This paper cites A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362(6419):1140–1144, 2018.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362(6419):1140–1144, 2018

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.593271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.469607Z digest=sha256:f6df4450e06cbb9fab5e66bab311fcaf1767fdcea37df477fecf1d787ff47f20

Observation 3dcbdbcf-1631-4f47-aaa6-d304346c9057 · outbound

This paper cites Sutton and Andrew G.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Sutton and Andrew G

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:15.473939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:53:15.473939Z digest=sha256:06984d3614ed76410292899e47261a343d9d70809c79a6a9f7131b02937200c2

Observation c23fe72e-c3e4-43d5-a58f-114fe4743eb8 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy gradient methods for reinforcement learning with function approximation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.572872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.478052Z digest=sha256:0397602cbd15e9a8755c80849a040f13f0bcc5c929c4d96d3e562d893d8298b3

Observation 95b5e946-8bc8-42ee-9e6e-91def601df2d · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead Policy gradient methods for reinforcement learning with function approximation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.561490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.482171Z digest=sha256:61059e85fac0560030ac5424a5bce81c26c7a126c8172dfbcdf8a2ef9bd64907

Observation a92f577c-8c28-4350-b9fc-696876e7c04e · outbound

This paper cites On the convergence rates of policy gradient methods, 2022.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the convergence rates of policy gradient methods, 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.549983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.485818Z digest=sha256:ce5723fd8169c1718e5f8a5e77668f9549222bf0735963826db4b3ddb65e9b16

Observation 86e9b9f1-d77b-4dec-a355-26195e177e7e · outbound

This paper cites On the convergence rates of policy gradient methods.Journal of Machine Learning Research, 23(282):1–36, 2022.

Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead On the convergence rates of policy gradient methods.Journal of Machine Learning Research, 23(282):1–36, 2022

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:53:15.539115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:53:15.489700Z digest=sha256:53ba3044bcc47530ff5823ae8c01ecd25a7affceeb775a89ac0cd7eef499f235

Pith citing papers

Observation ada1f7a4-89f3-4e80-98c2-0ba288ee1c01 · inbound

Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum cites this paper.

Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:17:36.525446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T08:12:57.430291Z digest=sha256:685ce388f5dadeb274347541f225d94b9e48e0c20ccd17afdb034e73f450fdd9