Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:01:47.191696Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.09523.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:01:47.191696Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 82aaab77-0338-46e0-bb06-92e787f30d06 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2eb5b45-b5dc-4f01-b32a-fa21b259404f · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49aac7d9-f0e4-4fd7-be52-3ff2dcc6abd5 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Sur les op \'e rations dans les ensembles abstraits et leur application aux \'e quations int \'e grales
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88c31345-0586-4a52-b1f5-a7038a19a76d · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Dynamic Programming
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7450fe15-d8bf-428c-9bda-665639b9468a · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Bertsekas and John N
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fbe8b06b-4a0d-4413-a3e5-73037c103b4f · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Multistep Credit Assignment in Deep Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49119f6f-aff3-4fbb-8560-2f69162cbcff · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values ChainerRL : A deep reinforcement learning library
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7dfa3418-b4fe-4d31-83b9-87ac0e446569 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d742ebfd-07fc-4337-90c9-25fd85e9f059 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Kingma and Jimmy Ba
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58a7a71b-afa8-4ac7-82b8-9c0fbb5dfb7a · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Orr, and Klaus-Robert M \"u ller
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a3e49df-698b-4a37-8b30-74688ca8b8b5 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Rusu, Joel Veness, Marc G
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9259be3-76b0-4b10-a6a0-c393a0c51e4d · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Towards model-free RL algorithms that scale well with unstructured data
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03ae75ac-1be6-466a-bae3-d767da002530 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Revisiting Rainbow : Promoting more insightful and inclusive deep reinforcement learning research
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b6d1f872-1467-4ef7-83ea-07d58b4a3290 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Rummery and Mahesan Niranjan
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49ad9446-5ee9-4e3d-b989-db0a0616c173 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8cff4409-b4e7-4c88-bd2e-fb44c043a8b3 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Littman, and Csaba Szepesv \'a ri
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58931b2f-9137-4d8e-a2c7-a1c07d8fb3ce · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90ab8c3d-19a9-4e2f-8f9d-8f470a925902 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Sutton and Andrew G
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 380e6ba9-f74a-4281-accb-a82f3045e7c2 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values VA-learning as a more efficient alternative to Q-learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41bad30e-a29f-44fd-993e-b5942ea11e2f · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Double Q-learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2413f22d-5b75-47ec-b6f8-1940c20fd601 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Insights in Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 344721d1-d006-4cd5-be36-51e75d1f6521 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Dueling network architectures for deep reinforcement learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3bc160fe-f2bf-477d-8e61-9713e3fe5b9d · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15cdb5f2-f4e9-43b5-91fa-fc55735cdd2f · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a390f672-f32c-42d0-aa0d-86a298d88ebc · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Wiering and Hado van Hasselt
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d490c7c4-2992-4a36-8d87-4706443e9f22 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values Wiering and Hado van Hasselt
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85b62dba-ea0f-4746-a7cd-3f2f496889b6 · outbound
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.