Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:42:59.250552Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2411.14019.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:42:59.250552Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:42:59.250552Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-12T15:42:59.404858Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b5c84053-f80e-4b8f-84c1-453c19f3cc2e · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecdca55a-7bb2-4742-aaef-1a34505bbcd0 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 381d1499-5fe6-4378-a685-f75d8aa4fea9 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Playing Atari with Deep Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24cd254b-6099-4857-9682-86463961ca40 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Dota 2 with Large Scale Deep Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24576388-8192-482d-9d48-c44095a59020 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fb9eaa60-77f6-4d12-9569-17dc64c5a582 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & V an Roy, B
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 99d7a4e1-9531-47c0-9a6c-36dc7f0e74f7 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Williams, R
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 83f9c814-d6f3-4b7a-a427-2955f06f1e62 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aca79d4-1d02-4420-aef9-2238ca940b1d · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Sutton, R
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f87e1e2e-1a36-4a68-acc0-825f438294a9 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Algorithms for reinforcement learning (Springer nature, 2022)
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7a9d7340-920e-4246-9be3-e5a2dc69fb58 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 32a26af1-7437-4fd2-8094-5ba61c633612 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9e936308-dedb-45db-bfe0-127856906239 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cec608e5-bf70-4f60-aa52-e78108fc2ecd · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Silver, D
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b2894724-a0e9-4ba2-9528-9381dd00129d · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Error bounds for approximate policy iteration
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5391d55f-7134-4852-851d-8a29e11141e1 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Prioritized Experience Replay
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76a94055-7c46-4536-8752-8daad7f98cb2 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Pilarski, P
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9b2e9293-a66c-4e0a-a9e8-bf86c7793da1 · outbound
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4e491de7-3065-4406-9cef-a33d81355653 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e042fc0-aab3-41b1-8c54-9faf9ea65bcd · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 15c7bbf3-4bd7-4f00-991f-56d3cfba2e47 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Optiongan: Learning joint reward-policy options using gen erative adversarial inverse reinforcement learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 25bf3420-43d4-42aa-a0ea-de2cbe694ad5 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Discovering hierarchy in reinforcement learnin g with hexq
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0de67471-4fdf-4303-af8b-bdaa098d74ad · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2ee35c95-28e3-4b5c-9f77-d23cdeac66fd · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition & Shimkin, N
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f9f7ba4d-e162-4c3c-94cd-6b9179d2d483 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 08db37ee-5cc7-4356-ac02-a81ee85b08a1 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 107ff731-acc9-4c2d-a8ee-4c29de087b79 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Human-level control through deep reinforcement learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4e8acf22-2e3c-486b-a08d-25bec08e83d0 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition A markovian decision process
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7040e14-bba6-4992-a7ad-7b1cee181c07 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e20b45e4-f732-4a9c-b065-32c7bdef6a0b · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition S., McAllester, D
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5ebf1381-c522-491a-b65c-3f76c5b8b3e9 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ff5edea6-8a09-4782-aeef-b1888b8d85ab · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Asynchronous methods for deep reinforcement learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2de305b-222f-454b-8927-97ea947fc5d7 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Proximal Policy Optimization Algorithms
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 870b958d-5b33-4066-97d1-fdef6dbd1879 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition S., Barto, A
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e860f8dc-57df-468f-a7eb-c6b993051fe4 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96c5fc6b-ea2c-407a-b573-4d9f66b394f2 · outbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 96c5fc6b-ea2c-407a-b573-4d9f66b394f2 · inbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.