Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:1512.02011.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:42:59.192879Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T07:59:40.143882Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 4e491de7-3065-4406-9cef-a33d81355653 · inbound
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b623dc0-70f2-4247-b409-a786b649530b · inbound
Segmenting Action-Value Functions Over Time-Scales in SARSA via TD($\Delta$) How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e87f761e-5451-4632-8dc7-05c2491bfcf2 · inbound
Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e81c2d7-32d5-4e48-9600-fe920c48e535 · inbound
Graph-Enhanced Policy Optimization in LLM Agent Training How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b8ecf15-d970-4ed6-975a-8360f4d5f3ae · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.