Pith. sign in

Paper Citation Record · LEDGER

How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:1512.02011.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1512.02011 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:42:59.192879Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T07:59:40.143882Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4e491de7-3065-4406-9cef-a33d81355653 · inbound

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition cites this paper.

Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:42:59.192879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:42:59.192879Z digest=sha256:553aeadadc22796ba889b59751f365e10a15d081e73d6ab189a6ad5cd62f46f5

Observation 5b623dc0-70f2-4247-b409-a786b649530b · inbound

Segmenting Action-Value Functions Over Time-Scales in SARSA via TD($\Delta$) cites this paper.

Segmenting Action-Value Functions Over Time-Scales in SARSA via TD($\Delta$) How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:27.318410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:27.318410Z digest=sha256:2347c370fe9324dd02e4032e87c015699716373522e5e4ee22da32cea9be23a8

Observation e87f761e-5451-4632-8dc7-05c2491bfcf2 · inbound

Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed cites this paper.

Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:52:50.975419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:52:50.975419Z digest=sha256:adb5fc850ea00a7341adbddda807eba97a4922eded29e38b7408a5f56aadc2a9

Observation 9e81c2d7-32d5-4e48-9600-fe920c48e535 · inbound

Graph-Enhanced Policy Optimization in LLM Agent Training cites this paper.

Graph-Enhanced Policy Optimization in LLM Agent Training How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:17:44.600398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:17:44.600398Z digest=sha256:b91db320d8a786c1ae5ac5cec5d34f7c646e6ce0a02239b20ea270428747978f

Observation 2b8ecf15-d970-4ed6-975a-8360f4d5f3ae · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-04T07:59:40.145194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:cc7ae11d94b0cc2172a2d6696b54fcdfbc793b65aef20fb6e58330ddc8db3893