Pith. sign in

REVIEW 5 cited by

How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1512.02011 v2 pith:HJN5HFXK submitted 2015-12-07 cs.LG cs.AI

classification cs.LGcs.AI
keywords learningdeepdiscountwhendynamicempiricallyfactorneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Using deep neural nets as function approximator for reinforcement learning tasks have recently been shown to be very powerful for solving problems approaching real-world complexity. Using these results as a benchmark, we discuss the role that the discount factor may play in the quality of the learning process of a deep Q-network (DQN). When the discount factor progressively increases up to its final value, we empirically show that it is possible to significantly reduce the number of learning steps. When used in conjunction with a varying learning rate, we empirically show that it outperforms original DQN on several experiments. We relate this phenomenon with the instabilities of neural networks when they are used in an approximate Dynamic Programming setting. We also describe the possibility to fall within a local optimum during the learning process, thus connecting our discussion with the exploration/exploitation dilemma.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Graph-Enhanced Policy Optimization in LLM Agent Training

    cs.AI 2025-10 conditional novelty 6.0 of 10

    GEPO adds graph-centrality-based intrinsic rewards, dynamic discounts, and two-level advantage shaping to group-based RL, improving LLM agent success on ALFWorld, WebShop, and a private Workbench benchmark.

  2. Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    Survey mapping RL techniques onto LLM training and highlighting gaps in value-based, off-policy, and bootstrapping methods.

  3. Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed

    cs.IR 2024-11 reject novelty 5.0 of 10

    SL-MGAC combines supervised reward prediction, user-group decomposition, and actor-critic RL to allocate live streams in a feed; offline and online tests report gains, but the reward predictor is partly fed the true r...

  4. Segmenting Action-Value Functions Over Time-Scales in SARSA via TD($\Delta$)

    cs.LG 2024-11 reject novelty 3.0 of 10

    SARSA(Delta) decomposes action-value estimates across discount factors, but the claimed gains are not consistently supported by the paper's own tables.

  5. Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition

    cs.LG 2024-11 reject novelty 3.0 of 10

    A proposed Q-learning port of TD(Delta) decomposes action values by discount factor, but the core Bellman equation for the delta components is derived incorrectly and the claimed experiments are missing.

Pith tools