REVIEW 5 cited by
How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Using deep neural nets as function approximator for reinforcement learning tasks have recently been shown to be very powerful for solving problems approaching real-world complexity. Using these results as a benchmark, we discuss the role that the discount factor may play in the quality of the learning process of a deep Q-network (DQN). When the discount factor progressively increases up to its final value, we empirically show that it is possible to significantly reduce the number of learning steps. When used in conjunction with a varying learning rate, we empirically show that it outperforms original DQN on several experiments. We relate this phenomenon with the instabilities of neural networks when they are used in an approximate Dynamic Programming setting. We also describe the possibility to fall within a local optimum during the learning process, thus connecting our discussion with the exploration/exploitation dilemma.
Forward citations
Cited by 5 Pith papers
-
Graph-Enhanced Policy Optimization in LLM Agent Training
GEPO adds graph-centrality-based intrinsic rewards, dynamic discounts, and two-level advantage shaping to group-based RL, improving LLM agent success on ALFWorld, WebShop, and a private Workbench benchmark.
-
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning
Survey mapping RL techniques onto LLM training and highlighting gaps in value-based, off-policy, and bootstrapping methods.
-
Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed
SL-MGAC combines supervised reward prediction, user-group decomposition, and actor-critic RL to allocate live streams in a feed; offline and online tests report gains, but the reward predictor is partly fed the true r...
-
Segmenting Action-Value Functions Over Time-Scales in SARSA via TD($\Delta$)
SARSA(Delta) decomposes action-value estimates across discount factors, but the claimed gains are not consistently supported by the paper's own tables.
-
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition
A proposed Q-learning port of TD(Delta) decomposes action values by discount factor, but the core Bellman equation for the delta components is derived incorrectly and the claimed experiments are missing.
Discussion (0). Continue with ORCID to comment.