Averaged-dqn: Variance reduction and stabiliza- tion for deep reinforcement learning

[Anschel et al · 2017

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

browse 1 citing papers

representative citing papers

In Hindsight: A Smooth Reward for Steady Exploration

cs.LG · 2019-06-24 · unverdicted · novelty 4.0

Adding a hindsight factor that integrates historic temporal differences into the Q-learning loss reduces overestimation and yields higher average scores than DQN, DDQN and dueling networks on ATARI games after 10 million frames.

citing papers explorer

Showing 1 of 1 citing paper.

In Hindsight: A Smooth Reward for Steady Exploration cs.LG · 2019-06-24 · unverdicted · none · ref 2
Adding a hindsight factor that integrates historic temporal differences into the Q-learning loss reduces overestimation and yields higher average scores than DQN, DDQN and dueling networks on ATARI games after 10 million frames.

Averaged-dqn: Variance reduction and stabiliza- tion for deep reinforcement learning

fields

years

verdicts

representative citing papers

citing papers explorer