EEDQN sets its bootstrap horizon using a Q-value difference threshold and then uses the ensemble mean for one-step targets and the ensemble minimum for multi-step targets, reporting the best final score on four of five MinAtar games.
Addressing function approximation error in actor-critic methods
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Ensemble Elastic DQN: A Step Dependent Ensemble Approach for Reducing Overestimation in Deep Value-Based Reinforcement Learning
EEDQN sets its bootstrap horizon using a Q-value difference threshold and then uses the ensemble mean for one-step targets and the ensemble minimum for multi-step targets, reporting the best final score on four of five MinAtar games.