Using the agent's current value estimate as a potential-based shaping function converges in tabular RL and accelerates DQN on Atari.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Bootstrapped Reward Shaping
Using the agent's current value estimate as a potential-based shaping function converges in tabular RL and accelerates DQN on Atari.