A proposed Q-learning port of TD(Delta) decomposes action values by discount factor, but the core Bellman equation for the delta components is derived incorrectly and the claimed experiments are missing.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition
A proposed Q-learning port of TD(Delta) decomposes action values by discount factor, but the core Bellman equation for the delta components is derived incorrectly and the claimed experiments are missing.