A gated mixture of Q-functions with different discount factors, trained with undiscounted Bellman error, adapts its temporal horizon in small MiniGrid tasks, but its theoretical justification is circular and baseline comparisons are missing.
Title resolution pending
1 Pith paper cite this work, alongside 539 external citations. Polarity classification is still indexing.
1
Pith paper citing it
539
external citations · OpenAlex
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Adaptive Multi-Horizon Reinforcement Learning
A gated mixture of Q-functions with different discount factors, trained with undiscounted Bellman error, adapts its temporal horizon in small MiniGrid tasks, but its theoretical justification is circular and baseline comparisons are missing.