In monitored MDPs, a deep reward model plus Q-learning can generalize to unmonitored states and reach near-optimal behavior, but can also overgeneralize; ensemble-based cautious policies reduce that overgeneralization.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Generalization in Monitored Markov Decision Processes (Mon-MDPs)
In monitored MDPs, a deep reward model plus Q-learning can generalize to unmonitored states and reach near-optimal behavior, but can also overgeneralize; ensemble-based cautious policies reduce that overgeneralization.