Two memory-saving variants of LSVI-UCB for linear MDPs are proposed; the fixed-reset variant has a sublinear space-regret trade-off proof, while the adaptive variant lacks a regret guarantee despite the abstract claiming sublinear regret.
Logarithmic online regret bounds for undiscounted reinforcement learning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs
Two memory-saving variants of LSVI-UCB for linear MDPs are proposed; the fixed-reset variant has a sublinear space-regret trade-off proof, while the adaptive variant lacks a regret guarantee despite the abstract claiming sublinear regret.