Presents an optimistic FTRL algorithm for online episodic tabular MDPs with unknown transitions that attains data-dependent regret bounds including first-order, second-order, path-length, and polylog(T) gap-dependent bounds in the stochastic case.
arXiv preprint arXiv:2504.07307 , year=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions
Presents an optimistic FTRL algorithm for online episodic tabular MDPs with unknown transitions that attains data-dependent regret bounds including first-order, second-order, path-length, and polylog(T) gap-dependent bounds in the stochastic case.