Introduces Good Policy Identification (GPI) and BEE-GPI algorithm whose sample complexity for positive instances has log(1/δ) coefficient O(H²/(V*−μ0)²) independent of state and action space sizes.
International Conference on Machine Learning , pages=
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.LG 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
Presents an optimistic FTRL algorithm for online episodic tabular MDPs with unknown transitions that attains data-dependent regret bounds including first-order, second-order, path-length, and polylog(T) gap-dependent bounds in the stochastic case.
citing papers explorer
-
Pure Exploration for a Good Policy in Reinforcement Learning with Bandit Feedback
Introduces Good Policy Identification (GPI) and BEE-GPI algorithm whose sample complexity for positive instances has log(1/δ) coefficient O(H²/(V*−μ0)²) independent of state and action space sizes.
-
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions
Presents an optimistic FTRL algorithm for online episodic tabular MDPs with unknown transitions that attains data-dependent regret bounds including first-order, second-order, path-length, and polylog(T) gap-dependent bounds in the stochastic case.