REVIEW 1 cited by
Near-optimal Optimistic Reinforcement Learning using Empirical Bernstein Inequalities
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We study model-based reinforcement learning in an unknown finite communicating Markov decision process. We propose a simple algorithm that leverages a variance based confidence interval. We show that the proposed algorithm, UCRL-V, achieves the optimal regret $\tilde{\mathcal{O}}(\sqrt{DSAT})$ up to logarithmic factors, and so our work closes a gap with the lower bound without additional assumptions on the MDP. We perform experiments in a variety of environments that validates the theoretical bounds as well as prove UCRL-V to be better than the state-of-the-art algorithms.
Forward citations
Cited by 1 Pith paper
-
Asymptotically optimal regret in communicating Markov decision processes
The paper claims the first asymptotically optimal regret algorithm, achieving the exact logarithmic constant K(M), for average-reward communicating Markov decision processes.
Discussion (0). Continue with ORCID to comment.