Pith. sign in

REVIEW

A Model-free Learning Algorithm for Infinite-horizon Average-reward MDPs with Near-optimal Regret

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.04354 v2 pith:WYIIWNO4 submitted 2020-06-08 cs.LG stat.ML

classification cs.LGstat.ML
keywords algorithmmodel-freelearningmdpsregretachievesapproximationaverage-reward
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Recently, model-free reinforcement learning has attracted research attention due to its simplicity, memory and computation efficiency, and the flexibility to combine with function approximation. In this paper, we propose Exploration Enhanced Q-learning (EE-QL), a model-free algorithm for infinite-horizon average-reward Markov Decision Processes (MDPs) that achieves regret bound of $O(\sqrt{T})$ for the general class of weakly communicating MDPs, where $T$ is the number of interactions. EE-QL assumes that an online concentrating approximation of the optimal average reward is available. This is the first model-free learning algorithm that achieves $O(\sqrt T)$ regret without the ergodic assumption, and matches the lower bound in terms of $T$ except for logarithmic factors. Experiments show that the proposed algorithm performs as well as the best known model-based algorithms.

Discussion (0). Continue with ORCID to comment.

Pith tools