Pith. sign in

REVIEW

Empirical Q-Value Iteration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1412.0180 v3 pith:B67GQFU3 submitted 2014-11-30 math.OC cs.LG

classification math.OCcs.LG
keywords algorithmq-valuealgorithmsapproximation-basedconvergenceempiricalfunctioniteration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a new simple and natural algorithm for learning the optimal Q-value function of a discounted-cost Markov Decision Process (MDP) when the transition kernels are unknown. Unlike the classical learning algorithms for MDPs, such as Q-learning and actor-critic algorithms, this algorithm doesn't depend on a stochastic approximation-based method. We show that our algorithm, which we call the empirical Q-value iteration (EQVI) algorithm, converges to the optimal Q-value function. We also give a rate of convergence or a non-asymptotic sample complexity bound, and also show that an asynchronous (or online) version of the algorithm will also work. Preliminary experimental results suggest a faster rate of convergence to a ball park estimate for our algorithm compared to stochastic approximation-based algorithms.

Discussion (0). Continue with ORCID to comment.

Pith tools