Pith. sign in

REVIEW 2 cited by

Efficient Model-free Reinforcement Learning in Metric Spaces

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.00475 v1 pith:GAP4K3PM submitted 2019-05-01 cs.LG stat.ML

classification cs.LGstat.ML
keywords algorithmsefficientmodel-freeq-learningmetricalgorithmlearningmdps
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model-free Reinforcement Learning (RL) algorithms such as Q-learning [Watkins, Dayan 92] have been widely used in practice and can achieve human level performance in applications such as video games [Mnih et al. 15]. Recently, equipped with the idea of optimism in the face of uncertainty, Q-learning algorithms [Jin, Allen-Zhu, Bubeck, Jordan 18] can be proven to be sample efficient for discrete tabular Markov Decision Processes (MDPs) which have finite number of states and actions. In this work, we present an efficient model-free Q-learning based algorithm in MDPs with a natural metric on the state-action space--hence extending efficient model-free Q-learning algorithms to continuous state-action space. Compared to previous model-based RL algorithms for metric spaces [Kakade, Kearns, Langford 03], our algorithm does not require access to a black-box planning oracle.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Reinforcement Learning Sample-Efficiency using Local Approximation

    cs.LG 2025-07 reject novelty 5.0 of 10

    The paper's claimed O(SA log A) sample-complexity bound for RL in metric state spaces is not supported due to a lower bound being misused as an upper bound in the proof.

  2. Bellman operator convergence enhancements in reinforcement learning algorithms

    cs.LG 2025-05 reject novelty 4.0 of 10

    A new advantage-weighted Bellman operator is claimed to speed up Q-learning convergence, but the proofs are flawed and experiments lack error bars.

Pith tools