Pith. sign in

REVIEW

Efficient Planning under Partial Observability with Unnormalized Q Functions and Spectral Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.05010 v2 pith:OG7FE6P7 submitted 2019-11-12 cs.AI cs.LG

classification cs.AIcs.LG
keywords learningalgorithmplanningproblemstimeapproachclassicaldomains
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning and planning in partially-observable domains is one of the most difficult problems in reinforcement learning. Traditional methods consider these two problems as independent, resulting in a classical two-stage paradigm: first learn the environment dynamics and then plan accordingly. This approach, however, disconnects the two problems and can consequently lead to algorithms that are sample inefficient and time consuming. In this paper, we propose a novel algorithm that combines learning and planning together. Our algorithm is closely related to the spectral learning algorithm for predicitive state representations and offers appealing theoretical guarantees and time complexity. We empirically show on two domains that our approach is more sample and time efficient compared to classical methods.

Discussion (0). Sign in to comment.

Pith tools