Pith. sign in

REVIEW

Multi-Timescale, Gradient Descent, Temporal Difference Learning with Linear Options

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1703.06471 v1 pith:EEQBEI3L submitted 2017-03-19 cs.AI

classification cs.AI
keywords temporaltimeabstractionmodelsalgorithmgeneratedlearninglinear
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deliberating on large or continuous state spaces have been long standing challenges in reinforcement learning. Temporal Abstraction have somewhat made this possible, but efficiently planing using temporal abstraction still remains an issue. Moreover using spatial abstractions to learn policies for various situations at once while using temporal abstraction models is an open problem. We propose here an efficient algorithm which is convergent under linear function approximation while planning using temporally abstract actions. We show how this algorithm can be used along with randomly generated option models over multiple time scales to plan agents which need to act real time. Using these randomly generated option models over multiple time scales are shown to reduce number of decision epochs required to solve the given task, hence effectively reducing the time needed for deliberation.

Discussion (0). Sign in to comment.

Pith tools