Pith. sign in

REVIEW 1 cited by

Effective Multi-step Temporal-Difference Learning for Non-Linear Function Approximation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1608.05151 v1 pith:FYTINHHO submitted 2016-08-18 cs.AI

classification cs.AI
keywords approximationfunctionmulti-stepnon-linearlearningmethodsreasonlambda
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Multi-step temporal-difference (TD) learning, where the update targets contain information from multiple time steps ahead, is one of the most popular forms of TD learning for linear function approximation. The reason is that multi-step methods often yield substantially better performance than their single-step counter-parts, due to a lower bias of the update targets. For non-linear function approximation, however, single-step methods appear to be the norm. Part of the reason could be that on many domains the popular multi-step methods TD($\lambda$) and Sarsa($\lambda$) do not perform well when combined with non-linear function approximation. In particular, they are very susceptible to divergence of value estimates. In this paper, we identify the reason behind this. Furthermore, based on our analysis, we propose a new multi-step TD method for non-linear function approximation that addresses this issue. We confirm the effectiveness of our method using two benchmark tasks with neural networks as function approximation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. To bootstrap or to rollout? An optimal and adaptive interpolation

    cs.LG 2024-11 conditional novelty 7.0 of 10

    Subgraph Bellman operators give a policy evaluation estimator whose finite-sample error nearly matches TD's optimal asymptotic variance while retaining MC's occupancy-adaptive sample complexity.

Pith tools