Pith. sign in

REVIEW 1 cited by

Analysis of Off-Policy Multi-Step TD-Learning with Linear Function Approximation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15781 v2 pith:E6XXOR2P submitted 2024-02-24 eess.SY cs.LGcs.SY

classification eess.SYcs.LGcs.SY
keywords algorithmstd-learningcounterpartslearninganalysisapproximationcontrolconverge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper analyzes multi-step TD-learning algorithms within the `deadly triad' scenario, characterized by linear function approximation, off-policy learning, and bootstrapping. In particular, we prove that n-step TD-learning algorithms converge to a solution as the sampling horizon n increases sufficiently. The paper is divided into two parts. In the first part, we comprehensively examine the fundamental properties of their model-based deterministic counterparts, including projected value iteration, gradient descent algorithms, and the control theoretic approach, which can be viewed as prototype deterministic algorithms whose analysis plays a pivotal role in understanding and developing their model-free reinforcement learning counterparts. In particular, we prove that these algorithms converge to meaningful solutions when n is sufficiently large. Based on these findings, two n-step TD-learning algorithms are proposed and analyzed, which can be seen as the model-free reinforcement learning counterparts of the gradient and control theoretic algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges

    cs.AI 2024-11 conditional novelty 2.0 of 10

    A comprehensive but flawed survey of RL algorithms that catalogs many methods and applications without rigorous comparative analysis.

Pith tools