Pith. sign in

REVIEW 1 cited by

Temporal Difference Learning with Continuous Time and State in the Stochastic Setting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.07960 v3 pith:LSMGYQCD submitted 2022-02-16 cs.LG cs.AImath.APmath.OC

Temporal Difference Learning with Continuous Time and State in the Stochastic Setting

classification cs.LG cs.AImath.APmath.OC
keywords learningstochasticcontinuous-timedifferentialequationsfunctionlinearmethods
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We consider the problem of continuous-time policy evaluation. This consists in learning through observations the value function associated with an uncontrolled continuous-time stochastic dynamic and a reward function. We propose two original variants of the well-known TD(0) method using vanishing time steps. One is model-free and the other is model-based. For both methods, we prove theoretical convergence rates that we subsequently verify through numerical simulations. Alternatively, those methods can be interpreted as novel reinforcement learning approaches for approximating solutions of linear PDEs (partial differential equations) or linear BSDEs (backward stochastic differential equations).

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Fast and Robust Convergence Rate for TD(0) with Linear Function Approximation, Universal Learning Steps and I.I.D. Samples

    stat.ML 2026-06 unverdicted novelty 7.0

    Establishes an O(1/k) MSE convergence rate for TD(0) with LFA that is independent of the smallest eigenvalue of the uncentered covariance matrix and robust to ill-conditioning.