Brownian-bridge self-normalized confidence regions for constant-stepsize TD learning under Markovian sampling, centered on the Richardson-Romberg stationary target or, in a horizon-indexed regime, the projected Bellman solution.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Self-Normalized Inference for Constant-Stepsize Temporal-Difference Learning under Markovian Sampling
Brownian-bridge self-normalized confidence regions for constant-stepsize TD learning under Markovian sampling, centered on the Richardson-Romberg stationary target or, in a horizon-indexed regime, the projected Bellman solution.