Quantile fixed-point estimators for distributional policy evaluation achieve the parametric √n rate, attain the semiparametric efficiency bound for fixed m, remain efficient as m→∞, and admit Berry–Esseen inference.
A finite sample analysis of distributional td learning with linear function approximation
3 Pith papers cite this work. Polarity classification is still indexing.
abstract
In this paper, we study the finite-sample statistical rates of distributional temporal difference (TD) learning with linear function approximation. The aim of distributional TD learning is to estimate the return distribution of a discounted Markov decision process for a given policy {\pi}. Previous works on statistical analysis of distributional TD learning mainly focus on the tabular case. In contrast, we first consider the linear function approximation setting and derive sharp finite-sample rates. Our theoretical results demonstrate that the sample complexity of linear distributional TD learning matches that of classic linear TD learning. This implies that, with linear function approximation, learning the full distribution of the return from streaming data is no more difficult than learning its expectation (value function). To derive tight sample complexity bounds, we conduct a fine-grained analysis of the linear-categorical Bellman equation and employ the exponential stability arguments for products of random matrices. Our results provide new insights into the statistical efficiency of distributional reinforcement learning algorithms.
citation-role summary
citation-polarity summary
years
2026 3roles
background 2polarities
background 2representative citing papers
Develops quotient-categorical representations that render the average-reward distributional Bellman operator well-defined, non-expansive, and convergent under i.i.d. and Markovian sampling.
citing papers explorer
-
Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning
Quantile fixed-point estimators for distributional policy evaluation achieve the parametric √n rate, attain the semiparametric efficiency bound for fixed m, remain efficient as m→∞, and admit Berry–Esseen inference.
-
Quotient-Categorical Representations for Bellman-Compatible Average-Reward Distributional Reinforcement Learning
Develops quotient-categorical representations that render the average-reward distributional Bellman operator well-defined, non-expansive, and convergent under i.i.d. and Markovian sampling.
- A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning