Pith. sign in

REVIEW 2 major objections 5 minor 60 references

Quantile return-distribution estimators achieve parametric sample rates and stay efficient even as the number of quantiles grows without bound.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 15:33 UTC pith:55JUSQDW

load-bearing objection Solid first non-asymptotic rates + efficiency + Berry–Esseen for the practical quantile parametrization of distributional RL; generative-model setting and density regularity are the usual limits, not fatal ones. the 2 major comments →

arxiv 2607.08444 v2 pith:55JUSQDW submitted 2026-07-09 stat.ML cs.LG

Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning

classification stat.ML cs.LG MSC 62M0562G2090C4068T05
keywords distributional reinforcement learningquantile representationsemiparametric efficiencyBerry–Esseen boundsWasserstein metricgenerative modelpolicy evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper studies how well one can recover the full distribution of discounted returns for a fixed policy when that distribution is represented by a finite collection of quantiles. With access to a generative model, the authors form an empirical MDP and solve its quantile-projected Bellman equation. For any fixed number of quantiles m they prove a high-probability error of order sqrt(m/n) in the supremum Wasserstein metric, matching the classical parametric rate. They further show that the estimated quantiles are asymptotically normal and attain the semiparametric efficiency bound. When m itself tends to infinity the same estimators remain efficient for the underlying infinite-dimensional problem, and a Berry–Esseen bound supplies a quantitative rate of Gaussian approximation that justifies confidence intervals for smooth functionals of the return distribution.

Core claim

The certainty-equivalence estimator obtained by solving the empirical quantile-projected Bellman equation is sample-efficient for fixed m, asymptotically normal and semiparametrically efficient, and continues to attain the nonparametric efficiency bound of the full return distribution in the large-m limit.

What carries the argument

The quantile fixed-point map η=Π_m T^π η together with its empirical counterpart; the argument is carried by rewriting the fixed-point equation as a finite-dimensional Z-estimator whose Jacobian G_m is invertible under a non-degeneracy condition, then embedding the growing-m matrices into operators on an L^{2} quantile space.

Load-bearing premise

No Bellman-shifted pair of quantile locations may land exactly on the endpoints of the reward support; otherwise the Jacobian of the projected Bellman map can become singular and the normality and efficiency claims fail.

What would settle it

In a simple two-state MDP with continuous reward densities, compute the empirical quantile fixed points for increasing n at fixed m and check whether the scaled error sqrt(n) times the Wasserstein distance concentrates around a constant of order sqrt(m) and whether the plug-in confidence intervals for a smooth functional attain the claimed coverage; also verify that the estimated asymptotic variance approaches the nonparametric efficiency bound as m grows.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper studies statistical efficiency of quantile-based distributional policy evaluation under a generative model. For fixed quantile level m it constructs the certainty-equivalence fixed point η_m^{(n)} of the empirical projected Bellman operator and proves a high-probability W̄_∞ bound of order Õ(√(m/n)) (Theorem 3.1), asymptotic normality of √n(θ_m^{(n)}−θ_m) with covariance equal to the semiparametric efficiency bound (Theorem 3.3), and a Berry–Esseen rate for smooth functionals (Theorem 3.5). As m→∞ the asymptotic variance of those functionals converges to the nonparametric efficiency bound of the infinite-dimensional return distribution (Theorem 3.4). The analysis relies on a Z-estimation formulation of the projected Bellman equation, operator embeddings between finite-dimensional quantile vectors and L^{2} spaces, and a stochastic expansion for the nonlinear remainder. Numerical experiments on a two-state MDP corroborate the rates, normality and coverage of the proposed confidence intervals.

Significance. If the results hold, the paper supplies the first non-asymptotic sample-complexity guarantee, efficiency theory and Berry–Esseen bound for quantile distributional RL, and shows that finite-dimensional quantile parametrizations preserve asymptotic efficiency in the infinite-dimensional limit. These contributions fill a clear gap relative to existing categorical analyses and nonparametric return-distribution estimators, and they justify the widespread practical use of quantile methods from a statistical standpoint. The proofs are fully written out in the appendices, the operator-theoretic limit argument is novel for this setting, and the numerical section provides clean validation of the main rates and inferential procedures.

major comments (2)
  1. Assumption 2 (θ_m(s,i)−γθ_m(s′,j)∉{0,1} for all pairs) is used to guarantee continuous differentiability of the projected Bellman map and invertibility of G_m (Lemma 4.3). While the set of violating configurations has measure zero under the standing density assumptions (Prop. 2.2 and Assumps. 1, 3/4), the manuscript never states this genericity claim explicitly. A short remark after Assumption 2 clarifying that the results hold on an open dense set of MDPs would remove any residual concern about the scope of Theorems 3.3–3.5.
  2. The Berry–Esseen rate in Theorem 3.5 is Õ(m^{13/4}n^{-1/4}). The n^{-1/4} exponent is correctly attributed to non-smoothness of the quantile projection, yet the polynomial dependence on m is quite high. The paper already flags this as a future direction (Section 6); it would strengthen the claim if the authors indicated whether the m-power can be improved by a more refined empirical-process argument or whether it is essentially sharp under the present expansion.
minor comments (5)
  1. In the abstract and Theorem 3.1 the rate is written Õ(√(m/n)); the same notation appears with different tilde conventions in the introduction. Standardize the asymptotic notation throughout.
  2. Figure 1 captions refer to “W distance” without specifying W_∞; adding the metric would improve readability.
  3. Several long displayed equations in Sections 3.4 and 4.3 (definitions of K and Σ̃) would benefit from a short verbal description of each operator immediately after the formula.
  4. Typographical inconsistencies appear in the arXiv version (e.g., “ηpnq m”, missing spaces around operators). A careful copy-edit pass is needed before final publication.
  5. The computational-cost paragraph after the definition of the quantile fixed-point estimator (Section 3.1) is useful but could be moved to a short remark so that it does not interrupt the statistical narrative.

Circularity Check

0 steps flagged

No significant circularity: quantile fixed-point analysis, Z-estimation, operator limits and efficiency bounds are derived from the generative-model MDP and explicit regularity assumptions; self-citations to the authors’ prior categorical/nonparametric DRL papers are comparative or foundational, not load-bearing reductions.

full rationale

The paper’s central claims (Thm 3.1 non-asymptotic W̄_∞ rate, Thm 3.3 asymptotic normality + semiparametric efficiency of θ_m^{(n)}, Thm 3.4 operator limit matching the nonparametric efficiency bound, Thm 3.5 Berry–Esseen) are obtained by (i) concentration of empirical quantiles of the projected Bellman operator (Bernstein + DKW + contraction), (ii) reformulation of the projected Bellman equation as a finite-dimensional Z-estimating equation whose Jacobian G_m is shown invertible by diagonal dominance under the standing density assumptions, (iii) pathwise differentiability of the fixed-point map to identify the efficient influence function, and (iv) an averaging/embedding operator framework that establishes strong operator convergence of the discrete matrices to the infinite-dimensional quantile Bellman and covariance operators. All steps are self-contained once Assumptions 1–5 are granted; no free parameters are fitted to data and then re-predicted, no uniqueness theorem is imported solely to forbid alternatives, and no ansatz is smuggled via citation. The only self-citations (Peng et al. on categorical DRL, Zhang et al. 2025 on nonparametric distributional RL) supply comparison results or the already-established nonparametric asymptotic variance expression that the present paper then matches via its own operator correspondence and re-derives the efficiency lower bound (Lemma 4.9). These citations are therefore ordinary foundational references, not circular reductions of the quantile claims. Numerical experiments are pure Monte-Carlo validation of the derived rates and coverage, not fitting. Hence the derivation chain does not collapse to its inputs by construction.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 2 invented entities

The paper is pure theory under a generative-model tabular MDP. The load-bearing ingredients are five regularity assumptions on reward densities and quantile locations, plus standard measure-theoretic and empirical-process tools. No free parameters are fitted; the invented objects are the usual projected operators and the averaging/embedding maps needed for the infinite-dimensional limit.

axioms (5)
  • domain assumption Reward distributions admit continuous densities bounded above (Assump. 1) and either bounded away from zero (Assump. 3) or vanishing at endpoints with monotonicity near boundaries (Assump. 4).
    Needed for positive density of the return at interior quantiles, for concentration of empirical quantiles, and for invertibility of the Jacobian G_m.
  • ad hoc to paper No Bellman-shifted quantile pair lands exactly on the reward-support boundary: θ_m(s,i)−γ θ_m(s′,j) ∉ {0,1} (Assump. 2).
    Technical non-degeneracy ensuring continuous differentiability of the projected Bellman map at the true parameter; without it the Z-estimation argument fails.
  • domain assumption Quantile functions of the true return satisfy mild integrability and modulus-of-continuity conditions near 0 and 1 (Assump. 5).
    Controls boundary singularities of the limiting operators K and Σ̃ when m→∞.
  • domain assumption Access to a generative model that returns independent (reward, next-state) pairs for every state-action pair.
    Standard offline generative-model setting; the empirical MDP is formed by n i.i.d. samples per (s,a).
  • standard math Standard empirical-process and concentration tools (Bernstein, DKW, Donsker classes, Berry–Esseen for nonlinear statistics).
    Used throughout the non-asymptotic and asymptotic arguments; cited from classical references.
invented entities (2)
  • Averaging and embedding operators R_m, E_m between L^{2}[0,1]^S and R^{S×m} no independent evidence
    purpose: To lift the finite-dimensional matrices G_m, Σ_m to operators on the infinite-dimensional Hilbert space and prove operator-norm convergence as m→∞.
    Technical device introduced in §4.3; no independent physical meaning, but necessary for the efficiency-preservation theorem.
  • Quantile Bellman operator K and covariance operator Σ̃ on ⊕_s L^{2}[0,1] no independent evidence
    purpose: Infinite-dimensional limits of the discrete Jacobians and covariances; used to identify the limiting asymptotic variance with the nonparametric efficiency bound.
    Defined in §3.4; shown to be bounded and to recover the distributional Bellman operator after a change of coordinates.

pith-pipeline@v1.1.0-grok45 · 52155 in / 3065 out tokens · 36190 ms · 2026-07-14T15:33:13.333528+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning." pith.science (2026). https://pith.science/paper/55JUSQDW

@misc{pith2026260708444,
  author       = {Pith},
  title        = {Pith review of: Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/55JUSQDW}},
  note         = {Machine review of arXiv:2607.08444}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency. We focus on distributional policy evaluation, whose goal is to characterize the return distribution, namely the distribution of discounted cumulative rewards under a given policy. To obtain a finite-dimensional representation of the return distribution, we consider the quantile fixed point $\eta_m$ induced by the quantile-projected distributional Bellman equation. Assuming access to a generative model, we construct an estimator $\eta_m^{(n)}$ based on an empirical Markov decision process. For a fixed number of quantiles $m$, we establish a non-asymptotic error bound for $\eta_m^{(n)}$ and $\eta_m$ under the supremum $W_\infty$ metric, showing that the estimation error scales as $\widetilde{O}(\sqrt{m/n})$ with respect to $m$ and $n$. This implies that the quantile-based distributional policy evaluation problem can be solved with sample efficiency, achieving the optimal parametric $\sqrt{n}$ convergence rate. We derive the asymptotic distribution of the quantile parameters $\sqrt{n}(\theta_m^{(n)}-\theta_m)$ and characterize the semiparametric efficiency bound, which is attained by our estimator. Beyond the fixed-dimensional setting, we investigate the asymptotic regime in which the number of quantiles diverges. We characterize the limit covariance structure and show that it matches the semiparametric efficiency bound of the nonparametric model for distributional policy evaluation, showing that quantile-based estimators remain asymptotically efficient in the infinite-dimensional limit. Finally, we establish a Berry--Esseen theorem for smooth functionals $\sqrt{n}(\eta_m^{(n)}(s)-\eta_m(s))f$, thereby providing a foundation for statistically valid inference on functionals of the quantile-projected return distribution.

Figures

Figures reproduced from arXiv: 2607.08444 by Yang Peng, Zhihua Zhang, Zijie Cheng.

Figure 1
Figure 1. Figure 1: The statistical error W¯ 8pη pnq m , ηmq with different sample sizes. From left to right: m “ 20, m “ 50, m “ 100. m slope intercept R2 20 -0.5022 0.5881 0.9999 50 -0.5095 0.6457 0.9998 100 -0.5114 0.6447 1.0000 [PITH_FULL_IMAGE:figures/full_fig_p034_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The QQ plot of standardized estimator ? npη pnq m ps1q ´ ηmps1qqf{σm,f,s1 for sample size n “ 10000. From left to right: m “ 20, m “ 50, m “ 100. m skewness excess kurtosis Shapiro–Wilk p-value 20 0.0347 -0.2063 0.2491 50 0.0485 -0.2085 0.4592 100 -0.1411 -0.0679 0.2873 [PITH_FULL_IMAGE:figures/full_fig_p035_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

60 extracted references · 7 linked inside Pith

  1. [1]

    R. R. Bahadur. A note on quantiles in large samples.The Annals of Mathematical Statistics, 37(3):577–580, 1966

  2. [2]

    M. G. Bellemare, W. Dabney, and R. Munos. A distributional perspective on reinforcement learning. InInternational conference on machine learning, pages 449–458. PMLR, 2017

  3. [3]

    M. G. Bellemare, S. Candido, P. S. Castro, J. Gong, M. C. Machado, S. Moitra, S. S. Ponda, 36 and Z. Wang. Autonomous navigation of stratospheric balloons using reinforcement learning. Nature, 588(7836):77–82, 2020

  4. [4]

    M. G. Bellemare, W. Dabney, and M. Rowland.Distributional Reinforcement Learning. MIT Press, 2023.http://www.distributional-rl.org

  5. [5]

    Bodnar, A

    C. Bodnar, A. Li, K. Hausman, P. Pastor, and M. Kalakrishnan. Quantile qt-opt for risk-aware vision-based robotic grasping.arXiv preprint arXiv:1910.02787, 2019

  6. [6]

    Boucheron, G

    S. Boucheron, G. Lugosi, and P. Massart.Concentration Inequalities: A Nonasymptotic Theory of Independence. Oxford University Press, 02 2013

  7. [7]

    Brézis.Functional analysis, Sobolev spaces and partial differential equations

    H. Brézis.Functional analysis, Sobolev spaces and partial differential equations. Springer, 2011

  8. [8]

    Chandak, S

    Y. Chandak, S. Niekum, B. da Silva, E. Learned-Miller, E. Brunskill, and P. S. Thomas. Universal off-policy evaluation.Advances in Neural Information Processing Systems, 34:27475–27490, 2021

  9. [9]

    L. Chen, G. Keilbar, and W. B. Wu. Smoothed sgd for quantiles: Bahadur representation and gaussian approximation, 2025

  10. [10]

    L. H. Chen and Q.-M. Shao. Normal approximation for nonlinear statistics using a concentration inequality approach.Bernoulli, 13(2), 2007

  11. [11]

    Dabney, G

    W. Dabney, G. Ostrovski, D. Silver, and R. Munos. Implicit quantile networks for distributional reinforcement learning. InInternational conference on machine learning, pages 1096–1105. PMLR, 2018

  12. [12]

    Dabney, M

    W. Dabney, M. Rowland, M. Bellemare, and R. Munos. Distributional reinforcement learning with quantile regression. InProceedings of the AAAI Conference on Artificial Intelligence, 2018

  13. [13]

    T. Doan, B. Mazoure, and C. Lyle. Gan q-learning.arXiv preprint arXiv:1805.04874, 2018

  14. [14]

    Fawzi, M

    A. Fawzi, M. Balog, A. Huang, T. Hubert, B. Romera-Paredes, M. Barekatain, A. Novikov, F. J. R Ruiz, J. Schrittwieser, G. Swirszcz, et al. Discovering faster matrix multiplication algorithms with reinforcement learning.Nature, 610(7930):47–53, 2022

  15. [15]

    Freirich, T

    D. Freirich, T. Shimkin, R. Meir, and A. Tamar. Distributional multivariate policy evaluation and exploration with the bellman gan. InInternational Conference on Machine Learning, pages 1983–1992. PMLR, 2019. 37

  16. [16]

    Ghysels, P

    E. Ghysels, P. Santa-Clara, and R. Valkanov. There is a risk-return trade-off after all.Journal of financial economics, 76(3):509–548, 2005

  17. [17]

    B. Hao, X. Ji, Y. Duan, H. Lu, C. Szepesvari, and M. Wang. Bootstrapping fitted q-evaluation for off-policy inference. InInternational Conference on Machine Learning, pages 4074–4084. PMLR, 2021

  18. [18]

    Y. Hua, R. Li, Z. Zhao, X. Chen, and H. Zhang. Gan-powered deep distributional reinforcement learning for resource management in network slicing.IEEE Journal on Selected Areas in Communications, 38(2):334–349, 2019

  19. [19]

    Huang, L

    A. Huang, L. Leqi, Z. Lipton, and K. Azizzadenesheli. Off-policy risk assessment for markov decision processes. InInternational Conference on Artificial Intelligence and Statistics, pages 5022–5050. PMLR, 2022

  20. [20]

    Jiang and L

    N. Jiang and L. Li. Doubly robust off-policy value evaluation for reinforcement learning. In International Conference on Machine Learning, pages 652–661. PMLR, 2016

  21. [21]

    J. Kiefer. On bahadur’s representation of sample quantiles.The Annals of Mathematical Statistics, 38(5):1323–1342, 1967

  22. [22]

    Kober, J

    J. Kober, J. A. Bagnell, and J. Peters. Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 32(11):1238–1274, 2013

  23. [23]

    R. Koenker. Quantile regression [m].Econometric Society Monographs, Cambridge University Press, Cambridge, 2005

  24. [24]

    S. N. Lahiri and S. Sun. A Berry–Esseen theorem for sample quantiles under weak dependence. The Annals of Applied Probability, 19(1):108 – 126, 2009

  25. [25]

    P. W. Lavori and R. Dawson. Dynamic treatment regimes: practical design considerations. Clinical trials, 1(1):9–20, 2004

  26. [26]

    X. Li, J. Liang, and Z. Zhang. Online statistical inference for nonlinear stochastic approximation with markovian data.arXiv preprint arXiv:2302.07690, 2023

  27. [27]

    X. Li, W. Yang, J. Liang, Z. Zhang, and M. I. Jordan. A statistical analysis of polyak-ruppert averaged q-learning. InInternational Conference on Artificial Intelligence and Statistics, pages 2207–2261. PMLR, 2023. 38

  28. [28]

    P. Massart. The tight constant in the dvoretzky-kiefer-wolfowitz inequality.The annals of Probability, pages 1269–1283, 1990

  29. [29]

    Morimura, M

    T. Morimura, M. Sugiyama, H. Kashima, H. Hachiya, and T. Tanaka. Nonparametric return distribution approximation for reinforcement learning. InProceedings of the 27th International Conference on Machine Learning (ICML-10), pages 799–806, 2010

  30. [30]

    Naeem, S

    F. Naeem, S. Seifollahi, Z. Zhou, and M. Tariq. A generative adversarial network enabled deep distributional reinforcement learning for transmission scheduling in internet of vehicles.IEEE Transactions on Intelligent Transportation Systems, 22(7):4550–4559, 2020

  31. [31]

    Gpt-4 technical report, 2023

    OpenAI. Gpt-4 technical report, 2023

  32. [32]

    Ouyang, J

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155, 2022

  33. [33]

    Y. Peng, L. Zhang, and Z. Zhang. Statistical efficiency of distributional temporal difference learning.Advances in Neural Information Processing Systems, 37:24724–24761, 2024

  34. [34]

    Y. Peng, K. Jin, L. Zhang, and Z. Zhang. A finite sample analysis of distributional td learning with linear function approximation.arXiv preprint arXiv:2502.14172, 2025

  35. [35]

    S. Portnoy. Nearly root-$n$ approximation for regression quantile processes.Annals of Statistics, 40:1714–1736, 2012

  36. [36]

    Z. Qi, C. Bai, Z. Wang, and L. Wang. Distributional off-policy evaluation in reinforcement learning.Journal of the American Statistical Association, 120(551):1517–1530, 2025

  37. [37]

    Rowland, M

    M. Rowland, M. Bellemare, W. Dabney, R. Munos, and Y. W. Teh. An analysis of categorical distributional reinforcement learning. InInternational Conference on Artificial Intelligence and Statistics, pages 29–37. PMLR, 2018

  38. [38]

    Rowland, R

    M. Rowland, R. Munos, M. G. Azar, Y. Tang, G. Ostrovski, A. Harutyunyan, K. Tuyls, M. G. Bellemare, and W. Dabney. An analysis of quantile temporal-difference learning.arXiv preprint arXiv:2301.04462, 2023. 39

  39. [39]

    Rowland, Y

    M. Rowland, Y. Tang, C. Lyle, R. Munos, M. G. Bellemare, and W. Dabney. The statistical benefits of quantile temporal-difference learning for value estimation. InInternational Conference on Machine Learning, pages 29210–29231. PMLR, 2023

  40. [40]

    D. W. Scott.Multivariate density estimation. Wiley Online Library, 2015

  41. [41]

    Shao and Z.-S

    Q.-M. Shao and Z.-S. Zhang. Berry–Esseen bounds for multivariate nonlinear statistics with applications to M-estimators and stochastic gradient descent algorithms.Bernoulli, 28(3):1548 – 1576, 2022

  42. [42]

    Shao and W.-X

    Q.-M. Shao and W.-X. Zhou. Cramér type moderate deviation theorems for self-normalized processes.Bernoulli, 22(4), 2016

  43. [43]

    C. Shi, S. Zhang, W. Lu, and R. Song. Statistical inference of the value function for reinforcement learning in infinite-horizon settings.Journal of the Royal Statistical Society Series B: Statistical Methodology, 84(3):765–793, 2022

  44. [44]

    Silver, T

    D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al. A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362(6419):1140–1144, 2018

  45. [45]

    H. A. Simon. Dynamic programming under uncertainty with a quadratic criterion function. Econometrica, Journal of the Econometric Society, pages 74–81, 1956

  46. [46]

    R. S. Sutton. The reward hypothesis, 2004. URL http://incompleteideas.net/rlai.cs. ualberta.ca/RLAI/rewardhypothesis.html

  47. [47]

    R. S. Sutton and A. G. Barto.Reinforcement learning: An introduction. MIT press, 2018

  48. [48]

    H. Theil. A note on certainty equivalence in dynamic planning.Econometrica: Journal of the Econometric Society, pages 346–349, 1957

  49. [49]

    Thomas, G

    P. Thomas, G. Theocharous, and M. Ghavamzadeh. High-confidence off-policy evaluation. In Proceedings of the AAAI Conference on Artificial Intelligence, 2015

  50. [50]

    van der Vaart.Asymptotic statistics, volume 3

    A. van der Vaart.Asymptotic statistics, volume 3. Cambridge university press, 2000

  51. [51]

    van der Vaart and J

    A. van der Vaart and J. A. Wellner.Weak Convergence and Empirical Processes: With Applications to Statistics. Springer Nature, 2023. 40

  52. [52]

    Vershynin.High-Dimensional Probability: An Introduction with Applications in Data Science

    R. Vershynin.High-Dimensional Probability: An Introduction with Applications in Data Science. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press,

  53. [53]

    doi: 10.1017/9781108231596

  54. [54]

    Vinyals, I

    O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al. Grandmaster level in starcraft ii using multi-agent reinforcement learning.Nature, 575(7782):350–354, 2019

  55. [55]

    R. Wu, M. Uehara, and W. Sun. Distributional offline policy evaluation with predictive error guarantees.arXiv preprint arXiv:2302.09456, 2023

  56. [56]

    D. Yang, L. Zhao, Z. Lin, T. Qin, J. Bian, and T.-Y. Liu. Fully parameterized quantile function for distributional reinforcement learning.Advances in neural information processing systems, 32, 2019

  57. [57]

    W. Yang, L. Zhang, and Z. Zhang. Toward theoretical understandings of robust markov decision processes: Sample complexity and asymptotics.The Annals of Statistics, 50(6):3223–3248, 2022

  58. [58]

    Zhang, Y

    L. Zhang, Y. Peng, J. Liang, W. Yang, and Z. Zhang. Estimation and inference in distributional reinforcement learning.The Annals of Statistics, 53(5):1987–2011, 2025

  59. [59]

    F. Zhou, J. Wang, and X. Feng. Non-crossing quantile regression for distributional reinforcement learning.Advances in Neural Information Processing Systems, 33, 2020

  60. [60]

    1 m mÿ j“1 p pF pnq s,a ´F s,aqpθmps, iq `u´γθ mps1, jqq and forλPR, we have EexppλZ pnq s,i q ď ÿ aPA πpa|sqEexp ˜ λ ÿ s1PS ws,i,a,s1

    Y. Zhu, J. Dong, and H. Lam. Uncertainty quantification and exploration for reinforcement learning.Operations Research, 2023. 41 A Preliminary Lemmas A.1 Proof of Proposition 2.2 Proof.Define Gπ H psq “ Hÿ t“0 γtRt, where S0 “s , At|St „πp¨ |S tq, Rt „P Rp¨ |S t, Atq and St`1 | pS t, Atq „Pp¨ |S t, Atq. The density ofG π H psqcan be written as pGH s pxq “...