Pith. sign in

REVIEW 2 major objections 5 minor 34 references

This paper proves that the sampling error of the empirical value function in infinite-horizon discounted stochastic control is asymptotically a Gaussian process, obtained by propagating a one-step empirical limit through the optimal closed-

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 07:11 UTC pith:ZPWMZWSA

load-bearing objection Genuinely new infinite-horizon CLT for SAA value functions, honestly conditional on an empirical linearization condition; the continuous-state applications have a patchable gap in one Donsker lemma. the 2 major comments →

arxiv 2607.21520 v1 pith:ZPWMZWSA submitted 2026-07-23 math.OC math.PR

Asymptotic Analysis of Empirical Dynamic Programming in Infinite-Horizon Stochastic Optimal Control

classification math.OC math.PR MSC 62E2090C4049L2093E20
keywords empirical dynamic programmingsample average approximationinfinite-horizon stochastic optimal controlfunctional central limit theoremrandom fixed point equationGaussian processinventory controlrenewable harvesting
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper establishes a statistical limit theory for the sample-average approximation of infinite-horizon discounted stochastic optimal control. Its central result is a functional central limit theorem: the square-root-of-N error between the SAA value function and the true value function converges to a Gaussian process G that solves the linear random fixed point equation G = H + γLG, where H is a static Gaussian process representing one-period sampling error and L propagates error along the optimal closed-loop dynamics. The result matters because it turns sampling noise in data-driven dynamic programming into a quantitative law: the error is Gaussian, its variance is computable from the model, and the same covariance kernel explains both value-based and policy-based estimators, whose variances can differ. The paper also proves that when optimal policies are nonunique, the limit can be non-Gaussian but still solves a random fixed point equation, and it verifies the theory in inventory control and renewable harvesting.

Core claim

The paper's central claim is Theorem 2.1: under compactness-Lipschitz assumptions (Assumption 1), uniqueness of optimal one-step costs and transitions (Assumption 2), and an empirical linearization condition (11), the SAA value function V̂_N satisfies √N(V̂_N - V) ⇝ G := (I-γL)^{-1}H in C(X), where H is a centered Gaussian process with covariance Cov_ξ[Φ*(x,ξ), Φ*(x',ξ)] and L is the closed-loop transition operator of an optimal policy. The proof separates the statistical step—a functional CLT for the empirical operator evaluated at the true value function—from the dynamical step, in which the error is pushed through the fixed point equation; the linearization condition asserts that the only

What carries the argument

The key machinery is the decomposition of √N(V̂_N - V) into a static empirical-process limit H plus a propagation term γL(V̂_N - V), with L defined by [LW](x) = E[W(F*(x,ξ))] for the optimal next-state map F*. The load-bearing identity is the fixed-point rearrangement (I-γL)√N(V̂_N - V) = √N(T̂_N V - TV) + o_p(1), which is valid exactly when the empirical linearization condition (11) holds; the operator I-γL is invertible because γL has norm at most γ < 1. Linearization of the population operator T at V (Hadamard differentiability with derivative γL) and stochastic equicontinuity of the empirical operator at random arguments are the two sufficient conditions that the paper verifies via compa

Load-bearing premise

The load-bearing premise is the empirical linearization condition (11): at the √N scale, the empirical DP operator evaluated at the estimated value function differs from its value at the true value function only by the linear term γL(V̂_N − V); the paper verifies it only under sufficient conditions such as compactness of the transition operator or finite control sets, and states that verifying it in general continuous-state models remains open.

What would settle it

Simulate the inventory model of Section 4 (which satisfies all sufficient conditions) for several large N and compare the empirical distribution of √N(V̂_N(x_0) − V(x_0)) with the Gaussian prediction N(0, σ²_DP,γ(x_0)) using (30). If the agreement deteriorates rather than improves with N, the functional CLT or its variance formula is wrong. More decisively, a continuous-state model satisfying Assumptions 1–2 but with non-compact transition operator and a verified failure of (11) would show the claimed Gaussian law does not hold in that regime.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • For every fixed initial state, √N(V̂_N(x0)−V(x0)) converges to N(0,σ²_DP,γ(x0)), so the theory provides a distributional approximation for the error of the SAA optimal value.
  • The limiting variance averages the same covariance kernel over two independent discounted occupation measures, while trajectory-based SAA averages only the diagonal; hence the two estimators can have different asymptotic variances, and Example 3.3 exhibits ratio regimes below, equal to, and above one.
  • In finite state-action problems the limit is (I−γP⋆)^{-1}H, a Gaussian vector whose covariance is computable from the optimal transition matrix P⋆.
  • When optimal policies are nonunique and induce different one-step outcomes, the limiting law solves a nonlinear random fixed point equation and need not be Gaussian (Theorem 2.7).
  • The hypotheses of the CLT are verified for a discounted inventory model with base-stock policy and for a renewable-harvesting model with escapement policy, yielding explicit Gaussian process limits in both applications.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A natural corollary, not developed in the paper, is that the same covariance kernel can be estimated from the sample, yielding plug-in standard errors for the value function at any state.
  • Because the two variances in Proposition 3.2 average the same kernel differently, model selection between value-based and trajectory-based estimators will generally change the width of confidence intervals, especially for discount factors near one.
  • The paper's vanishing-discount remark suggests a testable bridge: as γ→1, the normalized variances (1−γ)²σ²_DP and (1−γ²)σ²_π may converge to the same or related average-cost quantities; if they do, the Gaussian limits connect to average-cost control.
  • In nonunique-policy regimes the non-Gaussian limiting equation could imply heavy-tailed value-function estimators; one could test this by simulating a model with two globally optimal policies and checking for non-Gaussian quantiles.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper develops statistical limit theory for sample-average-approximation (SAA) value functions in infinite-horizon discounted stochastic optimal control. The main result, Theorem 2.1, states that under compactness/Lipschitz assumptions and a uniqueness condition on the one-step outcomes induced by population optimal policies, the static empirical error N^{1/2}(\hat T_N V - TV) converges to a Gaussian process H; if an additional empirical linearization condition (11) holds, then N^{1/2}(\hat V_N - V) converges to (I-\gamma L)^{-1}H in C(X). The paper proves this decomposition and gives sufficient conditions for the linearization via compactness of the transition operator, finite control sets, or Donsker-type conditions. It also compares those asymptotics with trajectory-based SAA policy optimization, derives closed-form variance identities, extends the theory to nonunique optimal policies with possibly non-Gaussian limits, and applies the results to inventory control and renewable harvesting.

Significance. If the advertised theorems are fully secured, the paper provides a principled large-sample distributional theory for empirical dynamic programming in infinite-horizon control, going beyond the finite-horizon recursions of [17] and complementing finite-state results. The variance identities contrasting SAA DP with trajectory-based policy optimization are interesting and potentially useful for uncertainty quantification. The paper is transparent about the key assumption (11), which is honestly identified as an additional hypothesis and not disguised as a consequence of the primitive assumptions. The applications to inventory and harvesting are well chosen and illustrate the scope of the sufficient conditions. However, the bridge from the conditional main theorem to the continuous-state applications is not fully proved as written; the gaps are localized but load-bearing.

major comments (2)
  1. [Section 2.3, Lemma 2.6] Lemma 2.6 is not self-contained as stated. It asserts that the class {ξ ↦ W(F(x,u,ξ)) : W ∈ V} is P-Donsker in cases (i)–(iii) without imposing any regularity on F as a function of ξ beyond continuity and without providing the bracketing/entropy calculation. For case (ii), a common Lipschitz constant on W∈V does not by itself make the composition W∘F Donsker for a merely continuous F. The inventory and harvesting applications do have F Lipschitz in ξ, so add an explicit hypothesis (e.g., F(·,·,ξ) Lipschitz in ξ with square-integrable constant) and prove the entropy bound. As written, the proof of (19), and hence the applications in Sections 4–5, is incomplete.
  2. [Section 4.2, Proposition 4.2] The proof of Proposition 4.2 says 'standard fixed point arguments ensure that V and, with probability one, \hat V_N are Lipschitz continuous with Lipschitz constant (1−γ)^{-1}(c+b+h)', but no argument is supplied. This membership in a common deterministic Lipschitz class is required by Lemma 2.6(ii) to verify (19). Since the SAA operator uses the same f and F, the bound can be proved by induction, but the manuscript should include the proof. The same issue affects Proposition 5.3.
minor comments (5)
  1. [Abstract and Introduction] The abstract states that the functional CLT holds 'under a uniqueness-type condition on population optimal policies' without mentioning the empirical linearization condition (11), which is not implied by Assumptions 1–2. The introduction states it correctly; please qualify the abstract to avoid overstatement.
  2. [Proof of Proposition 3.2] The notation E[H(x_t)|H] is misleading; since H(x_t) is H-measurable, read literally this equals H(x_t), which would invalidate the independent-copy covariance computation. The intended object is the t-step kernel applied to H, i.e. L^t H(x_0)=E[H(x_t)|x_0] with x_0 fixed. Please correct the notation.
  3. [Introduction, p. 2] The sentence 'Under regularity conditions, including a uniqueness-type condition on the one-step outcomes induced by population optimal policies and an empirical linearization condition, we characterize the error’s limiting law under a uniqueness-type condition ...' repeats 'under a uniqueness-type condition'; the second occurrence should be removed.
  4. [Figure 2] Panel (b) is labeled 'Limiting SAA policy density'; for consistency with Section 4.3 it should read 'Limiting trajectory-SAA policy-optimization density'.
  5. [Section 2.3, Lemma 2.6(iii)] The proof cites Example 2.10.9 in [30] rather tersely. Please spell out how the convex Lipschitz class has the required entropy bound in dimension at most three, since this is the only place the dimension restriction is justified.

Circularity Check

0 steps flagged

No significant circularity: the CLT is a conditional theorem with explicitly stated hypotheses, and no fitted parameter or self-citation chain forces the result.

full rationale

The derivation chain is self-contained given the model assumptions. Assumptions 1-2 define the population objects (V, L, Phi*, pi*) and are used to obtain the static functional CLT (9) through an empirical-process/Donsker argument and Hadamard differentiability of the infimum operator; no parameter is fitted to data and then renamed as a prediction. The fixed-point limit (12) is obtained by algebraically rearranging the SAA and population DP equations together with the explicitly stated empirical linearization condition (11), which the paper transparently identifies as an additional condition rather than concealing it as a consequence of Assumptions 1-2. The paper then devotes Section 2.3 and the appendices to verifiable sufficient conditions for (11), including compactness of the transition operator, finite control sets, and Donsker-type entropy conditions. The limiting covariance (10) and the operator (I-gamma L)^{-1} depend on the true value function and true optimal transition, which is standard target dependence in asymptotic statistics, not circular construction. Self-citations such as [17] and [18] are contextual or speculative extensions and are not load-bearing for the main theorem; external structural results from [22] and [25] are used as independent support in the applications. The noted gaps in verifying the sufficient conditions, such as the asserted common Lipschitz bound in the inventory application, are correctness/verification concerns rather than circularity. I therefore find no circular step.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

The central claim rests on the compact-Lipschitz model class, the uniqueness of optimal one-step outcomes, and the empirical linearization condition. No free parameters are fitted to data; discount factor, costs, and model primitives are inputs from the applications. The Gaussian processes H, G, and Z are standard asymptotic objects, not new physical entities.

axioms (5)
  • domain assumption Assumption 1: X, U, and Ξ are compact; f and F are continuous and Lipschitz in (x,u) with square-integrable modulus K(ξ).
    Defines the model class and is used throughout to obtain contraction, Hölder continuity of V, and Donsker properties.
  • domain assumption Assumption 2: all optimal actions at a given state induce the same one-step cost f*(x,ξ) and next-state map F*(x,ξ).
    Core uniqueness condition; makes L well-defined and Ψ Hadamard differentiable tangentially to S*(Q(V)).
  • ad hoc to paper Empirical linearization condition (11): N^{1/2}(\hat T_N \hat V_N - \hat T_N V - γL(\hat V_N - V)) = o_p(1).
    Stated as an additional hypothesis of Theorem 2.1 and verified in the applications through compactness or finite unique-policy conditions.
  • standard math Donsker/empirical process theory and the functional delta method (van der Vaart and Wellner), including entropy criteria for P-Donsker classes.
    Used in Lemma A.3 and Theorem 2.7 to obtain the empirical-process CLT and Hadamard-derivative propagation.
  • domain assumption Reed (1979), Theorem 1: under the monotonicity condition (33), the harvesting problem admits a base-stock-type optimal policy and the transformed value function is concave.
    Imported to verify Assumption 2 in the renewable-harvesting application.

pith-pipeline@v1.3.0-alltime-deepseek · 27034 in / 19522 out tokens · 182990 ms · 2026-08-01T07:11:19.909233+00:00 · methodology

0 comments
read the original abstract

We derive statistical limit theorems for sample-based approximations of infinite-horizon discounted stochastic optimal control problems in discrete time. Our first result is a functional central limit theorem for the sample-based value function under a uniqueness-type condition on population optimal policies. The limiting law is a mean-zero Gaussian process characterized by a linear fixed point equation that resembles a dynamic programming principle. We compare these asymptotics with those obtained from sample-based policy optimization and illustrate that their limiting variances can be different. We also derive a limit theorem for models with nonunique optimal policies, where the limiting law may be non-Gaussian. Applications to inventory control and renewable harvesting illustrate the theory.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

34 extracted references · 10 canonical work pages · 3 internal anchors

  1. [1]

    Arlotto and J

    A. Arlotto and J. M. Steele. A central limit theorem for temporally nonhomoge- nous Markov chains with applications to dynamic programming.Math. Oper. Res., 41(4):1448–1468, 2016.doi:10.1287/moor.2016.0784

  2. [2]

    Billingsley.Convergence of Probability Measures

    P . Billingsley.Convergence of Probability Measures. Wiley Ser. Probab. Stat. John Wiley & Sons, New York, NY, 2nd edition, 1999.doi:10.1002/9780470316962

  3. [4]

    Brezis.Functional Analysis, Sobolev Spaces and Partial Differential Equations

    H. Brezis.Functional Analysis, Sobolev Spaces and Partial Differential Equations. Uni- versitext. Springer, New York, NY, 2011.doi:10.1007/978-0-387-70914-7

  4. [5]

    Chen and J.-S

    F. Chen and J.-S. Song. Optimal policies for multiechelon inventory problems with Markov-modulated demand.Oper. Res., 49(2):226–234, 2001.doi:10.1287/opre. 49.2.226.13528

  5. [6]

    Eichhorn and W

    A. Eichhorn and W. R ¨omisch. Stochastic integer programming: Limit theorems and confidence intervals.Math. Oper. Res., 32(1):118–135, 2007.doi:10.1287/ moor.1060.0222

  6. [7]

    Ekerhovd, S

    N.-A. Ekerhovd, S. D. Fl ˚am, and S. I. Steinshamn. On shared use of renewable stocks.Eur. J. Oper. Res., 290(3):1125–1135, 2021

  7. [8]

    E. A. Feinberg, P . O. Kasyanov, and N. V . Zadoianchuk. Average cost Markov decision processes with weakly continuous transition probabilities.Math. Oper. Res., 37(4):591–607, 2012.doi:10.1287/moor.1120.0555

  8. [9]

    Ga ¨ıgi, V

    M. Ga ¨ıgi, V . Ly Vath, and S. Scotti. Optimal harvesting under marine reserves and uncertain environment.Eur. J. Oper. Res., 301(3):1181–1194, 2022

  9. [10]

    Gupta, R

    A. Gupta, R. Jain, and P . Glynn. Probabilistic contraction analysis of iterated ran- dom operators.IEEE Trans. Autom. Control, 69(9):5947–5962, 2024.doi:10.1109/ TAC.2024.3362686

  10. [11]

    W. B. Haskell, R. Jain, and D. Kalathil. Empirical dynamic programming.Math. Oper. Res., 41(2):402–429, 2016.doi:10.1287/moor.2015.0733

  11. [12]

    W. B. Haskell, R. Jain, H. Sharma, and P . Yu. A universal empirical dynamic programming algorithm for continuous state MDPs.IEEE Trans. Autom. Control, 65(1):115–129, 2020.doi:10.1109/TAC.2019.2907414

  12. [13]

    Hern ´andez-Lerma and J

    O. Hern ´andez-Lerma and J. B. Lasserre.Discrete-Time Markov Control Processes: Basic Optimality Criteria. Springer, New York, 1996.doi:10.1007/978-1-4612- 0729-0

  13. [14]

    W. W. Hogan. Point-to-set maps in mathematical programming.SIAM Rev., 15(3):591–603, 1973.doi:10.1137/1015073. 33

  14. [15]

    A. J. Kleywegt, A. Shapiro, and T. Homem-de Mello. The sample average approxi- mation method for stochastic discrete optimization.SIAM J. Optim., 12(2):479–502, 2002.doi:10.1137/S1052623499363220

  15. [16]

    Mannor, D

    S. Mannor, D. Simester, P . Sun, and J. N. Tsitsiklis. Bias and variance approximation in value function estimates.Manage. Sci., 53(2):308–322, February 2007.doi:10. 1287/mnsc.1060.0614

  16. [17]

    Central Limit Theorems for Sample Average Approximations in Stochastic Optimal Control

    J. Milz and A. Shapiro. Central limit theorems for sample average approximations in stochastic optimal control, 2025.arXiv:2508.01942,doi:10.48550/arxiv. 2508.01942

  17. [18]

    J. Milz, A. Shapiro, and E. Zhou. Stochastic optimal control with side informa- tion and Bayesian learning, 2026.arXiv:2602.22047,doi:10.48550/arxiv.2602. 22047

  18. [19]

    G. C. Nyassoke Titi, J. Sadefo Kamdem, and L. A. Fono. Optimal renewable re- source harvesting model using price and biomass stochastic variations: A utility based approach.Math. Methods Operations Res., 95(2):297–326, 2022

  19. [20]

    Pennanen and A.-P

    T. Pennanen and A.-P . Perkki ¨o. Dynamic programming and dimensionality in convex stochastic optimization and control.Optim. Lett., 20(2):257–273, August 2026.doi:10.1007/s11590-025-02230-4

  20. [21]

    W. B. Powell.Approximate dynamic programming. Solving the curses of dimensionality. Wiley Ser. Probab. Stat. Hoboken, NJ: John Wiley & Sons, 2nd ed. edition, 2011. doi:10.1002/9781118029176

  21. [22]

    W. J. Reed. Optimal escapement levels in stochastic and deterministic harvesting models.J. Environ. Econ. Manage., 6(4):350–363, 1979

  22. [23]

    S. P . Sethi and F. Cheng. Optimality of(s,S)policies in inventory models with Markovian demand.Oper. Res., 45(6):931–939, 1997.doi:10.1287/opre.45.6.931

  23. [24]

    A. Shapiro. Monte Carlo Sampling Methods. InStochastic Programming, Hand- books in Oper. Res. Manag. Sci. 10, pages 353–425. Elsevier, 2003.doi:10.1016/ S0927-0507(03)10006-0

  24. [25]

    Shapiro and Y

    A. Shapiro and Y. Cheng. Central limit theorem and sample complexity of station- ary stochastic programs.Oper. Res. Lett., 49(5):676–681, 2021.doi:10.1016/j.orl. 2021.06.019

  25. [26]

    Shapiro, D

    A. Shapiro, D. Dentcheva, and A. Ruszczy ´nski.Lectures on Stochastic Programming: Modeling and Theory. MOS-SIAM Ser. Optim. SIAM, Philadelphia, PA, 3rd edition, 2021.doi:10.1137/1.9781611976595

  26. [27]

    Shapiro and L

    A. Shapiro and L. Ding. Periodical multistage stochastic programs.SIAM J. Optim., 30(3):2083–2102, 2020.doi:10.1137/19M129406X

  27. [28]

    M. J. Sobel. The variance of discounted Markov decision processes.J. Appl. Probab., 19:794–802, 1982.doi:10.2307/3213832. 34

  28. [29]

    Z. Su, I. Banerjee, and D. Klabjan. Central limit theorems for transition probabili- ties of controlled Markov chains, 2025.doi:10.48550/ARXIV.2508.01517

  29. [30]

    A. W. van der Vaart and J. A. Wellner.Weak Convergence and Empirical Processes. With Applications to Statistics. Springer Ser. Stat. Cham: Springer, 2nd edition, 2023. doi:10.1007/978-3-031-29040-4

  30. [31]

    A. W. van der Vaart and J. A. Wellner. Empirical processes indexed by estimated functions. InAsymptotics: Particles, Processes and Inverse Problems, pages 234–252. Beachwood, OH: IMS, Institute of Mathematical Statistics, 2007.doi:10.1214/ 074921707000000382

  31. [32]

    Lipschitz Regularity in Wasserstein Robust Stochastic Optimal Control

    S. Wang and J. Blanchet. Lipschitz regularity in Wasserstein robust stochastic op- timal control, 2026.arXiv:2606.28633,doi:10.48550/arxiv.2606.28633

  32. [33]

    Zhang, Z.-S

    X. Zhang, Z.-S. Ye, and W. B. Haskell. Error propagation in asymptotic analysis of the data-driven(s,S)inventory policy.Oper. Res., 73(1):1–21, 2025.doi:10.1287/ opre.2020.0568

  33. [34]

    Y. Zhu, J. Dong, and H. Lam. Uncertainty quantification and exploration for rein- forcement learning.Oper. Res., 72(4):1689–1709, 2024.doi:10.1287/opre.2023. 2436

  34. [35]

    Zipkin.Foundation of inventory management

    P . Zipkin.Foundation of inventory management. McGraw-Hill, Boston, MA, 2000. 35