REVIEW 3 major objections 3 minor 1 cited by
Learning an unknown noise distribution during control yields a Bayesian value that is asymptotically normal around the true optimum at √N rate.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 20:48 UTC pith:K3GIQTL4
load-bearing objection The framework and consistency results are real; the advertised Bernstein–von Mises normality is conditional on an unproved √N-equivalence condition (23) plus a d-dimensional BvM limit the paper only states. the 3 major comments →
Stochastic Optimal Control with Side Information and Bayesian Learning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is Theorem 4.1: for a correctly specified, identifiable parametric model with a fixed initial state-context pair, √N(V*_N(x1,η1) − V*(x1,η1)) converges in distribution to a normal with mean zero and variance ∇g(θ*)^T I(θ*)^−1 ∇g(θ*), where V*_N is the Bayesian Bellman value, V* is the true optimal value, I(θ*) is the Fisher information of the observed context-randomness Markov chain, and ∇g(θ*) is the gradient of the optimal discounted cost with respect to the model parameter (given by the discounted cumulative cost times the score). The theorem holds conditional on an unproven coupling condition (23)—that the Bayesian value equals the posterior mean of the true-pol
What carries the argument
The Bayesian Bellman equation (7)—replacing the unknown conditional density q(ξ|η) by the posterior predictive expectation E_{θ∼p_N} E_{ξ∼f(·|η,θ)} in the Bellman recursion—is the central object. Its contraction property yields a unique Bayesian value function V*_N and policy π*_N. The asymptotic result hinges on Lemma 4.1, which differentiates the infinite-horizon discounted return with respect to θ, giving ∇g(θ) = E_{Pθ}[Σ_{t≥1} γ^{t−1} c_t S_{θ,t}]; this gradient formula converts a Bernstein–von Mises limit for θ into asymptotic normality for the value.
Load-bearing premise
The theorem assumes, without proof, that the Bayesian optimal value and the posterior mean of the true-policy value differ by o_p(N^{−1/2}) at the initial state-context pair; if this coupling fails, the advertised asymptotic normality for the Bayesian value does not follow.
What would settle it
Simulate the linear-Gaussian regression example (finite context set, a small discounted horizon), compute V*_N and E_{θ∼p_N}[V^{π*}_θ] exactly, and test whether √N times their difference vanishes as N grows; also compute the empirical variance of √N(V*_N − V*) and compare it to ∇g(θ*)^T I(θ*)^−1 ∇g(θ*). A persistent mismatch or a non-vanishing coupling would refute Theorem 4.1.
If this is right
- If the conditions hold, a practitioner can construct a confidence interval for the true optimal value from the posterior and the Fisher information, without resampling.
- The uniform convergence result means the Bayesian policy is asymptotically optimal on all states and contexts, not only at training points.
- Markov dependence in the context observations does not break posterior consistency; Bayesian learning remains valid for regime-switching or Markov-modulated control problems.
- The asymptotic variance formula separates statistical uncertainty (inverse Fisher information) from control-theoretic sensitivity (gradient of the discounted cost), allowing each to be studied independently.
Where Pith is reading between the lines
- The unproven condition (23) is the load-bearing coupling: without it, Theorem 4.1 describes the posterior mean of the true-policy value, not the Bayesian optimal value a controller would actually compute. A natural next step is to verify (23) for particular models (e.g., the linear-Gaussian example) or find a counterexample where the Bayesian policy's value deviates at the √N scale.
- The paper imports the Bernstein–von Mises limit for Markov chains (17) without proof; a self-contained proof of that limit for the specific chain (ξ_t, η_t) would close the gap and is needed for the theorem to be fully grounded.
- The result suggests a testable extension: in finite-horizon or episodic versions of this problem, uniform-in-(x,η) Bernstein–von Mises limits might be attainable with the authors' related CLT techniques, and one could check whether the uniform rate differs from √N.
- Replacing the posterior expectation in the Bayesian Bellman equation with a posterior risk measure would yield a different asymptotic variance, computable with the same gradient-sensitivity machinery, connecting to distributionally robust Bayesian control.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies infinite-horizon stochastic optimal control with finite-state Markovian side information (context), where the conditional distribution of the randomness given the context is unknown and learned from data via a parametric Bayesian model. The authors propose a Bayesian Bellman equation based on posterior predictive expectations, prove posterior consistency under Markov samples and uniform convergence of the Bayesian value function under correct specification and identifiability, and claim a Bernstein–von Mises-type asymptotic normality for the data-driven contextual optimal value at a fixed initial state. The main theorem (Theorem 4.1) is conditional on two unproved high-level ingredients: a d-dimensional Bernstein–von Mises limit for Markov chains (Section 4.1) and a √N-equivalence condition (23) between the Bayesian optimal value and the posterior mean of the true-policy value.
Significance. The modeling framework is timely and the consistency part is a reasonable extension of episodic Bayesian optimal control to a Markovian context setting. If the asymptotic normality result were fully established, it would be a meaningful contribution. However, the advertised BvM-type theorem is not actually proved: the key hypotheses (16)–(17) in Section 4.1 are stated without proof, and condition (23) is a high-level assumption of the same asymptotic character as the conclusion. The uniform LLN lemma also lacks a valid proof. As submitted, the manuscript does not deliver its central advertised claim, though the underlying ideas may be salvageable with substantial additional work.
major comments (3)
- [§4.1, Eqs. (16)-(17)] The Bernstein–von Mises limits for the Markov chain are stated without proof, and the cited references [3,10] apply only to one-dimensional parameter spaces, whereas Θ⊂R^d here. The Delta-method limit (17) is asserted following [2] without verifying its regularity conditions for g(θ)=V^{π*}_θ(x1,η1). Since this limit is a direct input to Theorem 4.1, the main theorem rests on an unproved and non-standard high-level condition.
- [§4.2, condition (23)] Condition (23) is assumed without proof: N^{1/2}(V*_N(x1,η1) − E_{θ∼p_N}[V^{π*}_θ(x1,η1)]) → 0 in P*-probability. The discussion after (27) only says the condition 'partially resembles' eq. (5.24) of [20] and notes that the right-hand side involves two distinct value functions. No contraction, envelope, or policy-differentiability argument is supplied. If (23) fails, the centered quantity in (24) differs from the posterior-mean quantity by a term that need not vanish, so the BvM conclusion does not follow. This is a load-bearing assumption of the same asymptotic order as the theorem's conclusion.
- [Lemma 3.1] The proof claims that pointwise LLN plus dominated integrability and compactness imply a uniform LLN over Θ. That implication is not generally valid; uniform convergence requires additional equicontinuity or bracketing/entropy conditions. Consequently, Lemma 3.1 is not established, and the results that rely on it (Theorem 3.1 and Proposition 3.1) are not proven as stated. This gap is correctable by adding suitable conditions, but as written the proof is insufficient.
minor comments (3)
- [Eq. (22)] The second displayed expression for I(θ*) has a typo: it should read ∑_{h∈H} ν_η(h) E_{ξ∼f(·|h,θ*)}[s s^T], not 'νη(η)'.
- [Section 3.2, paragraph after Definition 3.1] The phrase 'the posterior p_N almost surely converges to a θ*' is imprecise; it is the posterior measure P_N that concentrates on θ*, not the density itself. Please rephrase.
- [Lemma 4.1 proof] The application of Theorem 9.56 of [20] requires regularity conditions (domination, differentiability under the integral). These conditions are plausible under Assumption 4.1 but should be stated explicitly.
Circularity Check
No circular derivation: Theorem 4.1 is explicitly conditional on unproved high-level conditions (17) and (23); the consistency results are derived from stated assumptions; self-citations are not load-bearing.
full rationale
No circular step can be exhibited. The consistency portion (Lemmas 3.1–3.3, Theorem 3.1, Proposition 3.1) is a direct derivation: posterior consistency follows from a uniform LLN and exponential posterior decay, and value-function consistency follows from a contraction error bound plus L1 convergence of the posterior predictive density. The self-citations to [22] and [23] are used as proof templates or standard bounds, not as unverified premises that force the conclusion. The central asymptotic claim, Theorem 4.1, is an explicitly conditional statement: the proof applies the delta-method BvM limit (17) to the posterior mean of V^{π*}_θ and then invokes the explicit hypothesis (23) that V*_N and that posterior mean are asymptotically equivalent at the √N scale. The paper itself acknowledges the limits of its support: Section 4.1 says the two BvM limits are 'stated without proof' and that the cited theorems in [3,10] 'apply only to one-dimensional parameter spaces,' while Eq. (23) is introduced as an unproved 'suppose' condition. These are substantial gaps in support and make the abstract's 'we establish' an overstatement, but they are not a case of a prediction reducing to its inputs by construction: the assumptions are transparent, condition (23) is not definitionally identical to the conclusion, and the gradient computation in Lemma 4.1 is independent. Therefore the appropriate finding is no significant circularity, with minor self-citations that are not load-bearing.
Axiom & Free-Parameter Ledger
axioms (7)
- domain assumption Known and parameter-free context transition matrix ϖ used in the posterior (5) and Bellman equations.
- domain assumption Correct specification and identifiability of the parametric model at θ* (Assumption 3.2).
- domain assumption Assumption 3.1: compact Θ, prior bounded away from 0, positive continuous f, irreducible aperiodic η chain, dominated log-likelihood.
- domain assumption Unique optimal policy π* (Assumption 4.1(iii)).
- ad hoc to paper BvM limits (16)-(17) hold for the d-dimensional Markov chain (Section 4.1).
- ad hoc to paper Condition (23): N^{1/2}(V*_N − E_{θ∼p_N}[V^{π*}_θ]) → 0 in P*-probability.
- ad hoc to paper Uniform LLN in Lemma 3.1 follows from pointwise LLN and domination.
read the original abstract
We study infinite-horizon stochastic optimal control problems with observable side information: a Markov chain that modulates an unknown context-conditional randomness distribution. Since this distribution is unknown, we propose a Bayesian reformulation based on a parametric density model and posterior predictive dynamics, which yields a Bayesian Bellman equation. We prove posterior consistency under Markov samples and, under correct specification and identifiability, uniform convergence of the Bayesian value function. Finally, we establish Bernstein--von Mises-type asymptotic normality for the data-driven contextual optimal value.
Figures
Forward citations
Cited by 1 Pith paper
-
Asymptotic Analysis of Empirical Dynamic Programming in Infinite-Horizon Stochastic Optimal Control
The sample-based value function in discounted infinite-horizon stochastic control converges to a Gaussian process limit that solves a linear DP-type fixed-point equation.
Reference graph
Works this paper leans on
-
[1]
K. Arifo˘ glu and S.¨Ozekici. Optimal policies for inventory systems with finite capacity and partially observed Markov-modulated demand and supply processes.Eur. J. Oper. Res., 204(3):421–438, 2010.doi:10.1016/j.ejor.2009.10.029
-
[2]
P. J. Bickel and J. A. Yahav. Some contributions to the asymptotic theory of Bayes solutions.Z. Wahrscheinlichkeitstheorie und Verw. Gebiete, 11:257–276, 1969.doi:10.1007/BF00531650
-
[3]
J. Borwanker, G. Kallianpur, and B. L. S. Prakasa Rao. The Bernstein-von Mises theorem for Markov processes.Ann. Math. Statist., 42:1241–1253, 1971.doi:10.1214/aoms/1177693237
arXiv 1971
-
[4]
F. Chen and J.-S. Song. Optimal policies for multiechelon inventory problems with Markov- modulated demand.Oper. Res., 49(2):226–234, 2001.doi:10.1287/opre.49.2.226.13528
-
[5]
X. Chen, Y. Hu, and M. Zhao. Landscape of policy optimization for finite horizon MDPs with general state and action, 2024.arXiv:2409.17138
arXiv 2024
-
[6]
O. L. V. Costa, M. D. Fragoso, and R. P. Marques.Discrete-time Markov jump linear systems. Probability and its Applications. Springer, London, 2005.doi:10.1007/b138575
doi:10.1007/b138575 2005
-
[7]
J. Deng, Y. Cheng, S. Zou, and Y. Liang. Sample complexity characterization for linear con- textual MDPs. InProceedings of the 27th International Conference on Artificial Intelligence and Statistics (AISTATS), volume 238 ofPMLR, 2024. URL:https://proceedings.mlr.press/ v238/deng24a/deng24a.pdf
2024
-
[8]
G. Gallego and H. Hu. Optimal policies for production/inventory systems with finite capacity and Markov-modulated demand and supply processes.Ann. Oper. Res., 126:21–41, 2004.doi: 10.1023/B:ANOR.0000012274.69117.90
arXiv 2004
-
[9]
J. C. Geromel. Markov jump linear systems. InDifferential Linear Matrix Inequalities, pages 149–189. Springer, 2023.doi:10.1007/978-3-031-29754-0_6
-
[10]
H. Gillert. The Bernstein-von Mises theorem for nonstationary Markov processes. InTrans- actions of the ninth Prague conference on information theory, statistical decision functions, random processes, Vol. A (Prague, 1982), pages 253–256. Reidel, Dordrecht, 1983.doi: 10.1007/978-94-009-7013-7_30
-
[11]
Gupta and R
V. Gupta and R. M. Murray. Lecture summary: Markov jump linear systems. Lecture notes, Caltech/Notre Dame, 2007. Accessed: February 18, 2026. URL:https://murray.cds.caltech. edu/images/murray.cds/8/84/Lecture_mjls.pdf. 11
2007
-
[12]
A. Hallak, D. D. Castro, and S. Mannor. Contextual Markov decision processes.arXiv preprint arXiv:1502.02259, 2015. URL:https://arxiv.org/abs/1502.02259
Pith/arXiv arXiv 2015
-
[13]
Langford and T
J. Langford and T. Zhang. The epoch-greedy algorithm for multi-armed bandits with side information. In J. Platt, D. Koller, Y. Singer, and S. Roweis, ed- itors,Advances in Neural Information Processing Systems, volume 20. Curran Asso- ciates, Inc., 2007. URL:https://proceedings.neurips.cc/paper_files/paper/2007/file/ 4b04a686b0ad13dce35fa99fa4161c65-Paper.pdf
2007
-
[14]
O. Levy and Y. Mansour. Optimism in face of a context: Regret guarantees for stochastic contextual MDP. InProceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 8510–8517, 2023.doi:10.1609/aaai.v37i7.26025
-
[15]
Y. Lin, Y. Ren, and E. Zhou. Bayesian risk Markov decision processes. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors,Advances in Neural Information Processing Systems, volume 35, pages 17430–17442, 2022. URL:https://proceedings.neurips.cc/paper_ files/paper/2022/file/6f7d90b1198fec96defd80b5ebd5bc81-Paper-Conference.pdf
2022
-
[16]
S. S. Malladi, A. L. Erera, and C. C. I. White. Inventory control with modulated demand and a partially observed modulation process.Ann. Oper. Res., 321(1-2):343–369, 2023.doi: 10.1007/s10479-022-04932-9
-
[17]
J. Milz and A. Shapiro. Central limit theorems for sample average approximations in stochastic optimal control, August 2025.doi:10.48550/arXiv.2508.01942
-
[18]
U. Sadana, A. Chenreddy, E. Delage, A. Forel, E. Frejinger, and T. Vidal. A survey of contextual optimization methods for decision making under uncertainty.European Journal of Operational Research, 320(2):271–289, 2025.doi:10.1016/j.ejor.2024.03.020
-
[19]
S. P. Sethi and F. Cheng. Optimality of (s, S) policies in inventory models with Markovian demand.Oper. Res., 45(6):931–939, 1997.doi:10.1287/opre.45.6.931
-
[20]
Shapiro, D
A. Shapiro, D. Dentcheva, and A. Ruszczy´ nski.Lectures on Stochastic Programming: Modeling and Theory. MOS-SIAM Ser. Optim. SIAM, Philadelphia, PA, 3rd edition, 2021.doi:10.1137/ 1.9781611976595
2021
-
[21]
A. Shapiro and L. Ding. Periodical multistage stochastic programs.SIAM J. Optim., 30(3):2083– 2102, 2020.doi:10.1137/19M129406X
-
[22]
A. Shapiro, E. Zhou, and Y. Lin. Bayesian distributionally robust optimization.SIAM J. Optim., 33(2):1279–1304, 2023.doi:10.1137/21M1465548
-
[23]
A. Shapiro, E. Zhou, Y. Lin, and Y. Wang. Episodic bayesian optimal control with unknown randomness distributions.Operations Research, 2025.doi:10.1287/opre.2023.0446
arXiv 2025
-
[24]
J.-S. Song and P. Zipkin. Inventory control in a fluctuating demand environment.Oper. Res., 41(2):351–370, 1993.doi:10.1287/opre.41.2.351
-
[25]
Tennenholtz, N
G. Tennenholtz, N. Merlis, L. Shani, M. Mladenov, and C. Boutilier. Reinforcement learning with history-dependent dynamic contexts. InProceedings of the 40th International Conference on Ma- chine Learning (ICML), 2023. URL:https://proceedings.mlr.press/v202/tennenholtz23a/ tennenholtz23a.pdf
2023
-
[26]
A. W. van der Vaart.Asymptotic Statistics. Camb. Ser. Stat. Probab. Math. 3. Cambridge University Press, Cambridge, 1998.doi:10.1017/CBO9780511802256
-
[27]
D. Wu, H. Zhu, and E. Zhou. A Bayesian risk approach to data-driven stochastic optimization: for- mulations and asymptotics.SIAM J. Optim., 28(2):1588–1612, 2018.doi:10.1137/16M1101933. 12
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.