REVIEW 1 major objections 6 minor 84 references
On Pioneering Works of Albert Shiryaev on Markov Decision Processes and Some Later Developments
T0 review · 1 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This survey argues that three papers by Albert Shiryaev from the 1960s introduced the belief-state reduction for partially observable control and proved the existence of optimal stationary policies for average-cost Markov decision…
desk verdict A solid historical survey of Shiryaev's MDP work, but Theorem 6.5(d) has a genuine proof gap that needs to be fixed before the paper can stand. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central objects are the belief MDP and the canonical average-cost equations. A belief MDP is a fully observed Markov decision process whose state is the posterior distribution of the hidden state; the survey traces this reduction to Shiryaev's papers and uses it to transfer optimality results from complete-observation theory to partially observable problems. For average-cost problems, the canonical equations $w = P^\phi w$ and $w + u = T^\phi u$ characterize optimal deterministic policies, and later sections extend their validity to Borel state spaces under Assumptions (W*) and (S*), which use K-inf-compactness to drop compactness of action sets. A further mechanism is the semi-uniform Feller property for transition kernels, which is equivalent to weak continuity of the induced belief-state transition, and the Diffeomorphic Condition, which gives checkable sufficient conditions for total-variation continuity of filters defined by explicit state and observation equations.
What would settle it
Reading the original texts of Viskov and Shiryaev [82] and Shiryaev [74, 75] to check whether they state and prove the belief-state reduction and the existence of a deterministic stationary average-cost optimal policy; if, for example, [82] proves only epsilon-optimality or assumes a unichain structure, the survey's attribution of the general theorem would be wrong. Alternatively, a finite-state/action average-cost MDP with infinite one-step costs that admits no deterministic stationary optimal policy would directly contradict the theorem as described.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that Shiryaev's 1962–1967 work already contained the two structural ideas that later became the basis of partially observable control: control with incomplete observations can be reduced to a fully observed MDP on the space of posterior distributions, and average-cost MDPs with finite state and action sets admit an optimal deterministic stationary policy. The survey attributes the belief-space reduction to [74, 75] and the average-cost optimality proof to [82], noting that [12] and [19] reached the same conclusion independently. It then presents later results that make these ideas effective on infinite state and action spaces: with K-inf-compact cost functions and weak or setwise continuity assumptions, optimality equations hold, optimal policies exist, and value iteration converges for belief MDPs and for filtering problems with additive or multiplicative noise. The later sections report the resolution of Feller's problem on forward Kolmogorov equations for jump Markov processes and its use to prove that Markov policies suffice for continuous-time jump MDPs.
Load-bearing premise
The survey's narrative stands on the claim that the three 1960s papers, which are not reproduced in the survey, actually contain the belief-state reduction and the average-cost optimality proof attributed to them; if that reading is wrong, the historical claims would need to be revised.
Editorial extensions
If this is right
- If the belief-state reduction is valid, every POMDP can be solved as a fully observed MDP over probability distributions, so dynamic programming, value iteration, and linear programming methods apply directly to partially observable problems.
- The historical priority claim implies that the belief-state reduction predates the modern POMDP and reinforcement-learning literature by decades, not the reverse.
- Under the weak-continuity and K-inf-compactness conditions described in the survey, optimal policies exist and value iteration converges for belief MDPs even when action sets are non-compact, a setting that covers inventory control and many filtering models.
- For filtering problems written with explicit state and observation equations, the Diffeomorphic Condition guarantees weak continuity of the filter, so stochastic control and stochastic filtering can be solved by the same dynamic-programming machinery.
- The resolution of Feller's forward-equation problem and the Markov-policy sufficiency result for continuous-time jump MDPs extend these guarantees from discrete time to continuous-time controlled jump processes.
Reading between the lines
- Not in the paper: the survey does not compare Shiryaev's belief-state reduction in detail with the contemporaneous formulations it cites; a careful comparison could clarify how much of the modern POMDP framework is genuinely new.
- Not in the paper: the Diffeomorphic Condition is stated for Euclidean noise with absolutely continuous distributions; testing whether total-variation continuity survives for heavier-tailed noise or for noise on manifolds would be a natural extension.
- Not in the paper: the semi-uniform Feller property could be used as a practical certificate for POMDP solvers, since it guarantees weak continuity of the belief transition and therefore convergence of value iteration; this check is not discussed in the survey.
- Not in the paper: if the historical attributions are accepted, the survey implies that the belief-state idea should be presented in reinforcement-learning textbooks as a classical control-theoretic tool rather than a modern invention.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a survey of three historically important papers by A.N. Shiryaev (one with O.V. Viskov) on Markov decision processes and control with incomplete information. It recaps the existence of deterministic stationary optimal policies for finite average-cost MDPs, the reduction of POMDPs to belief MDPs, and later results on discounted and average-cost MDPs with Borel state/action spaces, weak continuity of belief transition kernels, and Kolmogorov equations for jump Markov processes. The last part of Section 6 states a new theorem (Theorem 6.5) giving sufficient conditions for weak continuity of the belief-MDP kernel in filtering problems defined by explicit equations.
Significance. If the historical attributions are correct, the survey performs a useful service by placing Shiryaev's 1962–1967 papers in the context of modern POMDP and reinforcement learning theory. The paper is clearly written, includes precise statements of many later results, and has a comprehensive bibliography. The main new mathematical claim (Theorem 6.5) is, however, not fully supported: part (d) is stated without a valid proof, and the Conclusion uses this theorem to advertise 'broad sufficient conditions' for filtering problems. Once that gap is repaired, the paper would be a solid contribution; in its present form it requires revision.
major comments (1)
- [Section 6, Theorem 6.5(d)] The proof of Theorem 6.5 says that part (d) follows from Theorems 5.1 and 6.4, but neither sufficient condition of Theorem 5.1 applies. Condition (ii) of Theorem 5.1 requires the transition kernel T to be continuous in total variation, whereas (d) only assumes F continuous in distribution (weak continuity of T). Condition (i) requires Q to be continuous in total variation jointly in (a,x); since g is only assumed measurable in a, Theorem 6.4 cannot be invoked for the pair (a,x), and continuity of g in x alone does not imply total-variation continuity of Q in a. Thus the weak continuity of p̄ is unproven under (d). This is load-bearing because the Conclusion relies on Theorem 6.5 to claim broad sufficient conditions for filtering problems; the statement needs a corrected proof (e.g., adding joint continuity of g or a different argument) or condition (d) should be withdrawn.
minor comments (6)
- [Section 4.3] Two different assumptions are both labeled 'Assumption (B)': the first (from Schäl [69]) uses sup_alpha u_alpha(x) < ∞, the second (from [34]) uses lim inf_alpha u_alpha(x) < ∞. Please rename them (e.g., (B1) and (B2)) to avoid ambiguity.
- [Section 5, Eq. (14)] The definition of semi-uniform Feller appears to have the product spaces reversed: as written, Ψ is a kernel from S3 to S1×S2 (since the conditioning argument is s3 and the output is a measure on S1 with a set B⊂S2), but the preceding text says Ψ is a transition probability from S1×S2 to S3. Please correct the direction and the notation.
- [Section 6, Theorem 6.4] In the statement of Theorem 6.4, the codomain of φ is written as 'R' but given S1 = R^n it should be 'R^n'; also 'satisfy' should be 'satisfies'.
- [Section 2] The definition of the history sets is garbled: 'Ht := X × (A × X),,' should be H_t = X × (A × X)^t (or an equivalent description), and the sentence lacks the verb 'let'.
- [Section 6, Theorem 6.5(d)] The hypothesis says g is 'measurable' and 'continuous in variable x_{t+1}'; please state explicitly that g is jointly Borel measurable and specify whether continuity in x is uniform over compacts or pointwise, because this affects the applicability of Theorem 6.4.
- [Throughout] There are several spelling/typographical errors, e.g., 'expexted' in Section 6, 'Cezaro' for 'Cesàro' in Section 2, and 'Corollart' in Corollary 6.3. Please proofread carefully.
Circularity Check
No circularity: the historical attributions are anchored in primary sources and the later-development theorems are cited with independent proofs; the only flagged issue is a non-circular proof gap in Theorem 6.5(d).
full rationale
This is an expository survey, not a paper with fitted parameters or a derivation chain whose conclusions are encoded in its assumptions. The central historical claims—that Viskov and Shiryaev [82] proved existence of deterministic optimal policies for average-cost MDPs and that Shiryaev [74, 75] formulated the reduction of partially observable problems to belief-state MDPs—are grounded in the cited primary sources and corroborated by independent classical references such as Blackwell [12], Derman [19], Aoki [1], Åström [3], and Dynkin [20]. These claims do not reduce to the author's own work. The many self-citations (e.g., [30, 32, 34, 36, 37, 40–44]) are used to report later developments; they point to published theorems with stated assumptions and proofs, which counts as independent support rather than a self-referential derivation chain. No fitted quantity is renamed as a prediction, no known result is merely re-coined, and no uniqueness or ansatz is imported solely through self-citation. One correctness concern emerged, but it is not circularity: the proof of Theorem 6.5(d) says it follows from Theorems 5.1 and 6.4, yet the hypotheses of (d)—g only measurable in a and continuous in x, with T only continuous in distribution—do not satisfy either sufficient route of Theorem 5.1, since route (i) requires Q continuous in total variation and route (ii) requires T continuous in total variation. That is a potential proof gap in the paper's own contribution, not a case where the conclusion is equivalent to the input by construction. Therefore the circularity score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The cited papers [74, 75, 82] contain the results and ideas attributed to them
- standard math Standard measure-theoretic and analytic results (Ionescu Tulcea theorem, Banach fixed point theorem, Berge's theorem, Hardy-Littlewood Tauberian theorem, Aumann's lemma) are used without proof
Cite this review
Pith. "Pith review of On Pioneering Works of Albert Shiryaev on Markov Decision Processes and Some Later Developments." pith.science (2026). https://pith.science/paper/MVAGTOO3
@misc{pith2026250604896,
author = {Pith},
title = {Pith review of: On Pioneering Works of Albert Shiryaev on Markov Decision Processes and Some Later Developments},
year = {2026},
howpublished = {\url{https://pith.science/paper/MVAGTOO3}},
note = {Machine review of arXiv:2506.04896}
}
read the original abstract
This article is dedicated to three fundamental papers on Markov Decision Processes and on control with incomplete observations published by Albert Shiryaev approximately sixty years ago. One of these papers was coauthored with O.V. Viskov. We discuss some of the results and some of many rich ideas presented in these papers and survey some later developments. At the end we mention some recent studies of Albert Shiryaev on Kolmogorov's equations for jump Markov processes and on control of continuous-time jump Markov processes.
Reference graph
Works this paper leans on
-
[1]
(1965) Optimal control of partially observable Markovian systems
Aoki, M. (1965) Optimal control of partially observable Markovian systems. J. Franklin Inst. 280 pp. 367–386
1965
-
[2]
(1993) Discrete time controlled Markov processes with average cost criterion: a survey, SIAM J
Arapostathis, A., Borkar, V.S., Fernandez-Gaucherand, E, Ghosh M.K., Marcus, S.I. (1993) Discrete time controlled Markov processes with average cost criterion: a survey, SIAM J. Control Optim. 31(2) 282–344
1993
-
[3]
˚Astr¨ om, K.J. (1965). Optimal control of Markov processes with incomplete state informa- tion. J. Math. Anal. Appl. 10 pp. 174–205
1965
-
[4]
(1964) Mixed and behavior strategies in infinite exstensive games
Aumann, R.J. (1964) Mixed and behavior strategies in infinite exstensive games. Advances in Game Theory 52 pp. 627–650
1964
-
[5]
(1973) Optimal decision procedures for finite Markov chains
Bather, J. (1973) Optimal decision procedures for finite Markov chains. Part I: Examples. Adv. in Appl. Probab. 5 328–339
work page 1973
-
[6]
(1973) Optimal decision procedures for finite Markov chains
Bather, J. (1973) Optimal decision procedures for finite Markov chains. Part II: Commu- nicating systems. Adv. in Appl. Probab. 5 521–540
work page 1973
- [7]
-
[8]
(1996) Stochastic Optimal Control: The Discrete-Time Case
Bertsekas, D.P., Shreve, S.E. (1996) Stochastic Optimal Control: The Discrete-Time Case. Athena Scientific, Belmont, MA
work page 1996
Show all 84 references
-
[9]
(1996) Neuro-Dynamic Programming, Athena Scientific, Belmont, MA
Bertsekas, D.P., Tsitsiklis, J.N. (1996) Neuro-Dynamic Programming, Athena Scientific, Belmont, MA
1996
-
[10]
(1987) An expected average rward criterion
Bierth, K.-J. (1987) An expected average rward criterion. Stochastic Processes and Appli- cations 26 pp. 133–140
1987
-
[11]
(2014) Examples concerning Abel and Cesaro limits
Bishop, C.J., Feinberg, E.A., Zhang J. (2014) Examples concerning Abel and Cesaro limits. J. Math. Anal. Appl. 420, pp. 1654–1661
2014
-
[12]
(1962) Discrete dynamic programming
Blackwell, D. (1962) Discrete dynamic programming. Ann. Math. Statist. 33(2) pp. 719– 726
1962
-
[13]
(1965) Discounted dynamic programming
Blackwell, D. (1965) Discounted dynamic programming. Ann. Math. Statist. 36 pp. 226– 235
1965
-
[14]
(1967) Positive dynamic programming
Blackwell, D. (1967) Positive dynamic programming. In Proceedings of the fifth Berkeley symposium on mathematical statististics and probability (Berkeley, CA, 21 June-18 July 1965), vol. I: Theory of statistics. Edited by L.M. Le Cam and J. Neyman. University of California Pre...
1967
-
[15]
(1974) The optimal reward operator in dy- namic programming, Ann
Blackwell, D., Freedman, D., and Orkin, M. (1974) The optimal reward operator in dy- namic programming, Ann. Probability 2 pp. 926–941. 14
1974
-
[16]
(1991) A counterexample on the optimality equation in Markov deci- sion chains with the average cost criterion
Cavazos-Cadena, R. (1991) A counterexample on the optimality equation in Markov deci- sion chains with the average cost criterion. Systems & Control Lett. 16(5) 387–392
1991
-
[17]
(1975) A controlled finite Markov chain with an arbitrary set of decisions
Chitashvili, R.Y. (1975) A controlled finite Markov chain with an arbitrary set of decisions. Theor. Probability Appl. 20(4) 839–847
1975
-
[18]
(1968) Multichain Markov renewal programs
Denardo, E.V., Fox, B.L. (1968) Multichain Markov renewal programs. SIAM J. Appl. Math. 15(3) pp. 468-487
1968
-
[19]
(1962) On sequential decisions and Markov chains
Derman, C. (1962) On sequential decisions and Markov chains. Management Sci . 9(1) 16–24
1962
-
[20]
(1965) Controlled random sequences
Dynkin, E.B. (1965) Controlled random sequences. Theory Probab. Appl. 10 pp. 1–14
1965
-
[21]
(1979) Controlled Markov Processes
Dynkin, E.B., Yushkevich, A.A. (1979) Controlled Markov Processes. Springer-Verlag, New York
1979
-
[22]
(1980) SuccessivA survey of asymptotic value iteration for undiscounted Markov decision problems
Federgruen, A., Schweitzer, P.J. (1980) SuccessivA survey of asymptotic value iteration for undiscounted Markov decision problems. In: R. Hartley, L.C. Thomas, D.J. White (eds.) Resent Development in Markov Decision Processes, Academic Press, New York, NY, pp. 73–109
1980
-
[23]
(1975) On controlled finite state Markov processes with compact control sets, Theor
Feinberg, E.A. (1975) On controlled finite state Markov processes with compact control sets, Theor. Probab. Appl. 20, pp. 856–862
1975
-
[24]
(1978) The existence of a stationary ϵ-optimal policy for a finite Markov chain Theor
Feinberg, E.A. (1978) The existence of a stationary ϵ-optimal policy for a finite Markov chain Theor. Probab. Appl. 23, pp. 297–313
1978
-
[25]
(1980) An ϵ-optimal control of a finite Markov chain
Feinberg, E.A. (1980) An ϵ-optimal control of a finite Markov chain. Theor. Probab. Appl. 25(1) 70–81
1980
-
[26]
(1986) Sufficient classes of strategies in discrete dynamic programming
Feinberg, E.A. (1986) Sufficient classes of strategies in discrete dynamic programming. I: Decomposition of randomized strategies and imbedded models. Theor. Probab. Appl. 31 pp. 478–493
1986
-
[27]
(1987) Sufficient classes of strategies in discrete dynamic programming
Feinberg, E.A. (1987) Sufficient classes of strategies in discrete dynamic programming. II: Locally stationary strategies. Theor. Probab. Appl. 32 pp. 658–668
1987
-
[28]
(2020) Complexity bounds for approximately solving discounted MDPs by value iterations, Operations Research Letters 48(5): 545–548
Feinberg, E.A., He, G. (2020) Complexity bounds for approximately solving discounted MDPs by value iterations, Operations Research Letters 48(5): 545–548
2020
-
[29]
(2014) The value iteration algorithm is not strongly polynomial for discounted dynamic programming, Oper
Feinberg, E.A., Huang, J. (2014) The value iteration algorithm is not strongly polynomial for discounted dynamic programming, Oper. Res. Lett. 42: 130–131
2014
-
[30]
(2025) Continuity of filters for discrete-time control problems defined by explicit equations
Feinberg, E.A., Ishizawa, S., Kasyanov, P.O., Kraemer, D.N. (2025) Continuity of filters for discrete-time control problems defined by explicit equations. SIAM J. Control Optim. 63(3): 1709-1735
2025
-
[31]
(2024) Sufficient conditions for solving statistical filtering problems by dynamic programming
Feinberg, E.A., Ishizawa, S., Kasyanov, P.O., Kraemer, D.N. (2024) Sufficient conditions for solving statistical filtering problems by dynamic programming. Proceedings of 63rd IEEE Conference on Decision and Control, December 16-19, 2024, Milan, Italy , pp. 4052– 4057
2024
-
[32]
(2021) MDPs with setwise continuous transition probabil- ities
Feinberg, E.A., Kasyanov, P.O. (2021) MDPs with setwise continuous transition probabil- ities. Oper. Res. Lett. 49(5), pp. 734–740. 15
2021
-
[33]
Voorneveld, M
Feinberg, E.A., Kasyanov P.O., and M. Voorneveld, M. (2013) Berge’s maximum theorem for noncompact image sets. J. Math. Anal. Appl. 397(1):255–259
2013
-
[34]
(2012) Average-cost Markov decision processes with weakly continuous transition probabilities, Math
Feinberg, E.A., Kasyanov, P.O., Zadoianchuk, N.V. (2012) Average-cost Markov decision processes with weakly continuous transition probabilities, Math. Oper. Res. 37, pp. 591- 607
2012
-
[35]
(2013) Berge’s theorem for noncompact image sets
Feinberg, E.A., Kasyanov P.O., Zadoianchuk, N.V. (2013) Berge’s theorem for noncompact image sets. J. Math. Anal. Appl. 397, pp. 255–259
2013
-
[36]
(2016) Partially observable total-cost Markov decision processes with weakly continuous transition probabilities
Feinberg, E.A., Kasyanov P.O., Zgurovsky, M.Z. (2016) Partially observable total-cost Markov decision processes with weakly continuous transition probabilities. Math. Oper. Res. 41(2): 656–681
2016
-
[37]
(2022) Markov decision processes with incomplete information and semi-uniform Feller transition probabilities
Feinberg, E.A., Kasyanov P.O., Zgurovsky, M.Z. (2022) Markov decision processes with incomplete information and semi-uniform Feller transition probabilities. SIAM J. Control Optim. 60(4):2488–2513
2022
-
[38]
(2023) Semi-uniform Feller stochastic kernels
Feinberg, E.A., Kasyanov P.O., Zgurovsky, M.Z. (2023) Semi-uniform Feller stochastic kernels. J. Theor. Probab. 36, pp. 2262–2283
2023
-
[39]
(2008) On polynomial classification problems for Markov decision processes
Feinberg, E.A., Yang, F. (2008) On polynomial classification problems for Markov decision processes. Oper. Res. Lett. 36 pp. 527–530. q
2008
-
[40]
(2014) On solutions of Kolmogorov’s equa- tions for jump Markov processes
Feinberg, E.A., Mandava, M., Shiryaev, A.N. (2014) On solutions of Kolmogorov’s equa- tions for jump Markov processes. J. Math. Anal. Appl. 411, pp. 261–270
2014
-
[41]
(2022) Kolmogorov’s equations for jump Markov processes with unbounded jump rates
Feinberg, E.A., Mandava, M., Shiryaev, A.N. (2022) Kolmogorov’s equations for jump Markov processes with unbounded jump rates. Ann. Oper. Res. 317(2), pp. 587–604
2022
-
[42]
(2022) Sufficiency of Markov policies for continuous-time jump Markov decision processes
Feinberg, E.A., Mandava, M., Shiryaev, A.N. (2022) Sufficiency of Markov policies for continuous-time jump Markov decision processes. Math. Oper. Res. 47(2), pp. 1266–1286
2022
-
[43]
(2022) Kolmogorov’s equations for jump Markov processes and their applications to control problems
Feinberg, E.A., Shiryaev, A.N. (2022) Kolmogorov’s equations for jump Markov processes and their applications to control problems. Theory Probab. Appl. 66(4), pp. 582–600, 2022
2022
-
[44]
(2024) On forward and backward Kolmogorov equations for pure jump Markov processes and their generalizations, Theor
Feinberg, E.A., Shiryaev, A.N. (2024) On forward and backward Kolmogorov equations for pure jump Markov processes and their generalizations, Theor. Probab. Appl. 68(4), pp. 643–656
2024
-
[45]
(1983) Stationary and Markov policies in countable state dy- namic programming
Feinberg, E.A., Sonin, I.M. (1983) Stationary and Markov policies in countable state dy- namic programming. Lecture Notes in Math. 1021 pp. 111–129
1983
-
[46]
(1940) On the integro-differential equations of purely discontinuous Markoff processes, Trans
Feller, W. (1940) On the integro-differential equations of purely discontinuous Markoff processes, Trans. Amer. Math. Soc., 48, pp. 488–515; Errata, Trans. Amer. Math. Soc., 58 (1945), p. 474
1940
-
[47]
(1974) The optimal reward operator in special classes of dynamic program- ming problems, Ann
Freedman, D. (1974) The optimal reward operator in special classes of dynamic program- ming problems, Ann. Probability 2, pp. 942-949
1974
-
[48]
(1979) Controlled Stochastic Processes.Springer, New York, NY
Gikhman, I.I., Skorohod, A.V. (1979) Controlled Stochastic Processes.Springer, New York, NY
1979
-
[49]
(2009)Continuous-Time Markov Decision Processes: The- ory and Applications
Guo, X., Hern´ andez-Lerma, O. (2009)Continuous-Time Markov Decision Processes: The- ory and Applications. Springer-Verlag, Berlin, 2009. 16
2009
-
[50]
(1989) Adaptive Markov Control Processes, Springer-Verlag, New York
Hern´ andez-Lerma, O. (1989) Adaptive Markov Control Processes, Springer-Verlag, New York
1989
-
[51]
Hern´ andez-Lerma, O. 1991. Averege optimality in dynamic programming on Borel spaces - Unbounded costs and controls. Systems & Control Lett. 17(3) pp. 237–242
1991
-
[52]
(1996)Discrete-Time Markov Control Processes: Ba- sic Optimality Criteria
Hern´ andez-Lerma, O., Lassere, J.B. (1996)Discrete-Time Markov Control Processes: Ba- sic Optimality Criteria . Springer, New York
1996
-
[53]
(1960) Dynamic Programming and Markov Processes
Howard, R.A. (1960) Dynamic Programming and Markov Processes. John Wiley & Sons, New York, NY
1960
-
[54]
(1983) Linear Programming and Finite Markovian Control Problems
Kallenberg, L.C.M. (1983) Linear Programming and Finite Markovian Control Problems. Mathematical Centre Tract 148, Mathematical Centre, Amsterdam
1983
-
[55]
(2002) Finite state and action MDPs
Kallenberg, L.C.M. (2002) Finite state and action MDPs. E.A. Feinberg, A. Shwartz, eds. Handbook of Markov Decision Processes. Methods and Applications. Kluwer, Boston, pp. 21–87
2002
-
[56]
(2019) Weak Feller property of non-linear filters.Systems & Control Letters 134, 104512
Kara, A.D., Saldi, N., Y¨ uksel, S. (2019) Weak Feller property of non-linear filters.Systems & Control Letters 134, 104512
2019
-
[57]
(1995) Controlled Queueing Systems
Kitaev, M.Yu., Rykov, V.V. (1995) Controlled Queueing Systems. CRC Press, Boca Raton
1995
-
[58]
(2014) A bound for the number of different basic solutions gen- erated by the simplex method, Math
Kitahara, T., Mizuno, S. (2014) A bound for the number of different basic solutions gen- erated by the simplex method, Math. Program. 137 pp. 579–586
2014
-
[59]
(1931) ¨Uber die analytischen Methoden in der Wahrscheinlichkeitsrech- nung, Math
Kolmogoroff A. (1931) ¨Uber die analytischen Methoden in der Wahrscheinlichkeitsrech- nung, Math. Ann. , 104 ), pp. 415–458; English transl.: A.N. Kolmogorov, On analytical methods in probability theory, in Selected Works of A.N. Kolmogorov, Vol. II: Probability Theory and Mat...
1931
-
[60]
Luque-V´ asquez, F., Hern´ andez-Lerma. O. (1995) A counterexample on the semicontinuity of minima. Proc. Amer. Math. Soc. (10) 3175–3176
1995
-
[61]
Zhang, Y
Piunovskiy, A.B. Zhang, Y. (2020) Continuous-Time Markov Decision Processes. Springer Nature, Switzerland
2020
-
[62]
Post, I. Ye, Y. (2015) The simplex method is strongly polynomial for deterministic Markov decision processes, Math. Oper. Res. 40 pp. 859–868
2015
-
[63]
(1994) Markov Decision Processes: Discrete Stochastic Dynamic Pro- gramming
Puterman, M.L. (1994) Markov Decision Processes: Discrete Stochastic Dynamic Pro- gramming. John Wiley & Sons, New York, NY
1994
-
[64]
(1974) Incomplete information in Markovian decision models
Rhenius, D. (1974) Incomplete information in Markovian decision models. Ann. Statist. 2(6) pp. 1327–1334
1974
-
[65]
Ross, S. M. (1968) Non-discounted denumerable Markovian decision model. Ann. Math. Statist. 39(2) 412–424
1968
-
[66]
Ross, S. M. (1968) Arbitrary state Markovian decision processes. Ann. Math. Statist. 39(6) 2118–2122
1968
-
[67]
(1994) Approximations of Discrete Time Partially Ob- served Control Problems, Applied Mathematics Monographs CNR, Giardini Editori, Pisa
Runggaldier, W.J., Stettner, L. (1994) Approximations of Discrete Time Partially Ob- served Control Problems, Applied Mathematics Monographs CNR, Giardini Editori, Pisa. 17
1994
-
[68]
(1975) Conditions for optimality and for the limit of n-stage optimal policies to be optimal
Sch¨ al, M. (1975) Conditions for optimality and for the limit of n-stage optimal policies to be optimal. Z. Wahrscheinlichkeitstheor. Verw. Geb. 32 pp. 179–196
1975
-
[69]
(1993) Average optimality in dynamic programming with general state space
Sch¨ al, M. (1993) Average optimality in dynamic programming with general state space. Math. Oper. Res . 18(1) 163–172
1993
-
[70]
(2016) Improved and generalized upper bounds on the complexity of policy iteration
Scherrer, B. (2016) Improved and generalized upper bounds on the complexity of policy iteration. Math. Oper. Res. 41 pp. 758–774
2016
-
[71]
(1999) Stochastic Dynamic Programming and the Control of Queueing Sys- tems
Sennott, L.I. (1999) Stochastic Dynamic Programming and the Control of Queueing Sys- tems. John Wiley and Sons, New York
1999
-
[72]
Sennott, L.I. 2002. Average reward optimization theory for denumerable state spaces. E.A. Feinberg, A. Shwartz, eds. Handbook of Markov Decision Processes. Methods and Applications. Kluwer, Boston, pp. 153–172
2002
-
[73]
(1963) Stochastic games, Proc
Shapley, L.S. (1963) Stochastic games, Proc. Natl. Acad. USA 39 pp. 1095–1100
1963
-
[74]
Shiryaev, A.N. On the theory of decision functions and control by an observation process with incomplete data, Transactions of the Third Prague Conference on Information The- ory, Statistical Decision Functions, Random Processes (Liblice, 1962), 1964, pp. 657-681 (in Russian);...
1966
-
[75]
Shiryaev, A.N. Some new results in the theory of controlled random processes.Transactions of the Fourth Prague Conference on Information Theory, Statistical Decision Functions, Random Processes (Prague, 1965), 1967, pp. 131-201 (in Russian); Engl. transl. in Select. Transl. Ma...
1969
-
[76]
(2018) A general reinforcement learning algorithm that masters chess, shogi, and Go through self- play, Science 3662(6419) pp
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., Hassabis D. (2018) A general reinforcement learning algorithm that masters chess, shogi, and Go through self- play, Scie...
2018
-
[77]
(1966) Positive Dynamic Programming
Strauch, R. (1966) Positive Dynamic Programming. Ann. Math. Statist. 37 pp. 871–890
1966
-
[78]
(2018) Reinforcement Learning: An Introduction (2nd ed.)
Sutton, R.S., Barto, A.G. (2018) Reinforcement Learning: An Introduction (2nd ed.). The MIT Press, Cambridge, MA
2018
-
[79]
Taylor, III, H. M.. 1965. Markovian sequential replacement processes. Ann. Math. Statist. 36(6) 1677–1694
1965
-
[80]
(1990) Solving h-horizon, stationary Markov decision problems in time propor- tional to log(h), Oper
Tseng, P. (1990) Solving h-horizon, stationary Markov decision problems in time propor- tional to log(h), Oper. Res. Lett. 9 pp. 287–297
1990
-
[81]
(2007) NP -hardness of checking the unichain condition in average cost MDPs
Tsitsiklis, J.N. (2007) NP -hardness of checking the unichain condition in average cost MDPs. Oper. Res. Lett. 35 pp. 319–323
2007
-
[82]
On controls leading to optimal stationary regimes, Proceed- ings of the Steklov Institute of Mathematics, 71 (1964), pp
Viskov, O.V., Shiryaev, A.N. On controls leading to optimal stationary regimes, Proceed- ings of the Steklov Institute of Mathematics, 71 (1964), pp. 35-45 (In Russian); English translation: Report Number FTD-HT-67-69, National Technical Information Service, U.S. Department of...
1964
-
[83]
(2011) The simplex and policy-iteration methods are strongly polynomial for the Markov decision problem with a fixed discount rate, Math
Ye, Y. (2011) The simplex and policy-iteration methods are strongly polynomial for the Markov decision problem with a fixed discount rate, Math. Oper. Res. 36 pp. 593–603
2011
-
[84]
(1976) Reduction of a controlled Markov model with incomplete data to a problem with complete information in the case of Borel state and control spaces
Yushkevich, A.A. (1976) Reduction of a controlled Markov model with incomplete data to a problem with complete information in the case of Borel state and control spaces. Theory Probab. Appl. 21 pp. 153–158. 18
1976
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.