Pith. sign in

REVIEW 1 major objections 6 minor 84 references

On Pioneering Works of Albert Shiryaev on Markov Decision Processes and Some Later Developments

T0 review · 1 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This survey argues that three papers by Albert Shiryaev from the 1960s introduced the belief-state reduction for partially observable control and proved the existence of optimal stationary policies for average-cost Markov decision…

desk verdict A solid historical survey of Shiryaev's MDP work, but Theorem 6.5(d) has a genuine proof gap that needs to be fixed before the paper can stand. read the letter →

arxiv 2506.04896 v1 pith:MVAGTOO3 submitted 2025-06-05 math.PR math.OC

classification math.PRmath.OC MSC 90C4093E2060J7590C39
keywords MarkovdecisionprocessespartiallyobservableMDPsbeliefstatesaverage-costoptimalityweakcontinuitystochasticfilteringKolmogorovequationsjump
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that the mathematical core of today's partially observable reinforcement learning was present in three papers published by Albert Shiryaev in the 1960s. The papers are said to formulate the reduction of a control problem with hidden states to an ordinary Markov decision process whose states are the observer's posterior beliefs, and to prove that finite average-cost problems have optimal non-randomized stationary policies. The survey then reviews how these ideas were extended to infinite state and action spaces, including conditions under which belief-state transitions are weakly continuous and value iteration converges, and it connects the same line of work to recent solutions of Kolmogorov's equations for jump Markov processes and to continuous-time control. If the historical attributions are correct, the paper shows that the conceptual backbone of modern reinforcement learning was in place decades before the algorithms that made it prominent.

What carries the argument

The central objects are the belief MDP and the canonical average-cost equations. A belief MDP is a fully observed Markov decision process whose state is the posterior distribution of the hidden state; the survey traces this reduction to Shiryaev's papers and uses it to transfer optimality results from complete-observation theory to partially observable problems. For average-cost problems, the canonical equations $w = P^\phi w$ and $w + u = T^\phi u$ characterize optimal deterministic policies, and later sections extend their validity to Borel state spaces under Assumptions (W*) and (S*), which use K-inf-compactness to drop compactness of action sets. A further mechanism is the semi-uniform Feller property for transition kernels, which is equivalent to weak continuity of the induced belief-state transition, and the Diffeomorphic Condition, which gives checkable sufficient conditions for total-variation continuity of filters defined by explicit state and observation equations.

What would settle it

Reading the original texts of Viskov and Shiryaev [82] and Shiryaev [74, 75] to check whether they state and prove the belief-state reduction and the existence of a deterministic stationary average-cost optimal policy; if, for example, [82] proves only epsilon-optimality or assumes a unichain structure, the survey's attribution of the general theorem would be wrong. Alternatively, a finite-state/action average-cost MDP with infinite one-step costs that admits no deterministic stationary optimal policy would directly contradict the theorem as described.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that Shiryaev's 1962–1967 work already contained the two structural ideas that later became the basis of partially observable control: control with incomplete observations can be reduced to a fully observed MDP on the space of posterior distributions, and average-cost MDPs with finite state and action sets admit an optimal deterministic stationary policy. The survey attributes the belief-space reduction to [74, 75] and the average-cost optimality proof to [82], noting that [12] and [19] reached the same conclusion independently. It then presents later results that make these ideas effective on infinite state and action spaces: with K-inf-compact cost functions and weak or setwise continuity assumptions, optimality equations hold, optimal policies exist, and value iteration converges for belief MDPs and for filtering problems with additive or multiplicative noise. The later sections report the resolution of Feller's problem on forward Kolmogorov equations for jump Markov processes and its use to prove that Markov policies suffice for continuous-time jump MDPs.

Load-bearing premise

The survey's narrative stands on the claim that the three 1960s papers, which are not reproduced in the survey, actually contain the belief-state reduction and the average-cost optimality proof attributed to them; if that reading is wrong, the historical claims would need to be revised.

Editorial extensions

If this is right

  • If the belief-state reduction is valid, every POMDP can be solved as a fully observed MDP over probability distributions, so dynamic programming, value iteration, and linear programming methods apply directly to partially observable problems.
  • The historical priority claim implies that the belief-state reduction predates the modern POMDP and reinforcement-learning literature by decades, not the reverse.
  • Under the weak-continuity and K-inf-compactness conditions described in the survey, optimal policies exist and value iteration converges for belief MDPs even when action sets are non-compact, a setting that covers inventory control and many filtering models.
  • For filtering problems written with explicit state and observation equations, the Diffeomorphic Condition guarantees weak continuity of the filter, so stochastic control and stochastic filtering can be solved by the same dynamic-programming machinery.
  • The resolution of Feller's forward-equation problem and the Markov-policy sufficiency result for continuous-time jump MDPs extend these guarantees from discrete time to continuous-time controlled jump processes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Not in the paper: the survey does not compare Shiryaev's belief-state reduction in detail with the contemporaneous formulations it cites; a careful comparison could clarify how much of the modern POMDP framework is genuinely new.
  • Not in the paper: the Diffeomorphic Condition is stated for Euclidean noise with absolutely continuous distributions; testing whether total-variation continuity survives for heavier-tailed noise or for noise on manifolds would be a natural extension.
  • Not in the paper: the semi-uniform Feller property could be used as a practical certificate for POMDP solvers, since it guarantees weak continuity of the belief transition and therefore convergence of value iteration; this check is not discussed in the survey.
  • Not in the paper: if the historical attributions are accepted, the survey implies that the belief-state idea should be presented in reinforcement-learning textbooks as a classical control-theoretic tool rather than a modern invention.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 6 minor

Summary. The paper is a survey of three historically important papers by A.N. Shiryaev (one with O.V. Viskov) on Markov decision processes and control with incomplete information. It recaps the existence of deterministic stationary optimal policies for finite average-cost MDPs, the reduction of POMDPs to belief MDPs, and later results on discounted and average-cost MDPs with Borel state/action spaces, weak continuity of belief transition kernels, and Kolmogorov equations for jump Markov processes. The last part of Section 6 states a new theorem (Theorem 6.5) giving sufficient conditions for weak continuity of the belief-MDP kernel in filtering problems defined by explicit equations.

Significance. If the historical attributions are correct, the survey performs a useful service by placing Shiryaev's 1962–1967 papers in the context of modern POMDP and reinforcement learning theory. The paper is clearly written, includes precise statements of many later results, and has a comprehensive bibliography. The main new mathematical claim (Theorem 6.5) is, however, not fully supported: part (d) is stated without a valid proof, and the Conclusion uses this theorem to advertise 'broad sufficient conditions' for filtering problems. Once that gap is repaired, the paper would be a solid contribution; in its present form it requires revision.

major comments (1)
  1. [Section 6, Theorem 6.5(d)] The proof of Theorem 6.5 says that part (d) follows from Theorems 5.1 and 6.4, but neither sufficient condition of Theorem 5.1 applies. Condition (ii) of Theorem 5.1 requires the transition kernel T to be continuous in total variation, whereas (d) only assumes F continuous in distribution (weak continuity of T). Condition (i) requires Q to be continuous in total variation jointly in (a,x); since g is only assumed measurable in a, Theorem 6.4 cannot be invoked for the pair (a,x), and continuity of g in x alone does not imply total-variation continuity of Q in a. Thus the weak continuity of p̄ is unproven under (d). This is load-bearing because the Conclusion relies on Theorem 6.5 to claim broad sufficient conditions for filtering problems; the statement needs a corrected proof (e.g., adding joint continuity of g or a different argument) or condition (d) should be withdrawn.
minor comments (6)
  1. [Section 4.3] Two different assumptions are both labeled 'Assumption (B)': the first (from Schäl [69]) uses sup_alpha u_alpha(x) < ∞, the second (from [34]) uses lim inf_alpha u_alpha(x) < ∞. Please rename them (e.g., (B1) and (B2)) to avoid ambiguity.
  2. [Section 5, Eq. (14)] The definition of semi-uniform Feller appears to have the product spaces reversed: as written, Ψ is a kernel from S3 to S1×S2 (since the conditioning argument is s3 and the output is a measure on S1 with a set B⊂S2), but the preceding text says Ψ is a transition probability from S1×S2 to S3. Please correct the direction and the notation.
  3. [Section 6, Theorem 6.4] In the statement of Theorem 6.4, the codomain of φ is written as 'R' but given S1 = R^n it should be 'R^n'; also 'satisfy' should be 'satisfies'.
  4. [Section 2] The definition of the history sets is garbled: 'Ht := X × (A × X),,' should be H_t = X × (A × X)^t (or an equivalent description), and the sentence lacks the verb 'let'.
  5. [Section 6, Theorem 6.5(d)] The hypothesis says g is 'measurable' and 'continuous in variable x_{t+1}'; please state explicitly that g is jointly Borel measurable and specify whether continuity in x is uniform over compacts or pointwise, because this affects the applicability of Theorem 6.4.
  6. [Throughout] There are several spelling/typographical errors, e.g., 'expexted' in Section 6, 'Cezaro' for 'Cesàro' in Section 2, and 'Corollart' in Corollary 6.3. Please proofread carefully.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the historical attributions are anchored in primary sources and the later-development theorems are cited with independent proofs; the only flagged issue is a non-circular proof gap in Theorem 6.5(d).

full rationale

This is an expository survey, not a paper with fitted parameters or a derivation chain whose conclusions are encoded in its assumptions. The central historical claims—that Viskov and Shiryaev [82] proved existence of deterministic optimal policies for average-cost MDPs and that Shiryaev [74, 75] formulated the reduction of partially observable problems to belief-state MDPs—are grounded in the cited primary sources and corroborated by independent classical references such as Blackwell [12], Derman [19], Aoki [1], Åström [3], and Dynkin [20]. These claims do not reduce to the author's own work. The many self-citations (e.g., [30, 32, 34, 36, 37, 40–44]) are used to report later developments; they point to published theorems with stated assumptions and proofs, which counts as independent support rather than a self-referential derivation chain. No fitted quantity is renamed as a prediction, no known result is merely re-coined, and no uniqueness or ansatz is imported solely through self-citation. One correctness concern emerged, but it is not circularity: the proof of Theorem 6.5(d) says it follows from Theorems 5.1 and 6.4, yet the hypotheses of (d)—g only measurable in a and continuous in x, with T only continuous in distribution—do not satisfy either sufficient route of Theorem 5.1, since route (i) requires Q continuous in total variation and route (ii) requires T continuous in total variation. That is a potential proof gap in the paper's own contribution, not a case where the conclusion is equivalent to the input by construction. Therefore the circularity score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No new parameters or entities are introduced; the paper is a survey. The free-parameter count is zero because the new theorem only assumes regularity conditions (e.g., the Diffeomorphic Condition) imported from [30].

assumptions (2)
  • domain assumption The cited papers [74, 75, 82] contain the results and ideas attributed to them
    The survey's historical and technical claims about Shiryaev's works are not reproduced or independently verified in this text.
  • standard math Standard measure-theoretic and analytic results (Ionescu Tulcea theorem, Banach fixed point theorem, Berge's theorem, Hardy-Littlewood Tauberian theorem, Aumann's lemma) are used without proof
    These background results are invoked in Sections 2 to 6 as standard tools in MDP theory.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On Pioneering Works of Albert Shiryaev on Markov Decision Processes and Some Later Developments." pith.science (2026). https://pith.science/paper/MVAGTOO3

@misc{pith2026250604896,
  author       = {Pith},
  title        = {Pith review of: On Pioneering Works of Albert Shiryaev on Markov Decision Processes and Some Later Developments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MVAGTOO3}},
  note         = {Machine review of arXiv:2506.04896}
}
read the original abstract

This article is dedicated to three fundamental papers on Markov Decision Processes and on control with incomplete observations published by Albert Shiryaev approximately sixty years ago. One of these papers was coauthored with O.V. Viskov. We discuss some of the results and some of many rich ideas presented in these papers and survey some later developments. At the end we mention some recent studies of Albert Shiryaev on Kolmogorov's equations for jump Markov processes and on control of continuous-time jump Markov processes.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

84 extracted references · 80 canonical work pages

  1. [1]

    (1965) Optimal control of partially observable Markovian systems

    Aoki, M. (1965) Optimal control of partially observable Markovian systems. J. Franklin Inst. 280 pp. 367–386

  2. [2]

    (1993) Discrete time controlled Markov processes with average cost criterion: a survey, SIAM J

    Arapostathis, A., Borkar, V.S., Fernandez-Gaucherand, E, Ghosh M.K., Marcus, S.I. (1993) Discrete time controlled Markov processes with average cost criterion: a survey, SIAM J. Control Optim. 31(2) 282–344

  3. [3]

    ˚Astr¨ om, K.J. (1965). Optimal control of Markov processes with incomplete state informa- tion. J. Math. Anal. Appl. 10 pp. 174–205

  4. [4]

    (1964) Mixed and behavior strategies in infinite exstensive games

    Aumann, R.J. (1964) Mixed and behavior strategies in infinite exstensive games. Advances in Game Theory 52 pp. 627–650

  5. [5]

    (1973) Optimal decision procedures for finite Markov chains

    Bather, J. (1973) Optimal decision procedures for finite Markov chains. Part I: Examples. Adv. in Appl. Probab. 5 328–339

  6. [6]

    (1973) Optimal decision procedures for finite Markov chains

    Bather, J. (1973) Optimal decision procedures for finite Markov chains. Part II: Commu- nicating systems. Adv. in Appl. Probab. 5 521–540

  7. [7]

    (1963) Topological Spaces

    Berge, E. (1963) Topological Spaces. Macmillan, New York

  8. [8]

    (1996) Stochastic Optimal Control: The Discrete-Time Case

    Bertsekas, D.P., Shreve, S.E. (1996) Stochastic Optimal Control: The Discrete-Time Case. Athena Scientific, Belmont, MA

Show all 84 references
  1. [9]

    (1996) Neuro-Dynamic Programming, Athena Scientific, Belmont, MA

    Bertsekas, D.P., Tsitsiklis, J.N. (1996) Neuro-Dynamic Programming, Athena Scientific, Belmont, MA

  2. [10]

    (1987) An expected average rward criterion

    Bierth, K.-J. (1987) An expected average rward criterion. Stochastic Processes and Appli- cations 26 pp. 133–140

  3. [11]

    (2014) Examples concerning Abel and Cesaro limits

    Bishop, C.J., Feinberg, E.A., Zhang J. (2014) Examples concerning Abel and Cesaro limits. J. Math. Anal. Appl. 420, pp. 1654–1661

  4. [12]

    (1962) Discrete dynamic programming

    Blackwell, D. (1962) Discrete dynamic programming. Ann. Math. Statist. 33(2) pp. 719– 726

  5. [13]

    (1965) Discounted dynamic programming

    Blackwell, D. (1965) Discounted dynamic programming. Ann. Math. Statist. 36 pp. 226– 235

  6. [14]

    (1967) Positive dynamic programming

    Blackwell, D. (1967) Positive dynamic programming. In Proceedings of the fifth Berkeley symposium on mathematical statististics and probability (Berkeley, CA, 21 June-18 July 1965), vol. I: Theory of statistics. Edited by L.M. Le Cam and J. Neyman. University of California Pre...

  7. [15]

    (1974) The optimal reward operator in dy- namic programming, Ann

    Blackwell, D., Freedman, D., and Orkin, M. (1974) The optimal reward operator in dy- namic programming, Ann. Probability 2 pp. 926–941. 14

  8. [16]

    (1991) A counterexample on the optimality equation in Markov deci- sion chains with the average cost criterion

    Cavazos-Cadena, R. (1991) A counterexample on the optimality equation in Markov deci- sion chains with the average cost criterion. Systems & Control Lett. 16(5) 387–392

  9. [17]

    (1975) A controlled finite Markov chain with an arbitrary set of decisions

    Chitashvili, R.Y. (1975) A controlled finite Markov chain with an arbitrary set of decisions. Theor. Probability Appl. 20(4) 839–847

  10. [18]

    (1968) Multichain Markov renewal programs

    Denardo, E.V., Fox, B.L. (1968) Multichain Markov renewal programs. SIAM J. Appl. Math. 15(3) pp. 468-487

  11. [19]

    (1962) On sequential decisions and Markov chains

    Derman, C. (1962) On sequential decisions and Markov chains. Management Sci . 9(1) 16–24

  12. [20]

    (1965) Controlled random sequences

    Dynkin, E.B. (1965) Controlled random sequences. Theory Probab. Appl. 10 pp. 1–14

  13. [21]

    (1979) Controlled Markov Processes

    Dynkin, E.B., Yushkevich, A.A. (1979) Controlled Markov Processes. Springer-Verlag, New York

  14. [22]

    (1980) SuccessivA survey of asymptotic value iteration for undiscounted Markov decision problems

    Federgruen, A., Schweitzer, P.J. (1980) SuccessivA survey of asymptotic value iteration for undiscounted Markov decision problems. In: R. Hartley, L.C. Thomas, D.J. White (eds.) Resent Development in Markov Decision Processes, Academic Press, New York, NY, pp. 73–109

  15. [23]

    (1975) On controlled finite state Markov processes with compact control sets, Theor

    Feinberg, E.A. (1975) On controlled finite state Markov processes with compact control sets, Theor. Probab. Appl. 20, pp. 856–862

  16. [24]

    (1978) The existence of a stationary ϵ-optimal policy for a finite Markov chain Theor

    Feinberg, E.A. (1978) The existence of a stationary ϵ-optimal policy for a finite Markov chain Theor. Probab. Appl. 23, pp. 297–313

  17. [25]

    (1980) An ϵ-optimal control of a finite Markov chain

    Feinberg, E.A. (1980) An ϵ-optimal control of a finite Markov chain. Theor. Probab. Appl. 25(1) 70–81

  18. [26]

    (1986) Sufficient classes of strategies in discrete dynamic programming

    Feinberg, E.A. (1986) Sufficient classes of strategies in discrete dynamic programming. I: Decomposition of randomized strategies and imbedded models. Theor. Probab. Appl. 31 pp. 478–493

  19. [27]

    (1987) Sufficient classes of strategies in discrete dynamic programming

    Feinberg, E.A. (1987) Sufficient classes of strategies in discrete dynamic programming. II: Locally stationary strategies. Theor. Probab. Appl. 32 pp. 658–668

  20. [28]

    (2020) Complexity bounds for approximately solving discounted MDPs by value iterations, Operations Research Letters 48(5): 545–548

    Feinberg, E.A., He, G. (2020) Complexity bounds for approximately solving discounted MDPs by value iterations, Operations Research Letters 48(5): 545–548

  21. [29]

    (2014) The value iteration algorithm is not strongly polynomial for discounted dynamic programming, Oper

    Feinberg, E.A., Huang, J. (2014) The value iteration algorithm is not strongly polynomial for discounted dynamic programming, Oper. Res. Lett. 42: 130–131

  22. [30]

    (2025) Continuity of filters for discrete-time control problems defined by explicit equations

    Feinberg, E.A., Ishizawa, S., Kasyanov, P.O., Kraemer, D.N. (2025) Continuity of filters for discrete-time control problems defined by explicit equations. SIAM J. Control Optim. 63(3): 1709-1735

  23. [31]

    (2024) Sufficient conditions for solving statistical filtering problems by dynamic programming

    Feinberg, E.A., Ishizawa, S., Kasyanov, P.O., Kraemer, D.N. (2024) Sufficient conditions for solving statistical filtering problems by dynamic programming. Proceedings of 63rd IEEE Conference on Decision and Control, December 16-19, 2024, Milan, Italy , pp. 4052– 4057

  24. [32]

    (2021) MDPs with setwise continuous transition probabil- ities

    Feinberg, E.A., Kasyanov, P.O. (2021) MDPs with setwise continuous transition probabil- ities. Oper. Res. Lett. 49(5), pp. 734–740. 15

  25. [33]

    Voorneveld, M

    Feinberg, E.A., Kasyanov P.O., and M. Voorneveld, M. (2013) Berge’s maximum theorem for noncompact image sets. J. Math. Anal. Appl. 397(1):255–259

  26. [34]

    (2012) Average-cost Markov decision processes with weakly continuous transition probabilities, Math

    Feinberg, E.A., Kasyanov, P.O., Zadoianchuk, N.V. (2012) Average-cost Markov decision processes with weakly continuous transition probabilities, Math. Oper. Res. 37, pp. 591- 607

  27. [35]

    (2013) Berge’s theorem for noncompact image sets

    Feinberg, E.A., Kasyanov P.O., Zadoianchuk, N.V. (2013) Berge’s theorem for noncompact image sets. J. Math. Anal. Appl. 397, pp. 255–259

  28. [36]

    (2016) Partially observable total-cost Markov decision processes with weakly continuous transition probabilities

    Feinberg, E.A., Kasyanov P.O., Zgurovsky, M.Z. (2016) Partially observable total-cost Markov decision processes with weakly continuous transition probabilities. Math. Oper. Res. 41(2): 656–681

  29. [37]

    (2022) Markov decision processes with incomplete information and semi-uniform Feller transition probabilities

    Feinberg, E.A., Kasyanov P.O., Zgurovsky, M.Z. (2022) Markov decision processes with incomplete information and semi-uniform Feller transition probabilities. SIAM J. Control Optim. 60(4):2488–2513

  30. [38]

    (2023) Semi-uniform Feller stochastic kernels

    Feinberg, E.A., Kasyanov P.O., Zgurovsky, M.Z. (2023) Semi-uniform Feller stochastic kernels. J. Theor. Probab. 36, pp. 2262–2283

  31. [39]

    (2008) On polynomial classification problems for Markov decision processes

    Feinberg, E.A., Yang, F. (2008) On polynomial classification problems for Markov decision processes. Oper. Res. Lett. 36 pp. 527–530. q

  32. [40]

    (2014) On solutions of Kolmogorov’s equa- tions for jump Markov processes

    Feinberg, E.A., Mandava, M., Shiryaev, A.N. (2014) On solutions of Kolmogorov’s equa- tions for jump Markov processes. J. Math. Anal. Appl. 411, pp. 261–270

  33. [41]

    (2022) Kolmogorov’s equations for jump Markov processes with unbounded jump rates

    Feinberg, E.A., Mandava, M., Shiryaev, A.N. (2022) Kolmogorov’s equations for jump Markov processes with unbounded jump rates. Ann. Oper. Res. 317(2), pp. 587–604

  34. [42]

    (2022) Sufficiency of Markov policies for continuous-time jump Markov decision processes

    Feinberg, E.A., Mandava, M., Shiryaev, A.N. (2022) Sufficiency of Markov policies for continuous-time jump Markov decision processes. Math. Oper. Res. 47(2), pp. 1266–1286

  35. [43]

    (2022) Kolmogorov’s equations for jump Markov processes and their applications to control problems

    Feinberg, E.A., Shiryaev, A.N. (2022) Kolmogorov’s equations for jump Markov processes and their applications to control problems. Theory Probab. Appl. 66(4), pp. 582–600, 2022

  36. [44]

    (2024) On forward and backward Kolmogorov equations for pure jump Markov processes and their generalizations, Theor

    Feinberg, E.A., Shiryaev, A.N. (2024) On forward and backward Kolmogorov equations for pure jump Markov processes and their generalizations, Theor. Probab. Appl. 68(4), pp. 643–656

  37. [45]

    (1983) Stationary and Markov policies in countable state dy- namic programming

    Feinberg, E.A., Sonin, I.M. (1983) Stationary and Markov policies in countable state dy- namic programming. Lecture Notes in Math. 1021 pp. 111–129

  38. [46]

    (1940) On the integro-differential equations of purely discontinuous Markoff processes, Trans

    Feller, W. (1940) On the integro-differential equations of purely discontinuous Markoff processes, Trans. Amer. Math. Soc., 48, pp. 488–515; Errata, Trans. Amer. Math. Soc., 58 (1945), p. 474

  39. [47]

    (1974) The optimal reward operator in special classes of dynamic program- ming problems, Ann

    Freedman, D. (1974) The optimal reward operator in special classes of dynamic program- ming problems, Ann. Probability 2, pp. 942-949

  40. [48]

    (1979) Controlled Stochastic Processes.Springer, New York, NY

    Gikhman, I.I., Skorohod, A.V. (1979) Controlled Stochastic Processes.Springer, New York, NY

  41. [49]

    (2009)Continuous-Time Markov Decision Processes: The- ory and Applications

    Guo, X., Hern´ andez-Lerma, O. (2009)Continuous-Time Markov Decision Processes: The- ory and Applications. Springer-Verlag, Berlin, 2009. 16

  42. [50]

    (1989) Adaptive Markov Control Processes, Springer-Verlag, New York

    Hern´ andez-Lerma, O. (1989) Adaptive Markov Control Processes, Springer-Verlag, New York

  43. [51]

    Hern´ andez-Lerma, O. 1991. Averege optimality in dynamic programming on Borel spaces - Unbounded costs and controls. Systems & Control Lett. 17(3) pp. 237–242

  44. [52]

    (1996)Discrete-Time Markov Control Processes: Ba- sic Optimality Criteria

    Hern´ andez-Lerma, O., Lassere, J.B. (1996)Discrete-Time Markov Control Processes: Ba- sic Optimality Criteria . Springer, New York

  45. [53]

    (1960) Dynamic Programming and Markov Processes

    Howard, R.A. (1960) Dynamic Programming and Markov Processes. John Wiley & Sons, New York, NY

  46. [54]

    (1983) Linear Programming and Finite Markovian Control Problems

    Kallenberg, L.C.M. (1983) Linear Programming and Finite Markovian Control Problems. Mathematical Centre Tract 148, Mathematical Centre, Amsterdam

  47. [55]

    (2002) Finite state and action MDPs

    Kallenberg, L.C.M. (2002) Finite state and action MDPs. E.A. Feinberg, A. Shwartz, eds. Handbook of Markov Decision Processes. Methods and Applications. Kluwer, Boston, pp. 21–87

  48. [56]

    (2019) Weak Feller property of non-linear filters.Systems & Control Letters 134, 104512

    Kara, A.D., Saldi, N., Y¨ uksel, S. (2019) Weak Feller property of non-linear filters.Systems & Control Letters 134, 104512

  49. [57]

    (1995) Controlled Queueing Systems

    Kitaev, M.Yu., Rykov, V.V. (1995) Controlled Queueing Systems. CRC Press, Boca Raton

  50. [58]

    (2014) A bound for the number of different basic solutions gen- erated by the simplex method, Math

    Kitahara, T., Mizuno, S. (2014) A bound for the number of different basic solutions gen- erated by the simplex method, Math. Program. 137 pp. 579–586

  51. [59]

    (1931) ¨Uber die analytischen Methoden in der Wahrscheinlichkeitsrech- nung, Math

    Kolmogoroff A. (1931) ¨Uber die analytischen Methoden in der Wahrscheinlichkeitsrech- nung, Math. Ann. , 104 ), pp. 415–458; English transl.: A.N. Kolmogorov, On analytical methods in probability theory, in Selected Works of A.N. Kolmogorov, Vol. II: Probability Theory and Mat...

  52. [60]

    Luque-V´ asquez, F., Hern´ andez-Lerma. O. (1995) A counterexample on the semicontinuity of minima. Proc. Amer. Math. Soc. (10) 3175–3176

  53. [61]

    Zhang, Y

    Piunovskiy, A.B. Zhang, Y. (2020) Continuous-Time Markov Decision Processes. Springer Nature, Switzerland

  54. [62]

    Post, I. Ye, Y. (2015) The simplex method is strongly polynomial for deterministic Markov decision processes, Math. Oper. Res. 40 pp. 859–868

  55. [63]

    (1994) Markov Decision Processes: Discrete Stochastic Dynamic Pro- gramming

    Puterman, M.L. (1994) Markov Decision Processes: Discrete Stochastic Dynamic Pro- gramming. John Wiley & Sons, New York, NY

  56. [64]

    (1974) Incomplete information in Markovian decision models

    Rhenius, D. (1974) Incomplete information in Markovian decision models. Ann. Statist. 2(6) pp. 1327–1334

  57. [65]

    Ross, S. M. (1968) Non-discounted denumerable Markovian decision model. Ann. Math. Statist. 39(2) 412–424

  58. [66]

    Ross, S. M. (1968) Arbitrary state Markovian decision processes. Ann. Math. Statist. 39(6) 2118–2122

  59. [67]

    (1994) Approximations of Discrete Time Partially Ob- served Control Problems, Applied Mathematics Monographs CNR, Giardini Editori, Pisa

    Runggaldier, W.J., Stettner, L. (1994) Approximations of Discrete Time Partially Ob- served Control Problems, Applied Mathematics Monographs CNR, Giardini Editori, Pisa. 17

  60. [68]

    (1975) Conditions for optimality and for the limit of n-stage optimal policies to be optimal

    Sch¨ al, M. (1975) Conditions for optimality and for the limit of n-stage optimal policies to be optimal. Z. Wahrscheinlichkeitstheor. Verw. Geb. 32 pp. 179–196

  61. [69]

    (1993) Average optimality in dynamic programming with general state space

    Sch¨ al, M. (1993) Average optimality in dynamic programming with general state space. Math. Oper. Res . 18(1) 163–172

  62. [70]

    (2016) Improved and generalized upper bounds on the complexity of policy iteration

    Scherrer, B. (2016) Improved and generalized upper bounds on the complexity of policy iteration. Math. Oper. Res. 41 pp. 758–774

  63. [71]

    (1999) Stochastic Dynamic Programming and the Control of Queueing Sys- tems

    Sennott, L.I. (1999) Stochastic Dynamic Programming and the Control of Queueing Sys- tems. John Wiley and Sons, New York

  64. [72]

    Sennott, L.I. 2002. Average reward optimization theory for denumerable state spaces. E.A. Feinberg, A. Shwartz, eds. Handbook of Markov Decision Processes. Methods and Applications. Kluwer, Boston, pp. 153–172

  65. [73]

    (1963) Stochastic games, Proc

    Shapley, L.S. (1963) Stochastic games, Proc. Natl. Acad. USA 39 pp. 1095–1100

  66. [74]

    Shiryaev, A.N. On the theory of decision functions and control by an observation process with incomplete data, Transactions of the Third Prague Conference on Information The- ory, Statistical Decision Functions, Random Processes (Liblice, 1962), 1964, pp. 657-681 (in Russian);...

  67. [75]

    Shiryaev, A.N. Some new results in the theory of controlled random processes.Transactions of the Fourth Prague Conference on Information Theory, Statistical Decision Functions, Random Processes (Prague, 1965), 1967, pp. 131-201 (in Russian); Engl. transl. in Select. Transl. Ma...

  68. [76]

    (2018) A general reinforcement learning algorithm that masters chess, shogi, and Go through self- play, Science 3662(6419) pp

    Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., Hassabis D. (2018) A general reinforcement learning algorithm that masters chess, shogi, and Go through self- play, Scie...

  69. [77]

    (1966) Positive Dynamic Programming

    Strauch, R. (1966) Positive Dynamic Programming. Ann. Math. Statist. 37 pp. 871–890

  70. [78]

    (2018) Reinforcement Learning: An Introduction (2nd ed.)

    Sutton, R.S., Barto, A.G. (2018) Reinforcement Learning: An Introduction (2nd ed.). The MIT Press, Cambridge, MA

  71. [79]

    Taylor, III, H. M.. 1965. Markovian sequential replacement processes. Ann. Math. Statist. 36(6) 1677–1694

  72. [80]

    (1990) Solving h-horizon, stationary Markov decision problems in time propor- tional to log(h), Oper

    Tseng, P. (1990) Solving h-horizon, stationary Markov decision problems in time propor- tional to log(h), Oper. Res. Lett. 9 pp. 287–297

  73. [81]

    (2007) NP -hardness of checking the unichain condition in average cost MDPs

    Tsitsiklis, J.N. (2007) NP -hardness of checking the unichain condition in average cost MDPs. Oper. Res. Lett. 35 pp. 319–323

  74. [82]

    On controls leading to optimal stationary regimes, Proceed- ings of the Steklov Institute of Mathematics, 71 (1964), pp

    Viskov, O.V., Shiryaev, A.N. On controls leading to optimal stationary regimes, Proceed- ings of the Steklov Institute of Mathematics, 71 (1964), pp. 35-45 (In Russian); English translation: Report Number FTD-HT-67-69, National Technical Information Service, U.S. Department of...

  75. [83]

    (2011) The simplex and policy-iteration methods are strongly polynomial for the Markov decision problem with a fixed discount rate, Math

    Ye, Y. (2011) The simplex and policy-iteration methods are strongly polynomial for the Markov decision problem with a fixed discount rate, Math. Oper. Res. 36 pp. 593–603

  76. [84]

    (1976) Reduction of a controlled Markov model with incomplete data to a problem with complete information in the case of Borel state and control spaces

    Yushkevich, A.A. (1976) Reduction of a controlled Markov model with incomplete data to a problem with complete information in the case of Borel state and control spaces. Theory Probab. Appl. 21 pp. 153–158. 18

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.