The paper proposes Bayesian composite risk stochastic control and MDP models with belief-dependent policies, and proves dynamic programming and asymptotic convergence results.
Mean-Variance Optimization in Markov Decision Processes
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
abstract
We consider finite horizon Markov decision processes under performance measures that involve both the mean and the variance of the cumulative reward. We show that either randomized or history-based policies can improve performance. We prove that the complexity of computing a policy that maximizes the mean reward under a variance constraint is NP-hard for some cases, and strongly NP-hard for others. We finally offer pseudopolynomial exact and approximation algorithms.
citation-role summary
background 1
citation-polarity summary
fields
math.OC 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
A Bayesian Composite Risk Approach for Stochastic Optimal Control and Markov Decision Processes
The paper proposes Bayesian composite risk stochastic control and MDP models with belief-dependent policies, and proves dynamic programming and asymptotic convergence results.