REVIEW 3 cited by
Mean-Variance Optimization in Markov Decision Processes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We consider finite horizon Markov decision processes under performance measures that involve both the mean and the variance of the cumulative reward. We show that either randomized or history-based policies can improve performance. We prove that the complexity of computing a policy that maximizes the mean reward under a variance constraint is NP-hard for some cases, and strongly NP-hard for others. We finally offer pseudopolynomial exact and approximation algorithms.
Forward citations
Cited by 3 Pith papers
-
Consistent Variance Estimation for Q-Function Estimators in Finite-Horizon MDP Tree Search
The variance of MCTS Q-estimators decomposes into reward, transition, and successor-value uncertainty; a new recursive estimator corrects the asymptotically biased i.i.d. sample-variance formula.
-
A Bayesian Composite Risk Approach for Stochastic Optimal Control and Markov Decision Processes
The paper proposes Bayesian composite risk stochastic control and MDP models with belief-dependent policies, and proves dynamic programming and asymptotic convergence results.
-
Practical Risk Measures in Reinforcement Learning
An actor-critic algorithm with a Monte Carlo risk critic is proposed for optimizing reinforcement learning policies under arbitrary, possibly non-coherent risk measures, with a risk function fitted from simulated data.
Discussion (0). Continue with ORCID to comment.