Pith. sign in

REVIEW 3 cited by

Mean-Variance Optimization in Markov Decision Processes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1104.5601 v1 pith:D6RO276Q submitted 2011-04-29 cs.LG cs.AI

classification cs.LGcs.AI
keywords decisionmarkovmeannp-hardperformanceprocessesrewardunder
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We consider finite horizon Markov decision processes under performance measures that involve both the mean and the variance of the cumulative reward. We show that either randomized or history-based policies can improve performance. We prove that the complexity of computing a policy that maximizes the mean reward under a variance constraint is NP-hard for some cases, and strongly NP-hard for others. We finally offer pseudopolynomial exact and approximation algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Consistent Variance Estimation for Q-Function Estimators in Finite-Horizon MDP Tree Search

    eess.SY 2026-07 conditional novelty 6.0 of 10

    The variance of MCTS Q-estimators decomposes into reward, transition, and successor-value uncertainty; a new recursive estimator corrects the asymptotically biased i.i.d. sample-variance formula.

  2. A Bayesian Composite Risk Approach for Stochastic Optimal Control and Markov Decision Processes

    math.OC 2024-12 conditional novelty 6.0 of 10

    The paper proposes Bayesian composite risk stochastic control and MDP models with belief-dependent policies, and proves dynamic programming and asymptotic convergence results.

  3. Practical Risk Measures in Reinforcement Learning

    cs.LG 2019-08 reject novelty 5.0 of 10

    An actor-critic algorithm with a Monte Carlo risk critic is proposed for optimizing reinforcement learning policies under arbitrary, possibly non-coherent risk measures, with a risk function fitted from simulated data.

Pith tools