madupite is a distributed, customizable MDP solver based on inexact policy iteration that scales to millions of states on HPC clusters.
Semismooth Newton Methods for Risk-Averse Markov Decision Processes
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Inspired by semismooth Newton methods, we propose a general framework for designing solution methods with convergence guarantees for risk-averse Markov decision processes. Our approach accommodates a wide variety of risk measures by leveraging the assumption of Markovian coherent risk measures. To demonstrate the versatility and effectiveness of this framework, we design three distinct solution methods, each with proven convergence guarantees and competitive empirical performance. Validation results on benchmark problems demonstrate the competitive performance of our methods. Furthermore, we establish that risk-averse policy iteration can be interpreted as an instance of semismooth Newton's method. This insight explains its superior convergence properties compared to risk-averse value iteration. The core contribution of our work, however, lies in developing an algorithmic framework inspired by semismooth Newton methods, rather than evaluating specific risk measures or advocating for risk-averse approaches over risk-neutral ones in particular applications.
citation-role summary
citation-polarity summary
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Inside madupite: Technical Design and Performance
madupite is a distributed, customizable MDP solver based on inexact policy iteration that scales to millions of states on HPC clusters.