In a controlled lifecycle benchmark, per-date backward neural policies beat single and two-regime networks on welfare and Bellman residuals, and a shape penalty fixes negative MPC violations.
Deep Reinforcement Learning and Mean-Variance Strategies for Responsible Portfolio Optimization
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Portfolio optimization involves determining the optimal allocation of portfolio assets in order to maximize a given investment objective. Traditionally, some form of mean-variance optimization is used with the aim of maximizing returns while minimizing risk, however, more recently, deep reinforcement learning formulations have been explored. Increasingly, investors have demonstrated an interest in incorporating ESG objectives when making investment decisions, and modifications to the classical mean-variance optimization framework have been developed. In this work, we study the use of deep reinforcement learning for responsible portfolio optimization, by incorporating ESG states and objectives, and provide comparisons against modified mean-variance approaches. Our results show that deep reinforcement learning policies can provide competitive performance against mean-variance approaches for responsible portfolio allocation across additive and multiplicative utility functions of financial and ESG responsibility objectives.
citation-role summary
citation-polarity summary
fields
math.OC 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Simulation-Based Neural Policies for Portfolio Choice: Architecture, Training, and Interpretability
In a controlled lifecycle benchmark, per-date backward neural policies beat single and two-regime networks on welfare and Bellman residuals, and a shape penalty fixes negative MPC violations.