A two-stage actor-critic RL algorithm learns deterministic equilibrium policies for general time-inconsistent control problems by combining DPG on an auxiliary time-consistent problem with fixed-point iteration on auxiliary functions.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2verdicts
UNVERDICTED 2representative citing papers
Reformulates risk-sensitive benchmarked asset allocation as an LQG stochastic differential game via free energy-entropy duality and develops a continuous-time q-learning actor-critic algorithm that learns optimal policies with high accuracy in a proof-of-concept.
citing papers explorer
-
Deterministic Policy Gradient for Learning Equilibrium in Time-Inconsistent Control Problems
A two-stage actor-critic RL algorithm learns deterministic equilibrium policies for general time-inconsistent control problems by combining DPG on an auxiliary time-consistent problem with fixed-point iteration on auxiliary functions.
-
Reinforcement Learning for Risk-Sensitive Investment Management: a Free Energy--Entropy Duality Approach
Reformulates risk-sensitive benchmarked asset allocation as an LQG stochastic differential game via free energy-entropy duality and develops a continuous-time q-learning actor-critic algorithm that learns optimal policies with high accuracy in a proof-of-concept.