A two-stage actor-critic RL algorithm learns deterministic equilibrium policies for general time-inconsistent control problems by combining DPG on an auxiliary time-consistent problem with fixed-point iteration on auxiliary functions.
Being serious about non-commitment: subgame perfect equilibrium in continuous time
3 Pith papers cite this work. Polarity classification is still indexing.
abstract
This paper characterizes differentiable subgame perfect equilibria in a continuous time intertemporal decision optimization problem with non-constant discounting. The equilibrium equation takes two different forms, one of which is reminescent of the classical Hamilton-Jacobi-Bellman equation of optimal control, but with a non-local term. We give a local existence result, and several examples in the consumption saving problem. The analysis is then applied to suggest that non constant discount rates generate an indeterminacy of the steady state in the Ramsey growth model. Despite its indeterminacy, the steady state level is robust to small deviations from constant discount rates.
years
2026 3representative citing papers
A mean-field equilibrium exists for a mean-variance portfolio game in which each agent's risk aversion switches discontinuously as wealth crosses the population average.
PG-DPO is a new variational framework that replaces Bellman recursion with a Pontryagin-guided adjoint-MC projection for RL under non-exponential discounting and shows gains on hyperbolic and survival benchmarks.
citing papers explorer
-
Deterministic Policy Gradient for Learning Equilibrium in Time-Inconsistent Control Problems
A two-stage actor-critic RL algorithm learns deterministic equilibrium policies for general time-inconsistent control problems by combining DPG on an auxiliary time-consistent problem with fixed-point iteration on auxiliary functions.
-
Mean-field game of mean-variance portfolio optimization with peer-based risk aversion
A mean-field equilibrium exists for a mean-variance portfolio game in which each agent's risk aversion switches discontinuously as wealth crosses the population average.
-
Beyond the Bellman Recursion: A Pontryagin-Guided Framework for Non-Exponential Discounting
PG-DPO is a new variational framework that replaces Bellman recursion with a Pontryagin-guided adjoint-MC projection for RL under non-exponential discounting and shows gains on hyperbolic and survival benchmarks.