Robust Q-learning algorithm with convergence and finite-time bounds for mean-field control under Wasserstein uncertainty in common noise.
Distributionally Robust Deep Q-Learning
2 Pith papers cite this work. Polarity classification is still indexing.
abstract
We propose a novel distributionally robust $Q$-learning algorithm for the non-tabular case accounting for continuous state spaces where the state transition of the underlying Markov decision process is subject to model uncertainty. The uncertainty is taken into account by considering the worst-case transition from a ball around a reference probability measure. To determine the optimal policy under the worst-case state transition, we solve the associated non-linear Bellman equation by dualising and regularising the Bellman operator with the Sinkhorn distance, which is then parameterized with deep neural networks. This approach allows us to modify the Deep Q-Network algorithm to optimise for the worst case state transition. We illustrate the tractability and effectiveness of our approach through several applications, including a portfolio optimisation task based on S\&{P}~500 data.
years
2026 2representative citing papers
In high-frequency market making, action robustness (Sinkhorn regularization) dominates uncertainty tolerance in reshaping sequential quoting and inventory, while excessive robustness can cut execution opportunities in illiquid names.
citing papers explorer
-
Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise
Robust Q-learning algorithm with convergence and finite-time bounds for mean-field control under Wasserstein uncertainty in common noise.
-
Robustness in Sequential Decision Making under Evolving Uncertainty: Evidence from High-Frequency Market Making
In high-frequency market making, action robustness (Sinkhorn regularization) dominates uncertainty tolerance in reshaping sequential quoting and inventory, while excessive robustness can cut execution opportunities in illiquid names.