REVIEW 8 cited by
Unified continuous-time q-learning for mean-field game and mean-field control problems
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Unified continuous-time q-learning for mean-field game and mean-field control problems
read the original abstract
This paper studies the continuous-time q-learning in mean-field jump-diffusion models in a setting where the environment simulator does not provide direct access to the population distribution. We propose the integrated q-function in decoupled form (decoupled Iq-function) and establish its martingale characterization, which provides a unified policy evaluation rule for both mean-field game (MFG) and mean-field control (MFC) problems. Moreover, we consider the learning procedure where population distribution is updated based on the representative agent's state values. Depending on the task to solve the MFG or MFC problem, we can employ the decoupled Iq-function differently to characterize the mean-field equilibrium policy or the mean-field optimal policy respectively. Based on these theoretical findings, we devise a unified parametric q-learning algorithm for both MFG and MFC problems by utilizing test policies and the averaged martingale orthogonality condition. In two applications within and beyond LQ framework, we illustrate the effectiveness and efficiency of our unified parametric q-learning algorithm for both MFG and MFC learning tasks.
Forward citations
Cited by 8 Pith papers
-
Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise
Robust Q-learning algorithm with convergence and finite-time bounds for mean-field control under Wasserstein uncertainty in common noise.
-
Policy Gradient for Continuous-Time Mean-Field Control
Derives an explicit Gâteaux policy-gradient formula for entropy-regularized continuous-time mean-field control using the value function and cylindrical representations, then builds a model-based actor-critic scheme wi...
-
Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations
Establishes existence and uniqueness for optimal policies in continuous-time entropy-regularized mean-field control with common noise via an integrated q-function, plus explicit Gaussian characterization in the LQ setting.
-
Reinforcement learning for irreversible reinsurance problems: the randomized singular control approach
A randomized, entropy-regularized singular control law enables continuous-time reinforcement learning to solve irreversible reinsurance problems, with an explicit equilibrium boundary Γ(x)=e^{-βΦ(x)/λ}.
-
Continuous-Time Reinforcement Learning for $N$-Player Stochastic Differential Games with Exploratory Policies
For entropy-regularized N-player differential games, a Nash-type equilibrium exists exactly when the Gibbs conditional best responses are jointly compatible, checkable via a cross-partial criterion on the learned q-functions.
-
Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies
Model-free deterministic policy gradients and a continuous-time deep actor-critic algorithm solve extended mean-field control problems whose dynamics and rewards depend on the joint state-control law.
-
Continuous-time q-learning for mean-field control with common noise, part-II: q-learning algorithms
The authors propose actor-critic q-learning algorithms for mean-field control with common noise based on martingale orthogonality conditions and relaxed controls, establish convergence of inner iterations in the linea...
-
Continuous-time reinforcement learning for optimal switching over multiple regimes
An entropy-regularized exploratory formulation of multi-regime optimal switching is shown to admit well-posed HJB systems, fast-converging policy iteration, and a vanishing-entropy limit that recovers the classical problem.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.