Pith. sign in

REVIEW 8 cited by

Unified continuous-time q-learning for mean-field game and mean-field control problems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.04521 v3 pith:LCFFQAKS submitted 2024-07-05 math.OC cs.LGq-fin.CP

Unified continuous-time q-learning for mean-field game and mean-field control problems

classification math.OC cs.LGq-fin.CP
keywords mean-fieldq-learningunifieddecoupledpolicyproblemsalgorithmcontinuous-time
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This paper studies the continuous-time q-learning in mean-field jump-diffusion models in a setting where the environment simulator does not provide direct access to the population distribution. We propose the integrated q-function in decoupled form (decoupled Iq-function) and establish its martingale characterization, which provides a unified policy evaluation rule for both mean-field game (MFG) and mean-field control (MFC) problems. Moreover, we consider the learning procedure where population distribution is updated based on the representative agent's state values. Depending on the task to solve the MFG or MFC problem, we can employ the decoupled Iq-function differently to characterize the mean-field equilibrium policy or the mean-field optimal policy respectively. Based on these theoretical findings, we devise a unified parametric q-learning algorithm for both MFG and MFC problems by utilizing test policies and the averaged martingale orthogonality condition. In two applications within and beyond LQ framework, we illustrate the effectiveness and efficiency of our unified parametric q-learning algorithm for both MFG and MFC learning tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise

    math.OC 2026-06 unverdicted novelty 7.0

    Robust Q-learning algorithm with convergence and finite-time bounds for mean-field control under Wasserstein uncertainty in common noise.

  2. Policy Gradient for Continuous-Time Mean-Field Control

    math.OC 2026-05 conditional novelty 7.0

    Derives an explicit Gâteaux policy-gradient formula for entropy-regularized continuous-time mean-field control using the value function and cylindrical representations, then builds a model-based actor-critic scheme wi...

  3. Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations

    math.OC 2026-04 unverdicted novelty 7.0

    Establishes existence and uniqueness for optimal policies in continuous-time entropy-regularized mean-field control with common noise via an integrated q-function, plus explicit Gaussian characterization in the LQ setting.

  4. Reinforcement learning for irreversible reinsurance problems: the randomized singular control approach

    math.OC 2025-12 conditional novelty 7.0

    A randomized, entropy-regularized singular control law enables continuous-time reinforcement learning to solve irreversible reinsurance problems, with an explicit equilibrium boundary Γ(x)=e^{-βΦ(x)/λ}.

  5. Continuous-Time Reinforcement Learning for $N$-Player Stochastic Differential Games with Exploratory Policies

    math.OC 2026-07 conditional novelty 6.0

    For entropy-regularized N-player differential games, a Nash-type equilibrium exists exactly when the Gibbs conditional best responses are jointly compatible, checkable via a cross-partial criterion on the learned q-functions.

  6. Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies

    math.OC 2026-07 accept novelty 6.0

    Model-free deterministic policy gradients and a continuous-time deep actor-critic algorithm solve extended mean-field control problems whose dynamics and rewards depend on the joint state-control law.

  7. Continuous-time q-learning for mean-field control with common noise, part-II: q-learning algorithms

    math.OC 2026-04 unverdicted novelty 6.0

    The authors propose actor-critic q-learning algorithms for mean-field control with common noise based on martingale orthogonality conditions and relaxed controls, establish convergence of inner iterations in the linea...

  8. Continuous-time reinforcement learning for optimal switching over multiple regimes

    math.OC 2025-12 conditional novelty 5.0

    An entropy-regularized exploratory formulation of multi-regime optimal switching is shown to admit well-posed HJB systems, fast-converging policy iteration, and a vanishing-entropy limit that recovers the classical problem.