Pith. sign in

REVIEW 2 cited by

Convergence of Multi-Scale Reinforcement Q-Learning Algorithms for Mean Field Game and Control Problems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.06659 v2 pith:YVSZ5RA5 submitted 2023-12-11 math.OC

Convergence of Multi-Scale Reinforcement Q-Learning Algorithms for Mean Field Game and Control Problems

classification math.OC
keywords convergencefieldmeanalgorithmalgorithmscontrolgamelearning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We establish the convergence of the unified two-timescale Reinforcement Learning (RL) algorithm presented in a previous work by Angiuli et al. This algorithm provides solutions to Mean Field Game (MFG) or Mean Field Control (MFC) problems depending on the ratio of two learning rates, one for the value function and the other for the mean field term. Our proof of convergence highlights the fact that in the case of MFC several mean field distributions need to be updated and for this reason we present two separate algorithms, one for MFG and one for MFC. We focus on a setting with finite state and action spaces, discrete time and infinite horizon. The proofs of convergence rely on a generalization of the two-timescale approach of Borkar. The accuracy of approximation to the true solutions depends on the smoothing of the policies. We provide a numerical example illustrating the convergence.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Continuous-time q-learning for mean-field control with common noise, part-II: q-learning algorithms

    math.OC 2026-04 unverdicted novelty 6.0

    The authors propose actor-critic q-learning algorithms for mean-field control with common noise based on martingale orthogonality conditions and relaxed controls, establish convergence of inner iterations in the linea...

  2. Mean Field Reinforcement Learning

    math.OC 2026-07 unverdicted novelty 2.0

    A monograph develops the probabilistic and control-theoretic framework connecting multi-agent reinforcement learning to mean field control, including analyses of Q-learning, policy gradients, and numerical methods for...