Pith. sign in

REVIEW 1 cited by

Unified Reinforcement Q-Learning for Mean Field Game and Control Problems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.13912 v3 pith:UX2V6BZ2 submitted 2020-06-24 math.OC cs.LGcs.MA

classification math.OCcs.LGcs.MA
keywords fieldmeanalgorithmproblemsagentasymptoticcontroldistribution
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a Reinforcement Learning (RL) algorithm to solve infinite horizon asymptotic Mean Field Game (MFG) and Mean Field Control (MFC) problems. Our approach can be described as a unified two-timescale Mean Field Q-learning: The \emph{same} algorithm can learn either the MFG or the MFC solution by simply tuning the ratio of two learning parameters. The algorithm is in discrete time and space where the agent not only provides an action to the environment but also a distribution of the state in order to take into account the mean field feature of the problem. Importantly, we assume that the agent can not observe the population's distribution and needs to estimate it in a model-free manner. The asymptotic MFG and MFC problems are also presented in continuous time and space, and compared with classical (non-asymptotic or stationary) MFG and MFC problems. They lead to explicit solutions in the linear-quadratic (LQ) case that are used as benchmarks for the results of our algorithm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games

    stat.ML 2025-05 conditional novelty 6.0 of 10

    Exact and sample-based trust-region policy optimization provably converge to approximate Nash equilibria in finite mean-field games with Õ(1/ε^6) sample complexity.

Pith tools