Pith. sign in

REVIEW 1 cited by

Reinforcement Learning for SBM Graphon Games with Re-Sampling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.16326 v1 pith:MPVJZTAN submitted 2023-10-25 cs.GT cs.LG

classification cs.GTcs.LG
keywords modelmp-mfgdynamicsggr-sgraphonlearningre-samplingagents
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The Mean-Field approximation is a tractable approach for studying large population dynamics. However, its assumption on homogeneity and universal connections among all agents limits its applicability in many real-world scenarios. Multi-Population Mean-Field Game (MP-MFG) models have been introduced in the literature to address these limitations. When the underlying Stochastic Block Model is known, we show that a Policy Mirror Ascent algorithm finds the MP-MFG Nash Equilibrium. In more realistic scenarios where the block model is unknown, we propose a re-sampling scheme from a graphon integrated with the finite N-player MP-MFG model. We develop a novel learning framework based on a Graphon Game with Re-Sampling (GGR-S) model, which captures the complex network structures of agents' connections. We analyze GGR-S dynamics and establish the convergence to dynamics of MP-MFG. Leveraging this result, we propose an efficient sample-based N-player Reinforcement Learning algorithm for GGR-S without population manipulation, and provide a rigorous convergence analysis with finite sample guarantee.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Policy Optimization for Continuous-time Linear-Quadratic Graphon Mean Field Games

    math.OC 2025-06 accept novelty 7.0 of 10

    A bilevel policy optimization algorithm for continuous-time linear-quadratic graphon mean field games converges linearly to best-response policies and globally to the Nash equilibrium.

Pith tools