Pith. sign in

REVIEW 1 cited by

Model Free Reinforcement Learning Algorithm for Stationary Mean field Equilibrium for Multiple Types of Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.15377 v1 pith:A2Z5FYFP submitted 2020-12-31 cs.MA cs.AIcs.GTcs.LGcs.SYeess.SY

classification cs.MAcs.AIcs.GTcs.LGcs.SYeess.SY
keywords agentsstateagentstationaryalgorithmequilibriuminfiniteinteraction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We consider a multi-agent Markov strategic interaction over an infinite horizon where agents can be of multiple types. We model the strategic interaction as a mean-field game in the asymptotic limit when the number of agents of each type becomes infinite. Each agent has a private state; the state evolves depending on the distribution of the state of the agents of different types and the action of the agent. Each agent wants to maximize the discounted sum of rewards over the infinite horizon which depends on the state of the agent and the distribution of the state of the leaders and followers. We seek to characterize and compute a stationary multi-type Mean field equilibrium (MMFE) in the above game. We characterize the conditions under which a stationary MMFE exists. Finally, we propose Reinforcement learning (RL) based algorithm using policy gradient approach to find the stationary MMFE when the agents are unaware of the dynamics. We, numerically, evaluate how such kind of interaction can model the cyber attacks among defenders and adversaries, and show how RL based algorithm can converge to an equilibrium.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Meta-Inverse Reinforcement Learning for Mean Field Games via Probabilistic Context Variables

    cs.LG 2025-09 reject novelty 5.0 of 10

    A meta-IRL algorithm for mean field games that learns task-conditioned rewards from mixed-type expert trajectories using a latent context variable.

Pith tools