Pith. sign in

REVIEW 1 cited by

The Dynamics of Q-learning in Population Games: a Physics-Inspired Continuity Equation Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.01500 v1 pith:NFQB5SRK submitted 2022-03-03 cs.MA

classification cs.MA
keywords modeldynamicsq-learninggamesequationpopulationcontinuitydifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although learning has found wide application in multi-agent systems, its effects on the temporal evolution of a system are far from understood. This paper focuses on the dynamics of Q-learning in large-scale multi-agent systems modeled as population games. We revisit the replicator equation model for Q-learning dynamics and observe that this model is inappropriate for our concerned setting. Motivated by this, we develop a new formal model, which bears a formal connection with the continuity equation in physics. We show that our model always accurately describes the Q-learning dynamics in population games across different initial settings of MASs and game configurations. We also show that our model can be applied to different exploration mechanisms, describe the mean dynamics, and be extended to Q-learning in 2-player and n-player games. Last but not least, we show that our model can provide insights into algorithm parameters and facilitate parameter tuning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations

    cs.MA 2024-12 conditional novelty 6.0 of 10

    A frequency-aware mean-field model of incremental Boltzmann Q-learning in the Prisoner's Dilemma predicts that apparent stable cooperation is a long metastable transient and that high discount factors induce oscillati...

Pith tools