Pith. sign in

REVIEW 2 cited by

Independent Learning in Stochastic Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.11743 v1 pith:H5YXYWFN submitted 2021-11-23 cs.GT cs.LGmath.DS

classification cs.GTcs.LGmath.DS
keywords learninggamesindependentdynamicsstochasticagentsdynamicmulti-agent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement learning (RL) has recently achieved tremendous successes in many artificial intelligence applications. Many of the forefront applications of RL involve multiple agents, e.g., playing chess and Go games, autonomous driving, and robotics. Unfortunately, the framework upon which classical RL builds is inappropriate for multi-agent learning, as it assumes an agent's environment is stationary and does not take into account the adaptivity of other agents. In this review paper, we present the model of stochastic games for multi-agent learning in dynamic environments. We focus on the development of simple and independent learning dynamics for stochastic games: each agent is myopic and chooses best-response type actions to other agents' strategy without any coordination with her opponent. There has been limited progress on developing convergent best-response type independent learning dynamics for stochastic games. We present our recently proposed simple and independent learning dynamics that guarantee convergence in zero-sum stochastic games, together with a review of other contemporaneous algorithms for dynamic multi-agent learning in this setting. Along the way, we also reexamine some classical results from both the game theory and RL literature, to situate both the conceptual contributions of our independent learning dynamics, and the mathematical novelties of our analysis. We hope this review paper serves as an impetus for the resurgence of studying independent and natural learning dynamics in game theory, for the more challenging settings with a dynamic environment.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Aggregate Fictitious Play for Learning in Anonymous Polymatrix Games (Extended Version)

    cs.GT 2025-08 conditional novelty 6.0 of 10

    In anonymous polymatrix games, fictitious play can be run on aggregate action counts without changing agents' best responses or losing convergence to Nash equilibrium.

  2. Ego-centric Learning of Communicative World Models for Autonomous Driving

    cs.RO 2025-06 reject novelty 5.0 of 10

    Sharing compressed latent states and planned waypoints between agents, triggered by prediction errors, improves multi-agent driving performance in CARLA while cutting communication bandwidth by roughly 50x.

Pith tools