Pith. sign in

REVIEW 1 cited by

Generalized Individual Q-learning for Polymatrix Games with Partial Observations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.02663 v1 pith:AQA333LQ submitted 2024-09-04 cs.GT cs.SYeess.SY

classification cs.GTcs.SYeess.SY
keywords actionsaccessagentsindividualq-learningconvergencedynamicsgames
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper addresses the challenge of limited observations in non-cooperative multi-agent systems where agents can have partial access to other agents' actions. We present the generalized individual Q-learning dynamics that combine belief-based and payoff-based learning for the networked interconnections of more than two self-interested agents. This approach leverages access to opponents' actions whenever possible, demonstrably achieving a faster (guaranteed) convergence to quantal response equilibrium in multi-agent zero-sum and potential polymatrix games. Notably, the dynamics reduce to the well-studied smoothed fictitious play and individual Q-learning under full and no access to opponent actions, respectively. We further quantify the improvement in convergence rate due to observing opponents' actions through numerical simulations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Aggregate Fictitious Play for Learning in Anonymous Polymatrix Games (Extended Version)

    cs.GT 2025-08 conditional novelty 6.0 of 10

    In anonymous polymatrix games, fictitious play can be run on aggregate action counts without changing agents' best responses or losing convergence to Nash equilibrium.

Pith tools