Pith. sign in

REVIEW 17 cited by

Game-Theoretic Multiagent Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.00583 v5 pith:JA2ILPP5 submitted 2020-11-01 cs.MA cs.AI

classification cs.MAcs.AI
keywords marllearningadvancesfieldgame-theoreticmultiagentrecentcovers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Tremendous advances have been made in multiagent reinforcement learning (MARL). MARL corresponds to the learning problem in a multiagent system in which multiple agents learn simultaneously. It is an interdisciplinary field of study with a long history that includes game theory, machine learning, stochastic control, psychology, and optimization. Despite great successes in MARL, there is a lack of a self-contained overview of the literature that covers game-theoretic foundations of modern MARL methods and summarizes the recent advances. The majority of existing surveys are outdated and do not fully cover the recent developments since 2010. In this work, we provide a monograph on MARL that covers both the fundamentals and the latest developments on the research frontier. The goal of this monograph is to provide a self-contained assessment of the current state-of-the-art MARL techniques from a game-theoretic perspective. We expect this work to serve as a stepping stone for both new researchers who are about to enter this fast-growing field and experts in the field who want to obtain a panoramic view and identify new directions based on recent advances.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. On the Geometry of Games and their Solvers

    cs.AI 2026-05 unverdicted novelty 7.0 of 10

    Introduces a structure-aware solver synthesis method with a learned game representation that organizes solvability into a continuous geometry aligned with solver dynamics.

  2. Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

    cs.AI 2026-05 unverdicted novelty 7.0 of 10

    New benchmark ICRL4AHT reveals that history-conditioned ICRL methods fail to show robust adaptation in multi-agent Overcooked-V2, underperforming random baselines on unseen teammates and layouts.

  3. Understanding Dynamics of Adam in Zero-Sum Games: An ODE Approach

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Derives ODE limits of Adam-DA showing that first- and second-order momentum parameters reverse their convergence roles in zero-sum games compared to minimization, validated on GAN experiments.

  4. Sample-efficient inductive matrix completion with noise and inexact side-information

    stat.ML 2026-05 unverdicted novelty 7.0 of 10

    Nonconvex projected gradient descent for noisy inductive matrix completion achieves linear convergence and order-optimal error at sample complexity scaling with side-information dimension a instead of ambient dimension n.

  5. Sample-efficient inductive matrix completion with noise and inexact side-information

    stat.ML 2026-05 unverdicted novelty 7.0 of 10

    A projected gradient descent algorithm for noisy inductive matrix completion achieves linear convergence and stable recovery at sample complexity governed by side-information dimension, extending to inexact side-infor...

  6. Equilibrium and Pricing in Consumer Networks with Nonlinear Utilities: An Online Shape-Constrained Learning Approach

    math.ST 2026-05 unverdicted novelty 7.0 of 10

    The paper establishes equilibrium existence and uniqueness for nonlinear utility consumer networks under contraction conditions and proposes a shape-constrained isotonic regression approach with strict no-regret conve...

  7. Vulnerable Agent Identification in Large-Scale Multi-Agent Reinforcement Learning

    cs.MA 2025-09 unverdicted novelty 7.0 of 10

    Proposes HAD-MFC framework that decouples upper-level vulnerable agent selection from lower-level adversarial policy learning in large-scale MARL using Fenchel-Rockafellar transform and MDP reformulation with provable...

  8. Do Not Discretize, Optimize: Almost Greedy Fictitious Play

    cs.GT 2026-06 unverdicted novelty 6.0 of 10

    Almost Greedy Fictitious Play achieves an instance-dependent O(1/T) convergence rate to Nash equilibrium in zero-sum games by optimizing stepsizes without discretization.

  9. Trajectory-Aware Retrieval Agents for Temporal Decision- Making

    cs.AI 2026-07 reject novelty 5.0 of 10

    TLM reports large accuracy gains on medical and financial temporal-decision tasks by fitting linear trends to retrieved embeddings, but its monotonicity theorem is circular and its baselines omit plain fine-tuned RAG.

  10. WebCQ: Cooperative Multi-Agent Deep Reinforcement Learning for Scalable Web GUI Testing

    cs.SE 2026-06 unverdicted novelty 5.0 of 10

    WebCQ applies cooperative MARL with QTRAN and DQN on semantic action vectors to web GUI testing, exploring 33.3% more states and 42.2% more actions than MARG on eight commercial sites.

  11. Conservative Equilibrium Discovery in Offline Game-Theoretic Multiagent Reinforcement Learning

    cs.AI 2026-02 conditional novelty 5.0 of 10

    COffeE-PSRO combines conservative uncertainty penalties with robust replicator dynamics to extract lower-regret equilibrium profiles from offline multi-agent datasets.

  12. An Agent-Centric Dynamical Systems Perspective on Multi-Agent Reinforcement Learning

    cs.MA 2025-12 conditional novelty 5.0 of 10

    Treating MARL training as coupled stochastic dynamical systems lets Lyapunov exponents, recurrence plots, and fractal dimensions characterize individual-agent stability and sensitivity.

  13. Policy Gradient with Self-Attention for Model-Free Distributed Nonlinear Multi-Agent Games

    eess.SY 2025-09 conditional novelty 5.0 of 10

    A self-attention policy trained with policy gradients learns distributed feedback control for multi-team games without models of dynamics or costs.

  14. Dilution, Diffusion and Symbiosis in Spatial Prisoner's Dilemma with Reinforcement Learning

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Adding a no-op 'persist' action to independent Q-learning agents creates a mutualistic shield that lets cooperation survive in a diluted, mobile spatial prisoner's dilemma.

  15. SOM: Structured Opponent Modeling for LLM-based Agents via Structural Causal Model

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    SOM uses a Structural Causal Model to create an explicit graph of opponent observation-to-action links, allowing LLMs to reason along those paths for more accurate and stable predictions in multi-agent settings.

  16. Homing through Reinforcement Learning

    cond-mat.soft 2026-02 reject novelty 4.0 of 10

    In a 2D Q-learning homing model, mean homing time is reported to be non-monotonic in rotational diffusion with a crossover at D_r≈12, and the learned policy is claimed to beat a stochastic-resetting ABP baseline.

  17. GenAI-based Multi-Agent Reinforcement Learning towards Distributed Agent Intelligence: A Generative-RL Agent Perspective

    cs.AI 2025-07 unverdicted novelty 3.0 of 10

    A position paper claiming that generative-AI agents that model and predict multi-agent dynamics will replace today's reactive MARL approaches.

Pith tools