Pith. sign in

REVIEW 11 cited by

Multi-agent Reinforcement Learning: A Comprehensive Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.10256 v2 pith:LM5TSEEX submitted 2023-12-15 cs.MA cs.AIcs.LG

Multi-agent Reinforcement Learning: A Comprehensive Survey

classification cs.MA cs.AIcs.LG
keywords marlapplicationschallengeslearningmulti-agentsurveyagentscomprehensive
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Multi-agent systems (MAS) are widely prevalent and crucially important in numerous real-world applications, where multiple agents must make decisions to achieve their objectives in a shared environment. Despite their ubiquity, the development of intelligent decision-making agents in MAS poses several open challenges to their effective implementation. This survey examines these challenges, placing an emphasis on studying seminal concepts from game theory (GT) and machine learning (ML) and connecting them to recent advancements in multi-agent reinforcement learning (MARL), i.e. the research of data-driven decision-making within MAS. Therefore, the objective of this survey is to provide a comprehensive perspective along the various dimensions of MARL, shedding light on the unique opportunities that are presented in MARL applications while highlighting the inherent challenges that accompany this potential. Therefore, we hope that our work will not only contribute to the field by analyzing the current landscape of MARL but also motivate future directions with insights for deeper integration of concepts from related domains of GT and ML. With this in mind, this work delves into a detailed exploration of recent and past efforts of MARL and its related fields and describes prior solutions that were proposed and their limitations, as well as their applications.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Asymmetric physics enables efficient learning in quadrupedal robot swarms

    cs.RO 2026-06 unverdicted novelty 6.0

    Asymmetric physics (high-fidelity non-diff simulator plus differentiable surrogates) enables end-to-end training of decentralized vision-based policies for up to 512 quadrupeds that transfer zero-shot to real hardware.

  2. TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-LLM Reasoning

    cs.AI 2026-05 unverdicted novelty 6.0

    TRACER combines a controller-regret layer using regret matching for speak/skip decisions with a generation-credit layer using GSPO rewards to enable learned collaboration in multi-LLM reasoning.

  3. MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning

    cs.MA 2026-05 unverdicted novelty 6.0

    MAGIC extracts multi-step causal influences between agents using interventional conditional mutual information and advantage gating to generate intrinsic rewards that improve coordination in MARL.

  4. MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning

    cs.MA 2026-05 unverdicted novelty 6.0

    MAGIC estimates multi-step action effects between agents with counterfactual interventions, gates them by advantage, and converts them to intrinsic rewards, yielding 26.9% and 10.1% relative gains on MPE and SMAC benchmarks.

  5. Subjective-Graph LLM Agents for Simulating Uncertainty in Classroom Social Perception

    cs.AI 2026-03 conditional novelty 6.0

    Subjective-graph LLM agents on 12 real classrooms accumulate collective ranking error from 0.066 to 0.124 over six exams despite repeated score anchors.

  6. The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

    cs.AI 2025-09 accept novelty 6.0

    Survey that defines agentic RL for LLMs via POMDPs, introduces a taxonomy of planning/tool-use/memory/reasoning capabilities and domains, and compiles open environments from over 500 papers.

  7. Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning

    cs.LG 2026-06 unverdicted novelty 5.0

    Survey mapping RL techniques onto LLM training and highlighting gaps in value-based, off-policy, and bootstrapping methods.

  8. Trust Region On-Policy Distillation

    cs.LG 2026-05 unverdicted novelty 5.0

    TrOPD stabilizes on-policy distillation for LLMs with trust-region learning, outlier estimation, and off-policy guidance, outperforming prior OPD methods on reasoning and code benchmarks.

  9. Building Better Environments for Autonomous Cyber Defence

    cs.CR 2026-04 conditional novelty 5.0

    A workshop synthesis provides a decomposition framework for RL-cyber environment interfaces and best-practice guidelines for training and evaluating autonomous cyber defence agents.

  10. Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures

    cs.AI 2026-04 unverdicted novelty 4.0

    A survey comparing classical multi-agent systems with large foundation model-enabled multi-agent systems, showing how the latter enables semantic-level collaboration and greater adaptability.

  11. A Survey of Reinforcement Learning for Large Reasoning Models

    cs.CL 2025-09 accept novelty 3.0

    A survey compiling RL methods, challenges, data resources, and applications for enhancing reasoning in large language models and large reasoning models since DeepSeek-R1.