Pith. sign in

REVIEW 6 cited by

Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.11251 v2 pith:UNKTTGI2 submitted 2021-09-23 cs.AI cs.MA

classification cs.AIcs.MA
keywords policyhappolearningmulti-agentregiontrustagentshatrpo
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Trust region methods rigorously enabled reinforcement learning (RL) agents to learn monotonically improving policies, leading to superior performance on a variety of tasks. Unfortunately, when it comes to multi-agent reinforcement learning (MARL), the property of monotonic improvement may not simply apply; this is because agents, even in cooperative games, could have conflicting directions of policy updates. As a result, achieving a guaranteed improvement on the joint policy where each agent acts individually remains an open challenge. In this paper, we extend the theory of trust region learning to MARL. Central to our findings are the multi-agent advantage decomposition lemma and the sequential policy update scheme. Based on these, we develop Heterogeneous-Agent Trust Region Policy Optimisation (HATPRO) and Heterogeneous-Agent Proximal Policy Optimisation (HAPPO) algorithms. Unlike many existing MARL algorithms, HATRPO/HAPPO do not need agents to share parameters, nor do they need any restrictive assumptions on decomposibility of the joint value function. Most importantly, we justify in theory the monotonic improvement property of HATRPO/HAPPO. We evaluate the proposed methods on a series of Multi-Agent MuJoCo and StarCraftII tasks. Results show that HATRPO and HAPPO significantly outperform strong baselines such as IPPO, MAPPO and MADDPG on all tested tasks, therefore establishing a new state of the art.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 83 citations worldwide. Full citation record

  1. Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Online action-space factorization via Kalman-refined cross-capacitance lets shared multi-agent policies zero-shot tune larger quantum-dot arrays with near-constant steps.

  2. TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning

    cs.MA 2026-02 unverdicted novelty 6.0 of 10

    TABX is a JAX-based, GPU-accelerated, configurable multi-agent battle simulator that lets researchers vary units, terrain, and physics to benchmark cooperative MARL algorithms.

  3. Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints

    cs.LG 2026-08 conditional novelty 5.0 of 10

    HeLyMARL uses virtual queues and sequential HAPPO updates to pace BS energy and user handover budgets within an episode, outperforming greedy and Lagrangian baselines in simulations.

  4. Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    Adaptive per-agent KL-threshold allocation via KKT (HATRPO-W) and greedy (HATRPO-G) improves HATRPO's final reward by over 22.5% in MARL benchmarks.

  5. Improving monotonic optimization in heterogeneous multi-agent reinforcement learning with optimal marginal deterministic policy gradient

    cs.AI 2025-07 reject novelty 5.0 of 10

    OMDPG combines optimal marginal Q-values with pessimistic Q-critics to reconcile monotonic improvement with partial parameter sharing in heterogeneous multi-agent RL.

  6. Light Aircraft Game : Basic Implementation and training results analysis

    cs.LG 2025-06 reject novelty 5.0 of 10

    In the new LAG air-combat environment, HASAC scores higher than HAPPO in no-weapon coordination tasks while HAPPO scores higher in missile combat, but the results come from single runs without error bars.

Pith tools