REVIEW 3 cited by
Decentralized Policy Optimization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The study of decentralized learning or independent learning in cooperative multi-agent reinforcement learning has a history of decades. Recently empirical studies show that independent PPO (IPPO) can obtain good performance, close to or even better than the methods of centralized training with decentralized execution, in several benchmarks. However, decentralized actor-critic with convergence guarantee is still open. In this paper, we propose \textit{decentralized policy optimization} (DPO), a decentralized actor-critic algorithm with monotonic improvement and convergence guarantee. We derive a novel decentralized surrogate for policy optimization such that the monotonic improvement of joint policy can be guaranteed by each agent \textit{independently} optimizing the surrogate. In practice, this decentralized surrogate can be realized by two adaptive coefficients for policy optimization at each agent. Empirically, we compare DPO with IPPO in a variety of cooperative multi-agent tasks, covering discrete and continuous action spaces, and fully and partially observable environments. The results show DPO outperforms IPPO in most tasks, which can be the evidence for our theoretical results.
Forward citations
Cited by 3 Pith papers
-
Dynamic Graph Communication for Decentralised Multi-Agent Reinforcement Learning
A GAT-based aggregator and a learned iteration controller improve NetMon's decentralized packet routing in simulated dynamic networks by 9.5% reward while using 6.4% less communication.
-
Achieving Collective Welfare in Multi-Agent Reinforcement Learning via Suggestion Sharing
A suggestion-sharing MARL algorithm lets agents exchange optimized action proposals for each other, with a theoretical bound relating the surrogate objective to collective return.
-
CORD: Generalizable Cooperation via Role Diversity
CORD improves zero-shot cooperation in multi-agent games by learning diverse, causally informed role assignments through an entropy-based objective.
Discussion (0). Continue with ORCID to comment.