Pith. sign in

REVIEW 1 cited by

FP3O: Enabling Proximal Policy Optimization in Multi-Agent Cooperation with Parameter-Sharing Versatility

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.05053 v1 pith:OB26OL6I submitted 2023-10-08 cs.LG cs.AIcs.MA

classification cs.LGcs.AIcs.MA
keywords multi-agentfp3oalgorithmcooperativefull-pipelineinterconnectionsmarloptimization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing multi-agent PPO algorithms lack compatibility with different types of parameter sharing when extending the theoretical guarantee of PPO to cooperative multi-agent reinforcement learning (MARL). In this paper, we propose a novel and versatile multi-agent PPO algorithm for cooperative MARL to overcome this limitation. Our approach is achieved upon the proposed full-pipeline paradigm, which establishes multiple parallel optimization pipelines by employing various equivalent decompositions of the advantage function. This procedure successfully formulates the interconnections among agents in a more general manner, i.e., the interconnections among pipelines, making it compatible with diverse types of parameter sharing. We provide a solid theoretical foundation for policy improvement and subsequently develop a practical algorithm called Full-Pipeline PPO (FP3O) by several approximations. Empirical evaluations on Multi-Agent MuJoCo and StarCraftII tasks demonstrate that FP3O outperforms other strong baselines and exhibits remarkable versatility across various parameter-sharing configurations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    Adaptive per-agent KL-threshold allocation via KKT (HATRPO-W) and greedy (HATRPO-G) improves HATRPO's final reward by over 22.5% in MARL benchmarks.

Pith tools