REVIEW 5 cited by
Revisiting Parameter Sharing in Multi-Agent Deep Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Parameter sharing, where each agent independently learns a policy with fully shared parameters between all policies, is a popular baseline method for multi-agent deep reinforcement learning. Unfortunately, since all agents share the same policy network, they cannot learn different policies or tasks. This issue has been circumvented experimentally by adding an agent-specific indicator signal to observations, which we term "agent indication". Agent indication is limited, however, in that without modification it does not allow parameter sharing to be applied to environments where the action spaces and/or observation spaces are heterogeneous. This work formalizes the notion of agent indication and proves that it enables convergence to optimal policies for the first time. Next, we formally introduce methods to extend parameter sharing to learning in heterogeneous observation and action spaces, and prove that these methods allow for convergence to optimal policies. Finally, we experimentally confirm that the methods we introduce function empirically, and conduct a wide array of experiments studying the empirical efficacy of many different agent indication schemes for image based observation spaces.
Forward citations
Cited by 5 Pith papers
-
Hierarchical Message-Passing Policies for Multi-Agent Reinforcement Learning
A feudal hierarchical MARL method where lower-level policies are rewarded with the upper level's advantage function, with theoretical alignment guarantees and strong benchmark results.
-
Low-Rank Agent-Specific Adaptation (LoRASA) for Multi-Agent Policy Learning
Per-agent low-rank adapters on a shared backbone let multi-agent policies specialize at a fraction of the memory cost of separate networks, with competitive benchmark performance.
-
An Extended Benchmarking of Multi-Agent Reinforcement Learning Algorithms in Complex Fully Cooperative Tasks
Algorithms that set state of the art on SMAC and GRF often underperform standard baselines on fully cooperative benchmarks, including image-based tasks.
-
Multi-Agent Reinforcement Learning for Dynamic Mobility Resource Allocation with Hierarchical Adaptive Grouping
HAG-PS adds hierarchical adaptive grouping and identity embeddings to MARL parameter sharing, reaching 77.21% fulfilled service ratio on simulated January 2024 Manhattan bike sharing.
-
Wasserstein-Barycenter Consensus for Cooperative Multi-Agent Reinforcement Learning
A cooperative MARL algorithm that regularizes each agent's policy toward the Sinkhorn barycenter of the team's visitation distributions, with a claimed but insufficiently proven geometric convergence guarantee.
Discussion (0). Continue with ORCID to comment.