CLOVER augments value decomposition with a GNN mixer whose weights depend on the realized wireless communication graph, proving permutation invariance, monotonicity, and greater expressiveness than QMIX while showing gains on Predator-Prey and Lumberjacks under p-CSMA channels.
Multiagent Bidirectionally-Coordinated Nets: Emergence of Human-level Coordination in Learning to Play StarCraft Combat Games
5 Pith papers cite this work. Polarity classification is still indexing.
abstract
Many artificial intelligence (AI) applications often require multiple intelligent agents to work in a collaborative effort. Efficient learning for intra-agent communication and coordination is an indispensable step towards general AI. In this paper, we take StarCraft combat game as a case study, where the task is to coordinate multiple agents as a team to defeat their enemies. To maintain a scalable yet effective communication protocol, we introduce a Multiagent Bidirectionally-Coordinated Network (BiCNet ['bIknet]) with a vectorised extension of actor-critic formulation. We show that BiCNet can handle different types of combats with arbitrary numbers of AI agents for both sides. Our analysis demonstrates that without any supervisions such as human demonstrations or labelled data, BiCNet could learn various types of advanced coordination strategies that have been commonly used by experienced game players. In our experiments, we evaluate our approach against multiple baselines under different scenarios; it shows state-of-the-art performance, and possesses potential values for large-scale real-world applications.
verdicts
UNVERDICTED 5representative citing papers
SCALE-COMM uses contrastive alignment on latent embeddings to decouple and stabilize communication learning from policy optimization in decentralized MARL, showing gains on benchmarks and a warehouse task.
Delayed multi-agent messages are scored by communication gain minus delay cost; agents request and fuse them only when predicted net value is positive, with a value-loss bound.
An attention-augmented actor-critic agent learns to dynamically weight multiple environment views by importance and outperforms baselines on TORCS and three other 3D simulators under noise and partial observability.
HRL-IM/CBS encodes battlefield states via influence map hashing and uses cluster-based scripts in a multi-Q-table hierarchy for StarCraft micromanagement, claiming competitive results with improved sample efficiency and interpretability over deep RL baselines.
citing papers explorer
-
Wireless Communication Enhanced Value Decomposition for Multi-Agent Reinforcement Learning
CLOVER augments value decomposition with a GNN mixer whose weights depend on the realized wireless communication graph, proving permutation invariance, monotonicity, and greater expressiveness than QMIX while showing gains on Predator-Prey and Lumberjacks under p-CSMA channels.
-
SCALE-COMM: Shared, Contrastively-Aligned Latent Embeddings for MARL Communication
SCALE-COMM uses contrastive alignment on latent embeddings to decouple and stabilize communication learning from policy optimization in decentralized MARL, showing gains on benchmarks and a warehouse task.
-
Communication Gain and Delay Cost Under Cross-Timestep Delays in Cooperative Multi-Agent Reinforcement Learning
Delayed multi-agent messages are scored by communication gain minus delay cost; agents request and fuse them only when predicted net value is positive, with a value-loss bound.
-
An Actor-Critic-Attention Mechanism for Deep Reinforcement Learning in Multi-view Environments
An attention-augmented actor-critic agent learns to dynamically weight multiple environment views by importance and outperforms baselines on TORCS and three other 3D simulators under noise and partial observability.
-
Hierarchical Reinforcement Learning in StarCraft Micromanagement with Influence Maps and Cluster-based Scripts
HRL-IM/CBS encodes battlefield states via influence map hashing and uses cluster-based scripts in a multi-Q-table hierarchy for StarCraft micromanagement, claiming competitive results with improved sample efficiency and interpretability over deep RL baselines.