HiComm proposes a plug-in hierarchical communication protocol for cooperative MARL that performs structured information retrieval over observation hierarchies using receiver queries and three-stage decoding, matching or outperforming baselines while reducing volume by up to 23×.
Smacv2: An improved benchmark for cooperative multi-agent reinforcement learning.Advances in Neural Information Processing Systems, 36:37567–37593
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2verdicts
UNVERDICTED 2representative citing papers
A C++ Dec-POMDP simulator using data-oriented design and zero-copy PyTorch integration achieves up to 33 million steps per second on a 16-core CPU, enabling multi-agent policy training in minutes with PPO, DQN, and SAC.
citing papers explorer
-
HiComm: Hierarchical Communication for Multi-agent Reinforcement Learning
HiComm proposes a plug-in hierarchical communication protocol for cooperative MARL that performs structured information retrieval over observation hierarchies using receiver queries and three-stage decoding, matching or outperforming baselines while reducing volume by up to 23×.
-
A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations
A C++ Dec-POMDP simulator using data-oriented design and zero-copy PyTorch integration achieves up to 33 million steps per second on a 16-core CPU, enabling multi-agent policy training in minutes with PPO, DQN, and SAC.