The authors create the first large-scale dataset and taxonomy of failure modes in multi-agent LLM systems to explain their limited performance gains.
Benchmarl: Benchmarking multi-agent reinforcement learning.Journal of Machine Learning Research, 25(217):1–10
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 2
citation-polarity summary
verdicts
UNVERDICTED 2roles
background 2polarities
background 2representative citing papers
PC3D trains decentralized policies to recover and use personalized coordination context from local histories, enabling higher returns than baselines on variable-roster cooperative MARL tasks with both seen and unseen team sizes.
citing papers explorer
-
Why Do Multi-Agent LLM Systems Fail?
The authors create the first large-scale dataset and taxonomy of failure modes in multi-agent LLM systems to explain their limited performance gains.
-
PC3D: Zero-Shot Cooperation Across Variable Rosters via Personalized Context Distillation
PC3D trains decentralized policies to recover and use personalized coordination context from local histories, enabling higher returns than baselines on variable-roster cooperative MARL tasks with both seen and unseen team sizes.