Proposes a diffusion model guided by individual control barrier functions for safe offline multi-agent RL, recovering policies through inverse dynamics and showing safety gains on benchmarks.
arXiv preprint arXiv:1910.01465 , year=
5 Pith papers cite this work. Polarity classification is still indexing.
abstract
Many real world tasks require multiple agents to work together. Multi-agent reinforcement learning (RL) methods have been proposed in recent years to solve these tasks, but current methods often fail to efficiently learn policies. We thus investigate the presence of a common weakness in single-agent RL, namely value function overestimation bias, in the multi-agent setting. Based on our findings, we propose an approach that reduces this bias by using double centralized critics. We evaluate it on six mixed cooperative-competitive tasks, showing a significant advantage over current methods. Finally, we investigate the application of multi-agent methods to high-dimensional robotic tasks and show that our approach can be used to learn decentralized policies in this domain.
citation-role summary
citation-polarity summary
fields
cs.LG 5years
2026 5representative citing papers
An entity-graph MARL framework (RACHE) using R-GCN message passing and attention pooling over train-service nodes outperforms baseline algorithms in railway pricing revenue across two simulated market scenarios.
PC3D trains decentralized policies to recover and use personalized coordination context from local histories, enabling higher returns than baselines on variable-roster cooperative MARL tasks with both seen and unseen team sizes.
TRIDENT is a MARL framework using Richardson-Romberg gradient correction, Lyapunov-constrained trust-region updates, and a physics-informed residual critic that claims O(1/sqrt(K)) convergence to constrained Nash equilibrium with O(sqrt(K)) violation bounds and large reductions in training violation
ERPPO adds a DSA-based ambiguity estimator to MAPPO and switches between L1 and L2 entropy regularization to improve exploration and stability in non-stationary multi-dimensional observations.
citing papers explorer
-
ERPPO: Entropy Regularization-based Proximal Policy Optimization
ERPPO adds a DSA-based ambiguity estimator to MAPPO and switches between L1 and L2 entropy regularization to improve exploration and stability in non-stationary multi-dimensional observations.