Sim2O enables efficient offline-to-online MARL by dynamically blending offline and online action proposals across agents and selecting high-value combinations via a centralized value function without auxiliary objectives.
Deep multi- agent reinforcement learning for decentralized continuous cooperative control
5 Pith papers cite this work, alongside 106 external citations. Polarity classification is still indexing.
years
2026 5representative citing papers
MAVIC corrects Bellman backups at instruction boundaries by adjusting the incoming objective and restoring continuation value, enabling consistent estimation under stochastic instruction switching in cooperative MARL.
A decoupled estimator combining gated dynamics learning and recursive Kalman filtering improves robustness of pre-trained MARL policies under stale observations and message loss.
A shared consensus vector, generated before any action, lets cooperative agents act simultaneously and lets the whole joint policy be trained with single-agent PPO.
Proposes a clipping objective for sequential trust-region updates in independent-actor cooperative MARL that yields a monotonic improvement bound and sub-linear convergence to epsilon-Nash equilibria while reducing advantage variance.
citing papers explorer
-
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition
Sim2O enables efficient offline-to-online MARL by dynamically blending offline and online action proposals across agents and selecting high-value combinations via a centralized value function without auxiliary objectives.
-
Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning
MAVIC corrects Bellman backups at instruction boundaries by adjusting the incoming objective and restoring continuation value, enabling consistent estimation under stochastic instruction switching in cooperative MARL.
-
Decoupled Delay Compensation: Enhancing Pre-trained MARL Policies via Learned Dynamics Filtering
A decoupled estimator combining gated dynamics learning and recursive Kalman filtering improves robustness of pre-trained MARL policies under stale observations and message loss.
-
Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus
A shared consensus vector, generated before any action, lets cooperative agents act simultaneously and lets the whole joint policy be trained with single-agent PPO.
-
Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning
Proposes a clipping objective for sequential trust-region updates in independent-actor cooperative MARL that yields a monotonic improvement bound and sub-linear convergence to epsilon-Nash equilibria while reducing advantage variance.