Pith. sign in

REVIEW 18 cited by

Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1706.02275 v4 pith:6RD3BRGA submitted 2017-06-07 cs.LG cs.AIcs.NE

classification cs.LGcs.AIcs.NE
keywords multi-agentpoliciesmethodsableactor-criticagentagentscoordination
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We explore deep reinforcement learning methods for multi-agent domains. We begin by analyzing the difficulty of traditional algorithms in the multi-agent case: Q-learning is challenged by an inherent non-stationarity of the environment, while policy gradient suffers from a variance that increases as the number of agents grows. We then present an adaptation of actor-critic methods that considers action policies of other agents and is able to successfully learn policies that require complex multi-agent coordination. Additionally, we introduce a training regimen utilizing an ensemble of policies for each agent that leads to more robust multi-agent policies. We show the strength of our approach compared to existing methods in cooperative as well as competitive scenarios, where agent populations are able to discover various physical and informational coordination strategies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Quantum Advantage in Multi Agent Reinforcement Learning

    cs.LG 2026-05 conditional novelty 6.0 of 10

    Entangled QMARL agents approach the Tsirelson bound of 0.854 in CHSH while unentangled versions match classical baselines, and hybrid quantum-classical setups outperform both in CoopNav.

  2. A Distributionally Robust Reinforcement Learning Framework for Constrained Urban EV Dispatch

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    A robust semi-Markov RL agent with MILP feasibility projection and Wasserstein ambiguity set achieves $1.22M net profit on an NYC EV simulator with zero feeder violations, outperforming heuristic and other RL baselines.

  3. Scalable Neighborhood-Based Multi-Agent Actor-Critic

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    MADDPG-K scales centralized critics in multi-agent RL by limiting each critic to k-nearest neighbors under Euclidean distance, yielding constant input size and competitive performance.

  4. Coupling Smoothed Particle Hydrodynamics with Multi-Agent Deep Reinforcement Learning for Cooperative Control of Point Absorbers

    eess.SY 2026-01 conditional novelty 6.0 of 10

    A GPU-coupled SPH and multi-agent reinforcement learning platform learns cooperative PTO damping policies that increase simulated wave-energy capture by up to 23.8% over fixed damping.

  5. Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage

    cs.AI 2025-06 unverdicted novelty 6.0 of 10

    CL-MARL uses an adaptive curriculum scheduler called FlexDiff and Counterfactual Group Relative Policy Advantage to break static-difficulty training in MARL and achieve higher win rates on hard StarCraft maps.

  6. Asynchronous Cooperative Multi-Agent Reinforcement Learning with Limited Communication

    cs.MA 2025-02 unverdicted novelty 6.0 of 10

    AsynCoMARL is a new asynchronous MARL algorithm that matches leading baselines on success and collision rates while using 26% fewer messages via graph transformers on dynamic communication graphs.

  7. Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details

    cs.AI 2026-08 conditional novelty 5.0 of 10

    For Other-Play in Yokai, agents trained with different implementation details coordinate across implementations about as well as across seeds, supporting inter-seed cross-play as a proxy for cross-implementation evaluation.

  8. Integrated Altruistic and Fairness Preference Induces Advanced Mutual Cooperation in Sequential Social Dilemmas

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Integrating forward-looking altruistic and fairness preferences into agent utilities produces mutual cooperation and higher collective returns than egoistic or inequity-aversion baselines in sequential social dilemmas.

  9. Decoupled Delay Compensation: Enhancing Pre-trained MARL Policies via Learned Dynamics Filtering

    cs.MA 2026-05 unverdicted novelty 5.0 of 10

    A decoupled estimator combining gated dynamics learning and recursive Kalman filtering improves robustness of pre-trained MARL policies under stale observations and message loss.

  10. Coordination Architecture Shapes Continuous Demand Response Outcomes in Building Districts

    eess.SY 2026-05 unverdicted novelty 5.0 of 10

    In a 25-building district simulation, the hybrid MPC-SAC architecture delivered the strongest balance of load tracking accuracy (4.8% NMBE), thermal comfort (16.8% exceedance), and lowest spatial variability compared ...

  11. A Distributionally Robust Reinforcement Learning Framework for Constrained Urban EV Dispatch

    cs.AI 2026-04 unverdicted novelty 5.0 of 10

    PD-RSAC, a distributionally robust SAC variant with GCN encoder and MILP constraint projection, reports $1.22M net profit on an NYC taxi-based EV simulator while achieving zero feeder violations, outperforming heurist...

  12. Multi-agent DRL-based Lane Change Decision Model for Cooperative Platooning in Mixed Traffic

    cs.AI 2026-01 conditional novelty 5.0 of 10

    A CNN-QMIX lane-change controller lifts simulated cooperative-platoon formation over MOBIL and greedy baselines and keeps working as the number of connected agents varies.

  13. Conflict Resolution under Degraded Surveillance in Air Corridors Using Multi-Agent Reinforcement Learning

    cs.LG 2026-07 reject novelty 4.0 of 10

    Separate DQN policies for UAV and eVTOL agents are trained to keep separation in a simulated structured corridor with degraded surveillance, and their behavior is summarized as action shares and Pareto-optimal safety/...

  14. Fog of Love: Engineering Virtuous Agent Behavior with Affinity-based Reinforcement Learning in a Game Environment

    cs.AI 2026-06 unverdicted novelty 4.0 of 10

    Localized affinity regularization improves multi-agent performance on both competitive and cooperative objectives in a Fog of Love environment compared to standard MADDPG.

  15. Adversarial Agent Behavior Learning in Autonomous Driving Using Deep Reinforcement Learning

    cs.CV 2025-08 reject novelty 4.0 of 10

    An adversarial car trained with a collision-based reward reliably decreases the reward of a PPO-trained ego vehicle in Highway-Env, and a robust PPO policy trained against it recovers performance.

  16. A Communication-Efficient Multi-Agent Actor-Critic Algorithm for Distributed Reinforcement Learning

    cs.LG 2019-07 unverdicted novelty 4.0 of 10

    A communication-efficient multi-agent actor-critic algorithm solves distributed RL on strongly connected directed graphs by transmitting only two scalar values per communication step.

  17. Built Environment Reasoning from Remote Sensing Imagery Using Large Vision--Language Models

    cs.CL 2026-05 unverdicted novelty 3.0 of 10

    Large vision-language models applied to multi-scale remote sensing imagery can generate recommendations on built environment design, constructability, land use, and risks for smart city decision-making.

  18. Topology-Driven Anti-Entanglement Control for Soft Robots

    cs.RO 2026-05 unverdicted novelty 3.0 of 10

    TD-MARL uses shared topological states and invariants to coordinate soft robots and reduce entanglement risk, outperforming standard DRL in simulated convergence and anti-winding performance.

Pith tools