REVIEW 3 cited by
Multi-agent Reinforcement Learning for Networked System Control
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper considers multi-agent reinforcement learning (MARL) in networked system control. Specifically, each agent learns a decentralized control policy based on local observations and messages from connected neighbors. We formulate such a networked MARL (NMARL) problem as a spatiotemporal Markov decision process and introduce a spatial discount factor to stabilize the training of each local agent. Further, we propose a new differentiable communication protocol, called NeurComm, to reduce information loss and non-stationarity in NMARL. Based on experiments in realistic NMARL scenarios of adaptive traffic signal control and cooperative adaptive cruise control, an appropriate spatial discount factor effectively enhances the learning curves of non-communicative MARL algorithms, while NeurComm outperforms existing communication protocols in both learning efficiency and control performance.
Forward citations
Cited by 3 Pith papers
-
Towards General Language-Conditioned Latent Safety Filters
A single Hamilton-Jacobi safety filter conditioned on language constraints reduces violations in simulated pick-and-place, wiping, and stacking, with partial transfer to unseen constraint instances.
-
Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces
CDCPG claims epsilon^-2 shared-oracle complexity to a structural stationarity floor for continuous networked MARL under exponential decay and an assumed TD-excitation condition.
-
Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning
A sparse action dependency graph derived from the coordination graph is sufficient for a locally optimal policy to be globally optimal in cooperative multi-agent RL.
Discussion (0). Continue with ORCID to comment.