REVIEW 5 cited by
Multi-agent Reinforcement Learning for Networked System Control
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper considers multi-agent reinforcement learning (MARL) in networked system control. Specifically, each agent learns a decentralized control policy based on local observations and messages from connected neighbors. We formulate such a networked MARL (NMARL) problem as a spatiotemporal Markov decision process and introduce a spatial discount factor to stabilize the training of each local agent. Further, we propose a new differentiable communication protocol, called NeurComm, to reduce information loss and non-stationarity in NMARL. Based on experiments in realistic NMARL scenarios of adaptive traffic signal control and cooperative adaptive cruise control, an appropriate spatial discount factor effectively enhances the learning curves of non-communicative MARL algorithms, while NeurComm outperforms existing communication protocols in both learning efficiency and control performance.
Forward citations
Cited by 5 Pith papers
-
Towards General Language-Conditioned Latent Safety Filters
A single Hamilton-Jacobi safety filter conditioned on language constraints reduces violations in simulated pick-and-place, wiping, and stacking, with partial transfer to unseen constraint instances.
-
Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces
CDCPG claims epsilon^-2 shared-oracle complexity to a structural stationarity floor for continuous networked MARL under exponential decay and an assumed TD-excitation condition.
-
Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning
A sparse action dependency graph derived from the coordination graph is sufficient for a locally optimal policy to be globally optimal in cooperative multi-agent RL.
-
Dynamic Graph Communication for Decentralised Multi-Agent Reinforcement Learning
A GAT-based aggregator and a learned iteration controller improve NetMon's decentralized packet routing in simulated dynamic networks by 9.5% reward while using 6.4% less communication.
-
Achieving Collective Welfare in Multi-Agent Reinforcement Learning via Suggestion Sharing
A suggestion-sharing MARL algorithm lets agents exchange optimized action proposals for each other, with a theoretical bound relating the surrogate objective to collective return.
Discussion (0). Continue with ORCID to comment.