Pith. sign in

REVIEW 3 cited by

Multi-agent Reinforcement Learning for Networked System Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.01339 v2 pith:7LZIVHUI submitted 2020-04-03 cs.LG stat.ML

classification cs.LGstat.ML
keywords controllearningmarlnetworkednmarladaptiveagentcommunication
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper considers multi-agent reinforcement learning (MARL) in networked system control. Specifically, each agent learns a decentralized control policy based on local observations and messages from connected neighbors. We formulate such a networked MARL (NMARL) problem as a spatiotemporal Markov decision process and introduce a spatial discount factor to stabilize the training of each local agent. Further, we propose a new differentiable communication protocol, called NeurComm, to reduce information loss and non-stationarity in NMARL. Based on experiments in realistic NMARL scenarios of adaptive traffic signal control and cooperative adaptive cruise control, an appropriate spatial discount factor effectively enhances the learning curves of non-communicative MARL algorithms, while NeurComm outperforms existing communication protocols in both learning efficiency and control performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards General Language-Conditioned Latent Safety Filters

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A single Hamilton-Jacobi safety filter conditioned on language constraints reduces violations in simulated pick-and-place, wiping, and stacking, with partial transfer to unseen constraint instances.

  2. Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces

    cs.MA 2026-07 conditional novelty 6.0 of 10

    CDCPG claims epsilon^-2 shared-oracle complexity to a structural stationarity floor for continuous networked MARL under exponential decay and an assumed TD-excitation condition.

  3. Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A sparse action dependency graph derived from the coordination graph is sufficient for a locally optimal policy to be globally optimal in cooperative multi-agent RL.

Pith tools