Pith. sign in

REVIEW 1 cited by

Integrating independent and centralized multi-agent reinforcement learning for traffic signal network optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.10651 v1 pith:DYUI2ZU4 submitted 2019-09-23 cs.LG cs.MAstat.ML

classification cs.LGcs.MAstat.ML
keywords trafficlearningconditionscontrolglobalmarlmulti-agentreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Traffic congestion in metropolitan areas is a world-wide problem that can be ameliorated by traffic lights that respond dynamically to real-time conditions. Recent studies applying deep reinforcement learning (RL) to optimize single traffic lights have shown significant improvement over conventional control. However, optimization of global traffic condition over a large road network fundamentally is a cooperative multi-agent control problem, for which single-agent RL is not suitable due to environment non-stationarity and infeasibility of optimizing over an exponential joint-action space. Motivated by these challenges, we propose QCOMBO, a simple yet effective multi-agent reinforcement learning (MARL) algorithm that combines the advantages of independent and centralized learning. We ensure scalability by selecting actions from individually optimized utility functions, which are shaped to maximize global performance via a novel consistency regularization loss between individual utility and a global action-value function. Experiments on diverse road topologies and traffic flow conditions in the SUMO traffic simulator show competitive performance of QCOMBO versus recent state-of-the-art MARL algorithms. We further show that policies trained on small sub-networks can effectively generalize to larger networks under different traffic flow conditions, providing empirical evidence for the suitability of MARL for intelligent traffic control.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TACTIC: Task-Agnostic Contrastive pre-Training for Inter-Agent Communication

    cs.MA 2025-01 conditional novelty 6.0 of 10

    TACTIC uses offline contrastive pretraining, aligning integrated local observations and messages with each agent's egocentric state, to improve multi-agent coordination across varied sight ranges on SMACv2.

Pith tools