Pith. sign in

REVIEW 1 cited by

AC2C: Adaptively Controlled Two-Hop Communication for Multi-Agent Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.12515 v2 pith:QQBEVYXK submitted 2023-02-24 cs.MA cs.AI

classification cs.MAcs.AI
keywords communicationac2ctwo-hopagentagentslearningmulti-agentadaptive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning communication strategies in cooperative multi-agent reinforcement learning (MARL) has recently attracted intensive attention. Early studies typically assumed a fully-connected communication topology among agents, which induces high communication costs and may not be feasible. Some recent works have developed adaptive communication strategies to reduce communication overhead, but these methods cannot effectively obtain valuable information from agents that are beyond the communication range. In this paper, we consider a realistic communication model where each agent has a limited communication range, and the communication topology dynamically changes. To facilitate effective agent communication, we propose a novel communication protocol called Adaptively Controlled Two-Hop Communication (AC2C). After an initial local communication round, AC2C employs an adaptive two-hop communication strategy to enable long-range information exchange among agents to boost performance, which is implemented by a communication controller. This controller determines whether each agent should ask for two-hop messages and thus helps to reduce the communication overhead during distributed execution. We evaluate AC2C on three cooperative multi-agent tasks, and the experimental results show that it outperforms relevant baselines with lower communication costs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dynamic Graph Communication for Decentralised Multi-Agent Reinforcement Learning

    cs.MA 2024-12 conditional novelty 6.0 of 10

    A GAT-based aggregator and a learned iteration controller improve NetMon's decentralized packet routing in simulated dynamic networks by 9.5% reward while using 6.4% less communication.

Pith tools