Pith. sign in

REVIEW 1 cited by

From Explicit Communication to Tacit Cooperation:A Novel Paradigm for Cooperative MARL

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.14656 v1 pith:GZKCZV54 submitted 2023-04-28 cs.MA cs.AIcs.LG

classification cs.MAcs.AIcs.LG
keywords informationcommunicationcooperationparadigmtrainingagentscommunicatedcooperative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Centralized training with decentralized execution (CTDE) is a widely-used learning paradigm that has achieved significant success in complex tasks. However, partial observability issues and the absence of effectively shared signals between agents often limit its effectiveness in fostering cooperation. While communication can address this challenge, it simultaneously reduces the algorithm's practicality. Drawing inspiration from human team cooperative learning, we propose a novel paradigm that facilitates a gradual shift from explicit communication to tacit cooperation. In the initial training stage, we promote cooperation by sharing relevant information among agents and concurrently reconstructing this information using each agent's local trajectory. We then combine the explicitly communicated information with the reconstructed information to obtain mixed information. Throughout the training process, we progressively reduce the proportion of explicitly communicated information, facilitating a seamless transition to fully decentralized execution without communication. Experimental results in various scenarios demonstrate that the performance of our method without communication can approaches or even surpasses that of QMIX and communication-based methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tacit Learning with Adaptive Information Selection for Cooperative Multi-Agent Reinforcement Learning

    cs.MA 2024-12 conditional novelty 5.0 of 10

    SICA combines selective state-space filtering with attention-based training-time communication and a regeneration module to let MARL agents coordinate without messages at execution time.

Pith tools