Pith. sign in

REVIEW 1 cited by

Learning Complex Multi-Agent Policies in Presence of an Adversary

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.07698 v2 pith:Q7YDGFSE submitted 2020-08-18 cs.MA

classification cs.MA
keywords learningmulti-agentagentsadversarylearnreinforcementbehaviorscomplex
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, there has been some outstanding work on applying deep reinforcement learning to multi-agent settings. Often in such multi-agent scenarios, adversaries can be present. We address the requirements of such a setting by implementing a graph-based multi-agent deep reinforcement learning algorithm. In this work, we consider the scenario of multi-agent deception in which multiple agents need to learn to cooperate and communicate in order to deceive an adversary. We have employed a two-stage learning process to get the cooperating agents to learn such deceptive behaviors. Our experiments show that our approach allows us to employ curriculum learning to increase the number of cooperating agents in the environment and enables a team of agents to learn complex behaviors to successfully deceive an adversary. Keywords: Multi-agent system, Graph neural network, Reinforcement learning

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Verbalized Bayesian Persuasion

    cs.GT 2025-02 conditional novelty 7.0 of 10

    VBP solves Bayesian persuasion in natural language by treating LLMs as sender and receiver in a mediator-augmented game and searching prompt strategies with Prompt-PSRO.

Pith tools