Pith. sign in

REVIEW 1 cited by

R-MADDPG for Partially Observable Environments and Limited Communication

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.06684 v2 pith:SGYGQUZF submitted 2020-02-16 cs.MA cs.AI

R-MADDPG for Partially Observable Environments and Limited Communication

classification cs.MA cs.AI
keywords communicationmultiagentagentslimitedobservableresourcecoordinationframework
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

There are several real-world tasks that would benefit from applying multiagent reinforcement learning (MARL) algorithms, including the coordination among self-driving cars. The real world has challenging conditions for multiagent learning systems, such as its partial observable and nonstationary nature. Moreover, if agents must share a limited resource (e.g. network bandwidth) they must all learn how to coordinate resource use. This paper introduces a deep recurrent multiagent actor-critic framework (R-MADDPG) for handling multiagent coordination under partial observable set-tings and limited communication. We investigate recurrency effects on performance and communication use of a team of agents. We demonstrate that the resulting framework learns time dependencies for sharing missing observations, handling resource limitations, and developing different communication patterns among agents.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Value-Aware Prediction for Robust Multi-Agent Coordination Under Communication Loss

    cs.MA 2026-07 conditional novelty 4.0

    Advantage-weighted NLL for observation imputation in MARL prevents performance collapse under high communication loss in 3 of 5 MPE tasks.