Pith. sign in

REVIEW 1 cited by

Multi-Agent Reinforcement Learning for Visibility-based Persistent Monitoring

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.01129 v2 pith:6BNB6JP7 submitted 2020-11-02 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords policyagentsproblemenvironmentmulti-agentalgorithmattentioneffective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The Visibility-based Persistent Monitoring (VPM) problem seeks to find a set of trajectories (or controllers) for robots to persistently monitor a changing environment. Each robot has a sensor, such as a camera, with a limited field-of-view that is obstructed by obstacles in the environment. The robots may need to coordinate with each other to ensure no point in the environment is left unmonitored for long periods of time. We model the problem such that there is a penalty that accrues every time step if a point is left unmonitored. However, the dynamics of the penalty are unknown to us. We present a Multi-Agent Reinforcement Learning (MARL) algorithm for the VPM problem. Specifically, we present a Multi-Agent Graph Attention Proximal Policy Optimization (MA-G-PPO) algorithm that takes as input the local observations of all agents combined with a low resolution global map to learn a policy for each agent. The graph attention allows agents to share their information with others leading to an effective joint policy. Our main focus is to understand how effective MARL is for the VPM problem. We investigate five research questions with this broader goal. We find that MA-G-PPO is able to learn a better policy than the non-RL baseline in most cases, the effectiveness depends on agents sharing information with each other, and the policy learnt shows emergent behavior for the agents.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Progressive Sentences: Combining the Benefits of Word and Sentence Learning

    cs.HC 2025-07 reject novelty 5.0 of 10

    A small AR-learning study shows that inserting 4-second gaps in progressive sentence display improves recall during walking, but the claimed benefit of progressive over full-sentence presentation lacks a formal test.

Pith tools