REVIEW 6 cited by
Guided Deep Reinforcement Learning for Swarm Systems
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this paper, we investigate how to learn to control a group of cooperative agents with limited sensing capabilities such as robot swarms. The agents have only very basic sensor capabilities, yet in a group they can accomplish sophisticated tasks, such as distributed assembly or search and rescue tasks. Learning a policy for a group of agents is difficult due to distributed partial observability of the state. Here, we follow a guided approach where a critic has central access to the global state during learning, which simplifies the policy evaluation problem from a reinforcement learning point of view. For example, we can get the positions of all robots of the swarm using a camera image of a scene. This camera image is only available to the critic and not to the control policies of the robots. We follow an actor-critic approach, where the actors base their decisions only on locally sensed information. In contrast, the critic is learned based on the true global state. Our algorithm uses deep reinforcement learning to approximate both the Q-function and the policy. The performance of the algorithm is evaluated on two tasks with simple simulated 2D agents: 1) finding and maintaining a certain distance to each others and 2) locating a target.
Forward citations
Cited by 6 Pith papers
-
Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex
Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.
-
Concept Learning for Cooperative Multi-Agent Reinforcement Learning
CMQ is a concept bottleneck value decomposition method for cooperative MARL that claims better performance and test-time concept interventions on SMAC and LBF benchmarks.
-
Adaptive Episode Length Adjustment for Multi-agent Reinforcement Learning
AELA improves MARL training by starting with truncated episodes and lengthening them when action-entropy falls, showing gains over QMIX and VDN on SMAC and predator-prey tasks.
-
Augmenting the action space with conventions to improve multi-agent cooperation in Hanabi
Agents trained with an action space augmented by human Hanabi conventions learn faster and score higher than baseline Rainbow agents, and they cooperate better with unfamiliar partners.
-
VariAntNet: Learning Decentralized Control of Multi-Agent Systems
A learned decentralized controller with a graph-Laplacian cohesion loss gathers bearing-only ant robots up to 2.5x faster than the analytical baseline, but loses cohesion in the hardest settings.
-
Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise
LIGHT injects decision-tree-extracted human knowledge into individual intrinsic rewards and improves sparse-reward MARL performance over QMIX, VDN, QTRAN, LIIR, and MASER.
Discussion (0). Continue with ORCID to comment.