Pith. sign in

REVIEW 3 cited by

Guided Deep Reinforcement Learning for Swarm Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1709.06011 v1 pith:IZ6FRUHF submitted 2017-09-18 cs.MA cs.AIcs.LGcs.SYeess.SYstat.ML

classification cs.MAcs.AIcs.LGcs.SYeess.SYstat.ML
keywords learningagentscriticgrouponlypolicyreinforcementstate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we investigate how to learn to control a group of cooperative agents with limited sensing capabilities such as robot swarms. The agents have only very basic sensor capabilities, yet in a group they can accomplish sophisticated tasks, such as distributed assembly or search and rescue tasks. Learning a policy for a group of agents is difficult due to distributed partial observability of the state. Here, we follow a guided approach where a critic has central access to the global state during learning, which simplifies the policy evaluation problem from a reinforcement learning point of view. For example, we can get the positions of all robots of the swarm using a camera image of a scene. This camera image is only available to the critic and not to the control policies of the robots. We follow an actor-critic approach, where the actors base their decisions only on locally sensed information. In contrast, the critic is learned based on the true global state. Our algorithm uses deep reinforcement learning to approximate both the Q-function and the policy. The performance of the algorithm is evaluated on two tasks with simple simulated 2D agents: 1) finding and maintaining a certain distance to each others and 2) locating a target.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  2. Adaptive Episode Length Adjustment for Multi-agent Reinforcement Learning

    cs.MA 2025-05 conditional novelty 6.0 of 10

    AELA improves MARL training by starting with truncated episodes and lengthening them when action-entropy falls, showing gains over QMIX and VDN on SMAC and predator-prey tasks.

  3. VariAntNet: Learning Decentralized Control of Multi-Agent Systems

    cs.LG 2025-09 conditional novelty 5.0 of 10

    A learned decentralized controller with a graph-Laplacian cohesion loss gathers bearing-only ant robots up to 2.5x faster than the analytical baseline, but loses cohesion in the hardest settings.

Pith tools