Pith. sign in

REVIEW 1 cited by

Safe Deep Reinforcement Learning for Multi-Agent Systems with Continuous Action Spaces

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.03952 v2 pith:FMZFCQGW submitted 2021-08-09 cs.LG cs.RO

classification cs.LGcs.RO
keywords deepmulti-agentsafetyactionconstraintslearningsoftconstraint
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-agent control problems constitute an interesting area of application for deep reinforcement learning models with continuous action spaces. Such real-world applications, however, typically come with critical safety constraints that must not be violated. In order to ensure safety, we enhance the well-known multi-agent deep deterministic policy gradient (MADDPG) framework by adding a safety layer to the deep policy network. In particular, we extend the idea of linearizing the single-step transition dynamics, as was done for single-agent systems in Safe DDPG (Dalal et al., 2018), to multi-agent settings. We additionally propose to circumvent infeasibility problems in the action correction step using soft constraints (Kerrigan & Maciejowski, 2000). Results from the theory of exact penalty functions can be used to guarantee constraint satisfaction of the soft constraints under mild assumptions. We empirically find that the soft formulation achieves a dramatic decrease in constraint violations, making safety available even during the learning procedure.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Constrained Optimization of Charged Particle Tracking with Multi-Agent Reinforcement Learning

    physics.comp-ph 2025-01 conditional novelty 5.0 of 10

    A constrained multi-agent RL tracker with a linear-assignment safety layer and cost-margin gradients improves particle track reconstruction on simulated proton-CT data.

Pith tools