Pith. sign in

Paper Citation Record · LEDGER

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System

As of 11 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2501.13727.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13727 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:42:33.315453Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T01:58:48.241542Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3c110ce-d129-4def-ab48-464b8d88cc69 · outbound

This paper cites Constrained policy optimization.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Constrained policy optimization

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.709776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.202823Z digest=sha256:164867859ff4735cb958661b5a8e1e17a7621d45f82eb13183aa516e490efb19

Observation 247f7d00-5d7f-4ef9-8468-e4f1f4da5963 · outbound

This paper cites Learning transferable cooperative behavior in multi-agent teams.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Learning transferable cooperative behavior in multi-agent teams

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.696572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.207399Z digest=sha256:d0315722edf23d98498b59c37b81b7ec66bf56412f0a519701f17744c61441f2

Observation fecf5c0e-12be-4748-80b6-2d97106ff844 · outbound

This paper cites Safe learning in robotics: From learning-based control to safe reinforcement learning.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Safe learning in robotics: From learning-based control to safe reinforcement learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.679075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.211019Z digest=sha256:9a7a02191514ea0d8fa9e71f5e325284c8b9c294e720ab527d9016da9a85d0ff

Observation b51738e4-cdc8-421b-a621-f56e76a6c044 · outbound

This paper cites Drqn-based 3d obstacle avoidance with a limited field of view.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Drqn-based 3d obstacle avoidance with a limited field of view

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.665771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.214769Z digest=sha256:53c6c6f70f270932309a65068ece1ab8b2ae006fdc1c46badf0baca731b969d8

Observation 62b65941-6f06-44f2-9c4e-99f3c5b48195 · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Safe rlhf: Safe reinforcement learning from human feedback

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.652865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.218775Z digest=sha256:1ca110ce651096eadd8c3c701bb7448e0cf04e60e4273982e8df1faf8133a84b

Observation 31728b7e-2816-4e2a-93f3-4a31ac5411cd · outbound

This paper cites Safe multi-agent reinforcement learning for multi-robot control.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Safe multi-agent reinforcement learning for multi-robot control

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.641650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.222544Z digest=sha256:bd623ba77f8948d44e65f87a641dee0f95d4505231c28ab3acc5bee5b8aba2e1

Observation 197b62d4-6670-4e47-9843-8cd07326c57f · outbound

This paper cites Scalable communication for multi-agent reinforcement learning via transformer-based email mechanism.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Scalable communication for multi-agent reinforcement learning via transformer-based email mechanism

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.630075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.226434Z digest=sha256:cb36b1bd389a3025866198ead97837496cd31294850b3b33e93142b51a3ee120

Observation 2a51c528-284b-4c45-a2dc-a13c61f7efc5 · outbound

This paper cites Long short-term memory.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Long short-term memory

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T15:42:33.230135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:42:33.230135Z digest=sha256:27c28dbb8b7568d4b43f5c4e6a12e6f5c560ad82f3a5f8647ca33eeabfaff3e8

Observation 04850b18-228f-4e7a-a0e3-f2d44c3be83b · outbound

This paper cites Collision avoidance and navigation for a quadrotor swarm using end-to-end deep reinforcement learning.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Collision avoidance and navigation for a quadrotor swarm using end-to-end deep reinforcement learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.607387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.233955Z digest=sha256:34c2292332dfbfce23aaf833f60f2f58fb31ef642c568de5f6fbb8c049a880d0

Observation c4689a2f-b261-4ccb-aff5-082cf0b1838b · outbound

This paper cites Graph convolutional reinforcement learning.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Graph convolutional reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.595414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.237486Z digest=sha256:3fb0272a907b6385fa43708eef11fee13e98bd310ba9b2504bb58a4d57d69df6

Observation 9416445e-cbed-4af1-9b8e-cf6d3f0bc66a · outbound

This paper cites Semi-supervised classification with graph convolutional networks.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Semi-supervised classification with graph convolutional networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.583704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.240963Z digest=sha256:80fb635b804f5d0d2542302e0194e33e688eb215d2768d7b95444bf821a2bac1

Observation deccf4ed-a799-4387-bb9b-84db804f8ad2 · outbound

This paper cites Trust region policy optimisation in multi-agent reinforcement learning.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Trust region policy optimisation in multi-agent reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.571672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.244813Z digest=sha256:475d82b6ad4229f14465c9f5c37895d591220bde84a7ab334e8cfb3b4bb8490a

Observation b9e737b9-35cb-4eb4-a5d0-ef257bbd5dc6 · outbound

This paper cites Cmix: Deep multi-agent reinforcement learning with peak and average constraints.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Cmix: Deep multi-agent reinforcement learning with peak and average constraints

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.558989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.248325Z digest=sha256:68eaa629b12107eb680d8fc2454e0eaa1565ea0b64de16068e947f2d0ef0fcd4

Observation 463dc400-d9da-4470-aea6-a03698ee8d04 · outbound

This paper cites Multi-agent actor-critic for mixed cooperative-competitive environments.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Multi-agent actor-critic for mixed cooperative-competitive environments

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.547687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.251664Z digest=sha256:69fa3fd60b52de931e3206e4735df7e0501067d53f918a909c5308d395db6ed8

Observation c78f18e3-2e42-4dec-8301-4402589ead52 · outbound

This paper cites Deep learning for safe autonomous driving: Current challenges and future directions.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Deep learning for safe autonomous driving: Current challenges and future directions

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.534443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.255554Z digest=sha256:0a48ffed40452460445329a43925d28c0c7a9fb8236bed054ec1f7c33f737301

Observation 9cbd690d-4e16-42ea-8bcc-faa99229538c · outbound

This paper cites Scalable multi-agent reinforcement learning through intelligent information aggregation.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Scalable multi-agent reinforcement learning through intelligent information aggregation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.522874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.259166Z digest=sha256:bcef657368249c37f420d65b94c876f4b9a93ed825d284140876133fd50c1e28

Observation 46e0a22d-d97d-4ec0-9304-e80b70c6980d · outbound

This paper cites A concise introduction to decentralized POMDPs , volume 1.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System A concise introduction to decentralized POMDPs , volume 1

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.510938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.262802Z digest=sha256:0eb8033f89ed3a1cc1f19d0286f5c57fa753bb9d35f6ca272c0dbb994591e2f1

Observation 757a0885-56dc-49e2-b780-42628ca23c2b · outbound

This paper cites Dealing with Non-Stationarity in Multi-Agent Deep Reinforcement Learning.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Dealing with Non-Stationarity in Multi-Agent Deep Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T15:42:33.266206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:42:33.266206Z digest=sha256:5f5bd387140b024638de4fe396c176285808cbab48fb175e361e1b7331440a06

Observation 040f8e5b-c53b-4426-ba0d-511f051134f9 · outbound

This paper cites Monotonic value function factorisation for deep multi-agent reinforcement learning.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Monotonic value function factorisation for deep multi-agent reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.498727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.270098Z digest=sha256:1967b121f25ee45caa6b7d8b9c9cd2140c0f56fb713eab74ea437d24d8eec848

Observation 512aa79d-cd0c-4e7e-a916-740147faf020 · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T15:42:33.274421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:42:33.274421Z digest=sha256:e828c6205bd0b2cec2dd3cf58b0ff1908a61134a8e7c79ff6f4002ae4ac27074

Observation b96014ec-1f09-498d-aaf6-32e0bc00b692 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T15:42:33.278141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:42:33.278141Z digest=sha256:e2a75a633bce056b9b33845fac8c1c15ba1eb8f233325b6c0d0ebe59e785ed19

Observation 908368f2-f73b-4d7d-b9e8-13ab0fe776b2 · outbound

This paper cites Trust Region Policy Optimization.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Trust Region Policy Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T15:42:33.281942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:42:33.281942Z digest=sha256:2e16a45fb9b9792433855177d7fe9293abdef9f1c5e6ce25d1463a9adb708f9f

Observation a9bc3ede-9b43-4fd9-9957-e816f3560f63 · outbound

This paper cites Masked label prediction: Unified message passing model for semi-supervised classification.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Masked label prediction: Unified message passing model for semi-supervised classification

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.486300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.285636Z digest=sha256:6211ffbad0ec97160df51f054b4792074e0be19c50b8e180d6a801342a0b2c38

Observation fb0e634b-12e5-479e-862b-a4e87210f6a6 · outbound

This paper cites Learning multiagent communication with backpropagation.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Learning multiagent communication with backpropagation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.474377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.289266Z digest=sha256:e06b12a6f6f2b0bf3177274de752b247b4e531801a98f8d67d9c5fef6fcf86d7

Observation 760187a1-3036-47da-9699-d97f6bf33771 · outbound

This paper cites Relative distributed formation and obstacle avoidance with multi-agent reinforcement learning.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Relative distributed formation and obstacle avoidance with multi-agent reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.461247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.293009Z digest=sha256:91203af3336e25272481be019a68228850497511bcada0aae765ff80164e283d

Observation 3db52748-8219-4071-aa0d-1c598580f139 · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System The surprising effectiveness of ppo in cooperative multi-agent games

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.448848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.296827Z digest=sha256:431e8013d9716cbda09fa5f01d48259999f688a9c407561a237213e46a2e2072

Observation 4b4f57d8-16a7-4510-ba0d-676016dd04a3 · outbound

This paper cites Attention-based reinforcement learning for real-time uav semantic communication.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Attention-based reinforcement learning for real-time uav semantic communication

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.437154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.300375Z digest=sha256:192331a5bf3915393030d03cb319a3db229bae1462f51d55acc5591fa7b7465a

Observation 16e2c2fd-c0f7-4e7a-9a45-7870a63f5901 · outbound

This paper cites A survey of multi-agent deep reinforcement learning with communication.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System A survey of multi-agent deep reinforcement learning with communication

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.425535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.304387Z digest=sha256:180ead0eaa256f2b869d76368ae81a70268fde05f9d873d40a606e75f224daf5

Observation 60dd8506-e42c-4a1b-932b-02233dff5179 · outbound

This paper cites Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T15:42:33.307887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:42:33.307887Z digest=sha256:34c25ae27d808060eacbc5c8c5dd6350836f38a7050673744d53706e19f68601

Observation 2038acb6-4e53-4feb-8c36-cd29e2a5739a · outbound

This paper cites Settling the variance of multi-agent policy gradients.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Settling the variance of multi-agent policy gradients

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:42:33.413186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:42:33.311765Z digest=sha256:864d7309469b54f4a0b0cb8b3df1e661e0bc9f81a83373d83fe769aa98048c8f

Observation 26675881-98b1-4dee-9305-8bfe0b9012de · outbound

This paper cites write newline.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System write newline

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T15:42:33.315453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:42:33.315453Z digest=sha256:e9b4d87c3bab6cdd95b785a92a4a6dc92bde6b7c77e4d4ec6cb78548101956ac

Pith citing papers

Observation 0f220bbd-6724-4ea3-bb7c-f7b51849ea26 · inbound

High-Precision Formation Control for Heterogeneous Multi-Robot Systems via Hierarchical Hybrid Physics-Informed Deep Reinforcement Learning cites this paper.

High-Precision Formation Control for Heterogeneous Multi-Robot Systems via Hierarchical Hybrid Physics-Informed Deep Reinforcement Learning Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T01:58:48.241542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:58:48.241542Z digest=sha256:3443c4e5c73ea18a6e4db0cc4710715367ff370258fd4265d925f3c9103543ab