Pith. sign in

Paper Citation Record · LEDGER

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 3 inbound Pith citation observations for arXiv:2507.00485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00485 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:20:29.473633Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:11:29.709863Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:07:28.391121Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 811bcd8d-b9a8-448b-8aae-9c2e93087fd6 · outbound

This paper cites Constrained policy optimiza- tion.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Constrained policy optimiza- tion

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.673697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.382296Z digest=sha256:904dda953732b719e34de530e69a1a3007ba0c9a1f3aff21e52232aa91254851

Observation 7d838393-106d-4239-9e0d-dc73312485de · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.412845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.412845Z digest=sha256:9f8ae50cf4455aa72613d3426c7c8809694896561070dd0757a4eb0cfc8b20c9

Observation 0b917f32-af8c-465d-89e6-100f6087d487 · outbound

This paper cites Enhancing the robustness of qmix against state-adversarial attacks.Neurocomputing, 572:127191,.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Enhancing the robustness of qmix against state-adversarial attacks.Neurocomputing, 572:127191,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.674761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.422276Z digest=sha256:cb1c81493774adb2bd10d35a3eb6662ee8865b4ad67fdb3693bcbab798e2912f

Observation 2cb0052d-2385-4b95-8162-4339f3f95b34 · outbound

This paper cites Robust training in multiagent deep reinforcement learning against optimal adversary.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Robust training in multiagent deep reinforcement learning against optimal adversary

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.665359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.425982Z digest=sha256:c01ec21e1ba8fd74fe17ee93e2995d56000b4de7211d62e8b9042e0201879ef7

Observation cb575ed0-9d50-41b5-bbc9-64257d2373f2 · outbound

This paper cites Backdoor attacks on safe reinforcement learning- enabled cyber–physical systems.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Backdoor attacks on safe reinforcement learning- enabled cyber–physical systems

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.632972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.435260Z digest=sha256:e53420f698fdc47febe2c5d5fc4778d8320ee13ff6c49c23059faedb8e68289c

Observation a9bf24f5-be90-40d1-9626-508c8ce35f1a · outbound

This paper cites Trojdrl: Evaluation of back- door attacks on deep reinforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Trojdrl: Evaluation of back- door attacks on deep reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.621937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.437848Z digest=sha256:af24d8a2f7e46078c9f00426fedda1feb15b638fef764c55dddd3bf2f9801339

Observation 95836e9a-628d-41a0-afd8-5ef045ce71b3 · outbound

This paper cites Con- strained variational policy optimization for safe reinforce- ment learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Con- strained variational policy optimization for safe reinforce- ment learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.611753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.444969Z digest=sha256:3d086fed67454f2655be80491a020ac18a42d7a1da885e75c27be7df56bae5e1

Observation 7b450561-be08-4f34-a4e9-55a1dec67e0a · outbound

This paper cites Towards deep learning models resistant to adversarial attacks.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Towards deep learning models resistant to adversarial attacks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.601839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.448131Z digest=sha256:f9e2ee8bbc4a89ab51b045eafddaa34ce301eb46d2a6dc14f2bc78468a3b5365

Observation dd5e120c-00f2-4d39-b49c-0cc62563a593 · outbound

This paper cites Marl sim2real transfer: Merging physical reality with digital virtuality in meta- verse.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Marl sim2real transfer: Merging physical reality with digital virtuality in meta- verse

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.591243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.451164Z digest=sha256:9aa47ed239c168df2994f91f2b86659a9299cc8decf53a82984d7f400dbfae05

Observation 7b95e5ce-3118-4800-b57f-0e94d1ee0e4c · outbound

This paper cites Responsive safety in reinforcement learn- ing by PID lagrangian methods.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Responsive safety in reinforcement learn- ing by PID lagrangian methods

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.580573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.454841Z digest=sha256:50c302454044b66d21798f8d9721c4ba613d62b6a7c5183895c06da1ab4047db

Observation 23b3a5a9-d5fa-49ff-967e-d5a4953c7e70 · outbound

This paper cites Mankowitz, and Shie Mannor.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Mankowitz, and Shie Mannor

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.570764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.458015Z digest=sha256:4ce2a4499efa95d52f396f9ecfcf4650801795f6fc3f7a22956a7b171b3b292b

Observation 318e91ba-365c-41eb-aefa-2ba262f9fbc5 · outbound

This paper cites Backdoorl: Backdoor attack against competitive reinforcement learn- ing.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Backdoorl: Backdoor attack against competitive reinforcement learn- ing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.561421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.461226Z digest=sha256:9caf98b9dc398199ecaf067c51e4991f1a6aa6532fd96454bd1f2c6d493990e5

Observation 6c2d81ce-7fbd-4c6a-97b1-bac205a7e6c5 · outbound

This paper cites Partially observable mean field multi- agent reinforcement learning based on graph attention net- work for uav swarms.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Partially observable mean field multi- agent reinforcement learning based on graph attention net- work for uav swarms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.552002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.466839Z digest=sha256:d7870d7f3c7b4ea973ea0ec518a75469f202899d52861f8d3dac2782cd57b7eb

Observation ca4b967f-7f06-4d96-b152-23f71921dd26 · outbound

This paper cites First order constrained optimization in policy space.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning First order constrained optimization in policy space

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.542180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.469695Z digest=sha256:7f092b48810f810649d9952f3b13aeccba2427589f1e4491fd1135ea1ccb1aa8

Observation 97f76f8e-0a90-4f66-9a50-aacb331ddb06 · outbound

This paper cites A robust mean-field actor-critic rein- forcement learning against adversarial perturbations on agent states.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning A robust mean-field actor-critic rein- forcement learning against adversarial perturbations on agent states

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.531670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.473633Z digest=sha256:40204dade118369106595f67df05b6a9f3bfa38a4ff2e5e62933e636695dd675

Observation 3ebc93a3-82df-491e-a0d5-d1d7ab5a2b9c · outbound

This paper cites Safety gymna- sium: A unified safe reinforcement learning benchmark.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Safety gymna- sium: A unified safe reinforcement learning benchmark

Reference 1994

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.642657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.431771Z digest=sha256:32c50c27617fda4e037f1ff9ee16f89df153865915e6e05781a15645793e315b

Observation 6f5a24a6-b521-4a36-b29c-d2e8c3d88137 · outbound

This paper cites Constrained policy optimiza- tion via bayesian world models.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Constrained policy optimiza- tion via bayesian world models

Reference 1998

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.651583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.388852Z digest=sha256:021e551d9edd873cc124ee04b5daef3bbe30053c3bf81662bcf1ed51b0ebbfb1

Observation c2955608-c556-44a5-96dd-e1e86def435f · outbound

This paper cites Context-aware safe reinforcement learning for non- stationary environments.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Context-aware safe reinforcement learning for non- stationary environments

Reference 2005

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.618862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.399212Z digest=sha256:4b26388bcc705bb9f633f58c6dbf644914751497a885004244c7dddf99cdc2c7

Observation 0b3580d5-cdf5-45dd-8836-748dda812fe9 · outbound

This paper cites an unresolved cited work.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Unresolved cited work

Reference 2012

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:20:30.629309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.396002Z digest=sha256:e71c1188e4588f2aefd7d0ba53ea6dbce5fa51df6cb1233af4e1f251e446ad35

Observation b5722f24-2869-49cc-9651-df085326ca0d · outbound

This paper cites Policycleanse: Backdoor detection and mitiga- tion for competitive reinforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Policycleanse: Backdoor detection and mitiga- tion for competitive reinforcement learning

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.686052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.419570Z digest=sha256:6dae7619b4a5da96617ee5a6dd3806e30676155b37431d60800c6ae56c3b09ac

Observation 67b1924f-a84d-4766-b9f7-3a45409974a0 · outbound

This paper cites Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.662359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.385947Z digest=sha256:223b830f702a9a5005a4a92f1bf535ba072d1ab67165f3b1ad6df07c8deee3d6

Observation 66f54b19-7271-4baa-84ea-a223e1c319cc · outbound

This paper cites Badrl: Sparse targeted backdoor attack against reinforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Badrl: Sparse targeted backdoor attack against reinforcement learning

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.585085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.409436Z digest=sha256:01a1f19167cfe5ce3149d95740d842a708aded649e66dd0c5ebe4be371a75873

Observation 2eaad8fe-ae38-46ee-bc78-4adf6c952ebe · outbound

This paper cites Goodfellow, Jonathon Shlens, and Christian Szegedy.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Goodfellow, Jonathon Shlens, and Christian Szegedy

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.695760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.416677Z digest=sha256:a962171f2249227ebe7e9b67b486de4ff58eaf5298ef44484104a61b56942b81

Observation 7f1d205e-0fd2-4481-bc3e-a208e4e89b08 · outbound

This paper cites Accelerated Primal-Dual Policy Optimization for Safe Reinforcement Learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Accelerated Primal-Dual Policy Optimization for Safe Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.441050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.441050Z digest=sha256:3881bf412176d2771f468638f4b58d2794757e54df5d7e3a2bc2a031c45fd573

Observation c45262ad-fe4e-41fd-b2cf-23cf9243c6ec · outbound

This paper cites Projection-Based Constrained Policy Optimization.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Projection-Based Constrained Policy Optimization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.464007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.464007Z digest=sha256:39910f2630ecd17daeb1822fb8bf1d4b864075f2d27334fd8207be14a3c16806

Observation 07ac8b65-1c42-4f9a-994f-04fea91ff6ba · outbound

This paper cites An online actor–critic algorithm with function approximation for constrained markov decision processes.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning An online actor–critic algorithm with function approximation for constrained markov decision processes

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.639990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.392517Z digest=sha256:1f083b8ef2b63e79148373ec4d8ea6552fa74d6f727f0a951b6bac8289a21f69

Observation 3a957276-6d76-4cf5-89eb-9270ba8b54bd · outbound

This paper cites Robust multi- agent reinforcement learning method based on adversar- ial domain randomization for real-world dual-uav co- operation.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Robust multi- agent reinforcement learning method based on adversar- ial domain randomization for real-world dual-uav co- operation

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.607737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.403636Z digest=sha256:8f59323f69ae2b84fd44a2e9f5c550855f6333a591d83c0ef3cd734f986cef30

Observation f538e655-37da-4bcc-b5d4-43d5ddcaee51 · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Risk-constrained reinforcement learning with percentile risk criteria

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.597019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.406412Z digest=sha256:bd494c7294ebebf2f299310a37104d2e9e71a62f41827047ed51e390abab6f72

Observation a0ef48ee-e03a-45d8-917b-b50a1fb35b02 · outbound

This paper cites Consideration of risk in re- inforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Consideration of risk in re- inforcement learning

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.653990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:20:29.429077Z digest=sha256:d68e136f4dcacaed33d1e790cc09f93cddd1303dd54ee12ebe003fcac5c2e2d5

Pith citing papers

Observation e9f5d610-b76f-44a1-8cda-dbe946532e22 · inbound

Dataset Poisoning Attacks on Behavioral Cloning Policies cites this paper.

Dataset Poisoning Attacks on Behavioral Cloning Policies PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:11:29.709863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:11:29.709863Z digest=sha256:60e20d86dd55127120660a73a800a6d1878d23a6f19389e595aae0c2e1a85556

Observation ed111060-4680-4267-9b22-ee791dc26641 · inbound

Trojan Attacks on Neural Network Controllers for Robotic Systems cites this paper.

Trojan Attacks on Neural Network Controllers for Robotic Systems PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:24:01.980022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:24:01.980022Z digest=sha256:dea82474fca0cf71eb2864f12d38ee2be4635ed85d75fa1dbe90782138f2abc3

Observation a0217317-2144-4599-956a-f6d5e1852ac3 · inbound

Safe-RULE: Safe Reinforcement UnLEarning cites this paper.

Safe-RULE: Safe Reinforcement UnLEarning PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:07:28.392658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T17:28:09.686163Z digest=sha256:2af83078fc810a419ca4e77654537105800746a4ab8a80d51d698b9283e93687