Pith. sign in

Paper Citation Record · LEDGER

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 3 inbound Pith citation observations for arXiv:2507.00485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00485 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:20:29.473633Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:11:29.709863Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:07:28.391121Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 811bcd8d-b9a8-448b-8aae-9c2e93087fd6 · outbound

This paper cites Constrained policy optimiza- tion.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Constrained policy optimiza- tion

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.673697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.382296Z digest=sha256:4490f75565d1c50639241e47900f95471b18e7e9f890ba49d3d044ddbe915ce4

Observation 7d838393-106d-4239-9e0d-dc73312485de · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.412845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.412845Z digest=sha256:9f8ae50cf4455aa72613d3426c7c8809694896561070dd0757a4eb0cfc8b20c9

Observation 0b917f32-af8c-465d-89e6-100f6087d487 · outbound

This paper cites Enhancing the robustness of qmix against state-adversarial attacks.Neurocomputing, 572:127191,.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Enhancing the robustness of qmix against state-adversarial attacks.Neurocomputing, 572:127191,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.674761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.422276Z digest=sha256:0a0ee957dedd63a8432d576f245f91bdaa08b0ea293ebd6f05ea333331f0aa14

Observation 2cb0052d-2385-4b95-8162-4339f3f95b34 · outbound

This paper cites Robust training in multiagent deep reinforcement learning against optimal adversary.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Robust training in multiagent deep reinforcement learning against optimal adversary

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.665359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.425982Z digest=sha256:d25a8ccbe86d7dd0ec470661be5749e27e41e738e3ba3e6bc0933a64c505d998

Observation cb575ed0-9d50-41b5-bbc9-64257d2373f2 · outbound

This paper cites Backdoor attacks on safe reinforcement learning- enabled cyber–physical systems.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Backdoor attacks on safe reinforcement learning- enabled cyber–physical systems

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.632972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.435260Z digest=sha256:3ba7b7b3b6f9fb1af80f20e66bf981b617a93cb78da2caa0b5fca17bde02a9fd

Observation a9bf24f5-be90-40d1-9626-508c8ce35f1a · outbound

This paper cites Trojdrl: Evaluation of back- door attacks on deep reinforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Trojdrl: Evaluation of back- door attacks on deep reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.621937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.437848Z digest=sha256:7bdf1cc56e4747a6b4c15817b4d52cdceb4440ba30169e9777d41782c3af1ebb

Observation 95836e9a-628d-41a0-afd8-5ef045ce71b3 · outbound

This paper cites Con- strained variational policy optimization for safe reinforce- ment learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Con- strained variational policy optimization for safe reinforce- ment learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.611753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.444969Z digest=sha256:c98f83db3e6469e22c3e57bad6c0ba91d03132c37be2f447a2e1a3ff8df2f120

Observation 7b450561-be08-4f34-a4e9-55a1dec67e0a · outbound

This paper cites Towards deep learning models resistant to adversarial attacks.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Towards deep learning models resistant to adversarial attacks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.601839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.448131Z digest=sha256:84a601052519a79e8de5a603ff19b9cbec4e280bea1cd30b288d41d5024444e9

Observation dd5e120c-00f2-4d39-b49c-0cc62563a593 · outbound

This paper cites Marl sim2real transfer: Merging physical reality with digital virtuality in meta- verse.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Marl sim2real transfer: Merging physical reality with digital virtuality in meta- verse

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.591243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.451164Z digest=sha256:cabb206ef74d703ce9edd32ca4ac26cb9e890185646d71eaeba1800e7e248a3d

Observation 7b95e5ce-3118-4800-b57f-0e94d1ee0e4c · outbound

This paper cites Responsive safety in reinforcement learn- ing by PID lagrangian methods.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Responsive safety in reinforcement learn- ing by PID lagrangian methods

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.580573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.454841Z digest=sha256:170b312d42834799ee00ebcce858070dc177bec708a981f05ebf65fba2fd0e37

Observation 23b3a5a9-d5fa-49ff-967e-d5a4953c7e70 · outbound

This paper cites Mankowitz, and Shie Mannor.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Mankowitz, and Shie Mannor

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.570764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.458015Z digest=sha256:64cc35205dfe4866498578235665226e2d307923d431f806828db033564efc1b

Observation 318e91ba-365c-41eb-aefa-2ba262f9fbc5 · outbound

This paper cites Backdoorl: Backdoor attack against competitive reinforcement learn- ing.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Backdoorl: Backdoor attack against competitive reinforcement learn- ing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.561421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.461226Z digest=sha256:96663f6cac6ad34ec7712dc0c2795abe502e757b3138ab05b884a926ef483086

Observation 6c2d81ce-7fbd-4c6a-97b1-bac205a7e6c5 · outbound

This paper cites Partially observable mean field multi- agent reinforcement learning based on graph attention net- work for uav swarms.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Partially observable mean field multi- agent reinforcement learning based on graph attention net- work for uav swarms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.552002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.466839Z digest=sha256:781735a8d2a54a90efba35f8597baa51261fb48e1efecc959bd1e465ecb55233

Observation ca4b967f-7f06-4d96-b152-23f71921dd26 · outbound

This paper cites First order constrained optimization in policy space.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning First order constrained optimization in policy space

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.542180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.469695Z digest=sha256:d4ee85ed3e842cc84a5b0db6f151096e0a35807044442c454aa6f55a138c8fea

Observation 97f76f8e-0a90-4f66-9a50-aacb331ddb06 · outbound

This paper cites A robust mean-field actor-critic rein- forcement learning against adversarial perturbations on agent states.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning A robust mean-field actor-critic rein- forcement learning against adversarial perturbations on agent states

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.531670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.473633Z digest=sha256:eb2b93a7430d1ff89af5ce051526409d385a5a21574a51b3c31a927569b1a2cc

Observation 3ebc93a3-82df-491e-a0d5-d1d7ab5a2b9c · outbound

This paper cites Safety gymna- sium: A unified safe reinforcement learning benchmark.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Safety gymna- sium: A unified safe reinforcement learning benchmark

Reference 1994

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.642657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.431771Z digest=sha256:9947ce0b4721a327e7d216d7a4bc7bdf14647179216f822bbc00181af623478a

Observation 6f5a24a6-b521-4a36-b29c-d2e8c3d88137 · outbound

This paper cites Constrained policy optimiza- tion via bayesian world models.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Constrained policy optimiza- tion via bayesian world models

Reference 1998

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.651583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.388852Z digest=sha256:610d1337b0fc5439ab339d708bb5783785de9b53c81d0f375d6fca16b18352fc

Observation c2955608-c556-44a5-96dd-e1e86def435f · outbound

This paper cites Context-aware safe reinforcement learning for non- stationary environments.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Context-aware safe reinforcement learning for non- stationary environments

Reference 2005

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.618862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.399212Z digest=sha256:f946c2e25ccf6e4e57d17abc5bee1617a2989dbb931a79d8b5a4e8fbe9fa7a83

Observation 0b3580d5-cdf5-45dd-8836-748dda812fe9 · outbound

This paper cites an unresolved cited work.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Unresolved cited work

Reference 2012

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:20:30.629309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.396002Z digest=sha256:914b789c069c6845d8fdc50c79c99f0dea54140487c9799f5d074fee7eef20af

Observation b5722f24-2869-49cc-9651-df085326ca0d · outbound

This paper cites Policycleanse: Backdoor detection and mitiga- tion for competitive reinforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Policycleanse: Backdoor detection and mitiga- tion for competitive reinforcement learning

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.686052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.419570Z digest=sha256:26d67e3dfca4c53a8a7f8d84b5414f23a4cbe348c44b00797e33b157bbac8476

Observation 67b1924f-a84d-4766-b9f7-3a45409974a0 · outbound

This paper cites Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.662359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.385947Z digest=sha256:f0ebd2f93a75eddde52ca2999327b65661685c205c1febd7e3ed8e9aa114f6cd

Observation 66f54b19-7271-4baa-84ea-a223e1c319cc · outbound

This paper cites Badrl: Sparse targeted backdoor attack against reinforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Badrl: Sparse targeted backdoor attack against reinforcement learning

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.585085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.409436Z digest=sha256:65becf63a0421ed4c1946978227207437320197e8ad7e85ce8b1e492fc636bd3

Observation 2eaad8fe-ae38-46ee-bc78-4adf6c952ebe · outbound

This paper cites Goodfellow, Jonathon Shlens, and Christian Szegedy.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Goodfellow, Jonathon Shlens, and Christian Szegedy

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.695760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.416677Z digest=sha256:67c0cd4b619b462fa53c0dc394a28437c9b4d20a1177c7c1ed7725c131cd860f

Observation 7f1d205e-0fd2-4481-bc3e-a208e4e89b08 · outbound

This paper cites Accelerated Primal-Dual Policy Optimization for Safe Reinforcement Learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Accelerated Primal-Dual Policy Optimization for Safe Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.441050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.441050Z digest=sha256:3881bf412176d2771f468638f4b58d2794757e54df5d7e3a2bc2a031c45fd573

Observation c45262ad-fe4e-41fd-b2cf-23cf9243c6ec · outbound

This paper cites Projection-Based Constrained Policy Optimization.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Projection-Based Constrained Policy Optimization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.464007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.464007Z digest=sha256:39910f2630ecd17daeb1822fb8bf1d4b864075f2d27334fd8207be14a3c16806

Observation 07ac8b65-1c42-4f9a-994f-04fea91ff6ba · outbound

This paper cites An online actor–critic algorithm with function approximation for constrained markov decision processes.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning An online actor–critic algorithm with function approximation for constrained markov decision processes

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.639990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.392517Z digest=sha256:1c982498f8d692ed61c093232327f4599f315bacc36ca71cfaf380842416ce5f

Observation 3a957276-6d76-4cf5-89eb-9270ba8b54bd · outbound

This paper cites Robust multi- agent reinforcement learning method based on adversar- ial domain randomization for real-world dual-uav co- operation.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Robust multi- agent reinforcement learning method based on adversar- ial domain randomization for real-world dual-uav co- operation

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.607737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.403636Z digest=sha256:48af8eb374225351e43af81812884a36e964dd63248a732efbf88d4a21531d22

Observation f538e655-37da-4bcc-b5d4-43d5ddcaee51 · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Risk-constrained reinforcement learning with percentile risk criteria

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.597019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.406412Z digest=sha256:70369cd3e47b3e140c86986a9efd29689fecd03dada67b56d7d8e0c63d724395

Observation a0ef48ee-e03a-45d8-917b-b50a1fb35b02 · outbound

This paper cites Consideration of risk in re- inforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Consideration of risk in re- inforcement learning

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.653990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:20:29.429077Z digest=sha256:56fd82b57279f8cc652926a68f10b775c5697bceb7629b485a84f42c7304a651

Pith citing papers

Observation e9f5d610-b76f-44a1-8cda-dbe946532e22 · inbound

Dataset Poisoning Attacks on Behavioral Cloning Policies cites this paper.

Dataset Poisoning Attacks on Behavioral Cloning Policies PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:11:29.709863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:11:29.709863Z digest=sha256:60e20d86dd55127120660a73a800a6d1878d23a6f19389e595aae0c2e1a85556

Observation ed111060-4680-4267-9b22-ee791dc26641 · inbound

Trojan Attacks on Neural Network Controllers for Robotic Systems cites this paper.

Trojan Attacks on Neural Network Controllers for Robotic Systems PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:24:01.980022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:24:01.980022Z digest=sha256:dea82474fca0cf71eb2864f12d38ee2be4635ed85d75fa1dbe90782138f2abc3

Observation a0217317-2144-4599-956a-f6d5e1852ac3 · inbound

Safe-RULE: Safe Reinforcement UnLEarning cites this paper.

Safe-RULE: Safe Reinforcement UnLEarning PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:07:28.392658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:28:09.686163Z digest=sha256:3756d3c4f1fffb0a9414566dde4e665f695ad41409564fda2fa4af96535c5f14