REVIEW 1 cited by
Vulnerability of Deep Reinforcement Learning to Policy Induction Attacks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Deep learning classifiers are known to be inherently vulnerable to manipulation by intentionally perturbed inputs, named adversarial examples. In this work, we establish that reinforcement learning techniques based on Deep Q-Networks (DQNs) are also vulnerable to adversarial input perturbations, and verify the transferability of adversarial examples across different DQN models. Furthermore, we present a novel class of attacks based on this vulnerability that enable policy manipulation and induction in the learning process of DQNs. We propose an attack mechanism that exploits the transferability of adversarial examples to implement policy induction attacks on DQNs, and demonstrate its efficacy and impact through experimental study of a game-learning scenario.
Forward citations
Cited by 1 Pith paper
-
AdvIRL: Reinforcement Learning-Based Adversarial Attacks on 3D NeRF Models
AdvIRL uses PPO to adjust Instant-NGP parameters so that CLIP misclassifies rendered 3D objects, with results on banana, truck, horse, and lighthouse scenes.
Discussion (0). Continue with ORCID to comment.