REVIEW 4 cited by
Execute Order 66: Targeted Data Poisoning for Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Data poisoning for reinforcement learning has historically focused on general performance degradation, and targeted attacks have been successful via perturbations that involve control of the victim's policy and rewards. We introduce an insidious poisoning attack for reinforcement learning which causes agent misbehavior only at specific target states - all while minimally modifying a small fraction of training observations without assuming any control over policy or reward. We accomplish this by adapting a recent technique, gradient alignment, to reinforcement learning. We test our method and demonstrate success in two Atari games of varying difficulty.
Forward citations
Cited by 4 Pith papers
-
Provably Efficient Action-Manipulation Attack Against Continuous Reinforcement Learning
LCBT is a trajectory-only tree-search attack that claims to steer continuous-action RL agents to target policies with sublinear attack cost, but the proof of the claim has a serious importance-sampling flaw.
-
Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning
Backdoors in deep RL agents can be planted by compromising a training component or by editing pretrained weights with no training data, matching training-time attack success on six Atari games.
-
Online Poisoning Attack Against Reinforcement Learning under Black-box Environments
An online attacker knowing only which states are reachable can poison rewards and transitions to make a Q-learning agent follow a target policy in a maze.
-
Trading Devil RL: Backdoor attack via Stock market, Bayesian Optimization and Reinforcement Learning
A data-poisoning backdoor attack on audio transformers is claimed with 100 percent success on TIMIT, but the paper provides no reproducible derivation or evaluation.
Discussion (0). Continue with ORCID to comment.