Pith. sign in

REVIEW 1 cited by

Execute Order 66: Targeted Data Poisoning for Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.00762 v2 pith:B6N75HTU submitted 2022-01-03 cs.LG cs.AIcs.CR

classification cs.LGcs.AIcs.CR
keywords learningreinforcementpoisoningcontroldatapolicytargetedaccomplish
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data poisoning for reinforcement learning has historically focused on general performance degradation, and targeted attacks have been successful via perturbations that involve control of the victim's policy and rewards. We introduce an insidious poisoning attack for reinforcement learning which causes agent misbehavior only at specific target states - all while minimally modifying a small fraction of training observations without assuming any control over policy or reward. We accomplish this by adapting a recent technique, gradient alignment, to reinforcement learning. We test our method and demonstrate success in two Atari games of varying difficulty.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Backdoors in deep RL agents can be planted by compromising a training component or by editing pretrained weights with no training data, matching training-time attack success on six Atari games.

Pith tools