Backdoors in deep RL agents can be planted by compromising a training component or by editing pretrained weights with no training data, matching training-time attack success on six Atari games.
In Findings of the Association for Computational Linguistics ACL 2024, 5339–5352
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning
Backdoors in deep RL agents can be planted by compromising a training component or by editing pretrained weights with no training data, matching training-time attack success on six Atari games.