REVIEW 6 cited by
BACKDOORL: Backdoor Attack against Competitive Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent research has confirmed the feasibility of backdoor attacks in deep reinforcement learning (RL) systems. However, the existing attacks require the ability to arbitrarily modify an agent's observation, constraining the application scope to simple RL systems such as Atari games. In this paper, we migrate backdoor attacks to more complex RL systems involving multiple agents and explore the possibility of triggering the backdoor without directly manipulating the agent's observation. As a proof of concept, we demonstrate that an adversary agent can trigger the backdoor of the victim agent with its own action in two-player competitive RL systems. We prototype and evaluate BACKDOORL in four competitive environments. The results show that when the backdoor is activated, the winning rate of the victim drops by 17% to 37% compared to when not activated.
Forward citations
Cited by 6 Pith papers
-
State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space
A backdoor attack on vision-language-action robot policies uses the arm's initial joint configuration as the trigger, achieving >90% triggered failure with only small clean-task degradation.
-
Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning
Backdoors in deep RL agents can be planted by compromising a training component or by editing pretrained weights with no training data, matching training-time attack success on six Atari games.
-
Data Free Backdoor Attacks
DFBA injects a backdoor into a pre-trained image classifier by editing one neuron per layer and a few output weights, requiring no data or retraining.
-
Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies
The paper proposes a one-to-one mapping between six causes of distribution shift and several AI safety issues, arguing for mutual method transfer through aligned definitions.
-
Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models
ConceptGuard protects concept bottleneck models from concept-level backdoor attacks by clustering concepts, training an ensemble of sub-classifiers, and taking a majority vote.
-
Trading Devil RL: Backdoor attack via Stock market, Bayesian Optimization and Reinforcement Learning
A data-poisoning backdoor attack on audio transformers is claimed with 100 percent success on TIMIT, but the paper provides no reproducible derivation or evaluation.
Discussion (0). Continue with ORCID to comment.