Pith. sign in

REVIEW 2 cited by

Adversarial Inception Backdoor Attacks against Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.13995 v3 pith:RT2WJ7L4 submitted 2024-10-17 cs.LG cs.CR

classification cs.LGcs.CR
keywords attacksagentadversarialbackdoorconstraintsinceptionachievearbitrary
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent works have demonstrated the vulnerability of Deep Reinforcement Learning (DRL) algorithms against training-time, backdoor poisoning attacks. The objectives of these attacks are twofold: induce pre-determined, adversarial behavior in the agent upon observing a fixed trigger during deployment while allowing the agent to solve its intended task during training. Prior attacks assume arbitrary control over the agent's rewards, inducing values far outside the environment's natural constraints. This results in brittle attacks that fail once the proper reward constraints are enforced. Thus, in this work we propose a new class of backdoor attacks against DRL which are the first to achieve state of the art performance under strict reward constraints. These "inception" attacks manipulate the agent's training data -- inserting the trigger into prior observations and replacing high return actions with those of the targeted adversarial behavior. We formally define these attacks and prove they achieve both adversarial objectives against arbitrary Markov Decision Processes (MDP). Using this framework we devise an online inception attack which achieves an 100\% attack success rate on multiple environments under constrained rewards while minimally impacting the agent's task performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Backdoors in deep RL agents can be planted by compromising a training component or by editing pretrained weights with no training data, matching training-time attack success on six Atari games.

  2. ORAN-DEFEND: Subspace Detection and Sanitization of Backdoor DRL xApps in Open RAN

    cs.CR 2026-07 conditional novelty 5.0 of 10

    SVD projection onto a clean-KPI subspace recovers 100% DRL return against four O-RAN backdoor attacks whenever trigger energy lies in the orthogonal complement.

Pith tools