Pith. sign in

REVIEW 3 cited by

Policy Smoothing for Provably Robust Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.11420 v3 pith:JYOHGSL4 submitted 2021-06-21 cs.LG

classification cs.LG
keywords policyrobustnessactionsadaptiveadversarialadversarylearningprevious
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The study of provable adversarial robustness for deep neural networks (DNNs) has mainly focused on static supervised learning tasks such as image classification. However, DNNs have been used extensively in real-world adaptive tasks such as reinforcement learning (RL), making such systems vulnerable to adversarial attacks as well. Prior works in provable robustness in RL seek to certify the behaviour of the victim policy at every time-step against a non-adaptive adversary using methods developed for the static setting. But in the real world, an RL adversary can infer the defense strategy used by the victim agent by observing the states, actions, etc., from previous time-steps and adapt itself to produce stronger attacks in future steps. We present an efficient procedure, designed specifically to defend against an adaptive RL adversary, that can directly certify the total reward without requiring the policy to be robust at each time-step. Our main theoretical contribution is to prove an adaptive version of the Neyman-Pearson Lemma -- a key lemma for smoothing-based certificates -- where the adversarial perturbation at a particular time can be a stochastic function of current and previous observations and states as well as previous actions. Building on this result, we propose policy smoothing where the agent adds a Gaussian noise to its observation at each time-step before passing it through the policy function. Our robustness certificates guarantee that the final total reward obtained by policy smoothing remains above a certain threshold, even though the actions at intermediate time-steps may change under the attack. Our experiments on various environments like Cartpole, Pong, Freeway and Mountain Car show that our method can yield meaningful robustness guarantees in practice.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fox in the Henhouse: Supply-Chain Backdoor Attacks Against Reinforcement Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A malicious pre-trained opponent can, through legitimate in-game actions, embed a trigger-activated backdoor into a victim reinforcement learning agent.

  2. Robust Behavior Cloning Via Global Lipschitz Regularization

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Imposing a global Lipschitz constraint on a behavior cloning policy provides a provable upper bound on worst-case reward loss under bounded state perturbations.

  3. Position: Certified Robustness Does Not (Yet) Imply Model Security

    cs.CR 2025-06 conditional novelty 4.0 of 10

    A certified robustness radius says nothing about whether a sample is clean or correctly predicted, so certification does not yet imply model security.

Pith tools