Pith. sign in

REVIEW 3 cited by

Online Robustness Training for Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.00887 v3 pith:XLTD2EFN submitted 2019-11-03 cs.LG stat.ML

classification cs.LGstat.ML
keywords trainingadversarialagentattacksdeeplearningonlinereinforcement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In deep reinforcement learning (RL), adversarial attacks can trick an agent into unwanted states and disrupt training. We propose a system called Robust Student-DQN (RS-DQN), which permits online robustness training alongside Q networks, while preserving competitive performance. We show that RS-DQN can be combined with (i) state-of-the-art adversarial training and (ii) provably robust training to obtain an agent that is resilient to strong attacks during training and evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization

    cs.LG 2025-07 conditional novelty 6.0 of 10

    OA-PI trains RL policies against the worst-case bounded action perturbation through a modified Bellman operator, and its TD3 and PPO variants improve robustness and nominal performance on Mujoco and Box2d tasks.

  2. Towards Robust Deep Reinforcement Learning against Environmental State Perturbation

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Environmental perturbations to the initial state sharply reduce the rewards of PPO-trained Overcooked agents, and the proposed BAT defense, supervised kickstarting followed by adversarial fine-tuning, restores robustn...

  3. Advancing Robustness in Deep Reinforcement Learning with an Ensemble Defense Approach

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Averaging three observation filters (random noise, autoencoder, PCA) before action selection substantially improves a Highway-env DQN's reward and collision rate under FGSM attacks.

Pith tools