Pith. sign in

REVIEW 13 cited by

Uncovering the Limits of Adversarial Training against Norm-Bounded Adversarial Examples

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.03593 v3 pith:SBEEHKI5 submitted 2020-10-07 stat.ML cs.AIcs.LG

classification stat.MLcs.AIcs.LG
keywords adversarialperturbationssizetrainingaccuracyadditionalattackcifar-10
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Adversarial training and its variants have become de facto standards for learning robust deep neural networks. In this paper, we explore the landscape around adversarial training in a bid to uncover its limits. We systematically study the effect of different training losses, model sizes, activation functions, the addition of unlabeled data (through pseudo-labeling) and other factors on adversarial robustness. We discover that it is possible to train robust models that go well beyond state-of-the-art results by combining larger models, Swish/SiLU activations and model weight averaging. We demonstrate large improvements on CIFAR-10 and CIFAR-100 against $\ell_\infty$ and $\ell_2$ norm-bounded perturbations of size $8/255$ and $128/255$, respectively. In the setting with additional unlabeled data, we obtain an accuracy under attack of 65.88% against $\ell_\infty$ perturbations of size $8/255$ on CIFAR-10 (+6.35% with respect to prior art). Without additional data, we obtain an accuracy under attack of 57.20% (+3.46%). To test the generality of our findings and without any additional modifications, we obtain an accuracy under attack of 80.53% (+7.62%) against $\ell_2$ perturbations of size $128/255$ on CIFAR-10, and of 36.88% (+8.46%) against $\ell_\infty$ perturbations of size $8/255$ on CIFAR-100. All models are available at https://github.com/deepmind/deepmind-research/tree/master/adversarial_robustness.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Scissors Effect: When Resize-Based Input Diversity Helps or Hurts Transfer Attacks

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    Resize-based input diversity boosts transfer attacks from standard surrogates but harms them from robust ones on ImageNet by 10.3% on average, traced to gradient alignment and mitigated by a local gradient consistency check.

  2. Adversarial Robustness in One-Stage Learning-to-Defer

    stat.ML 2025-10 unverdicted novelty 7.0 of 10

    Develops the first adversarial robustness framework for one-stage learning-to-defer, including cost-sensitive surrogate losses and theoretical consistency guarantees for classification and regression.

  3. Towards Generalized Certified Robustness with Multi-Norm Training

    cs.LG 2024-10 unverdicted novelty 7.0 of 10

    CURE is the first multi-norm certified training method that improves union robustness across l_p norms and unseen perturbations on MNIST, CIFAR-10 and TinyImagenet.

  4. Measuring Model Robustness via Fisher Information: Spectral Bounds, Theoretical Guarantees, and Practical Algorithms

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Proposes spectral norm of Fisher Information Matrix as attack-agnostic robustness metric with closed-form bounds for common architectures and correlation to adversarial vulnerability.

  5. Detecting Adversarial Data via Provable Adversarial Noise Amplification

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    A provable adversarial noise amplification theorem under sufficient conditions enables a custom-trained detector that identifies adversarial examples at inference time using enhanced layer-wise noise signals.

  6. Sample-wise Adaptive Weighting for Transfer Consistency in Adversarial Distillation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    SAAD adaptively weights adversarial training samples by their transferability to the teacher, yielding higher AutoAttack robustness than prior distillation methods on CIFAR and Tiny-ImageNet without extra compute.

  7. Nearest Neighbor Projection Removal Adversarial Training

    cs.CV 2025-09 unverdicted novelty 6.0 of 10

    Nearest Neighbor Projection Removal Adversarial Training projects out inter-class dependencies in feature space during training, claims to reduce the Lipschitz constant and Rademacher complexity, and reports competiti...

  8. Improving Clean Accuracy via a Tangent-Space Perspective on Adversarial Training

    cs.LG 2024-08 unverdicted novelty 6.0 of 10

    TART improves clean accuracy in adversarial training by modulating perturbation bounds according to the tangential component of adversarial examples.

  9. Adversarial Robustness in One-Stage Learning-to-Defer

    stat.ML 2025-10 reject novelty 5.0 of 10

    New adversarial surrogate losses and claimed consistency guarantees for one-stage learning-to-defer in classification and regression, with experiments suggesting improved robustness.

  10. Theoretical Analysis of Relative Errors in Gradient Computations for Adversarial Attacks with CE Loss

    cs.LG 2025-07 conditional novelty 4.0 of 10

    T-MIFPE adaptively rescales logits with a theoretically motivated t* per attack phase to reduce floating-point gradient errors, edging out MIFPE in PGD robustness evaluation.

  11. Improving Adversarial Robustness Through Adaptive Learning-Driven Multi-Teacher Knowledge Distillation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A multi-teacher adversarial robustness distillation method (MTKD-AR) trains a clean-data student using cosine-similarity-weighted logits from adversarially trained teachers, reporting improved robustness on MNIST and ...

  12. Explaining Machine Learning and Memorization with Statistical Mechanics

    cs.LG 2026-06 unverdicted novelty 3.0 of 10

    Thesis uses statistical mechanics to study DAM and RBM models for understanding memorization, low-dimensional learning, and adversarial robustness in neural networks.

  13. RCR-AF: Enhancing Model Generalization via Rademacher Complexity Reduction Activation Function

    cs.LG 2025-07 reject novelty 2.0 of 10

    RCR-AF, a clipped scaled-softplus activation, is claimed to improve CIFAR-10 accuracy and robustness, but the evidence is undermined by test-set tuning and a flawed complexity derivation.

Pith tools