Pith. sign in

REVIEW 35 cited by

Ensemble Adversarial Training: Attacks and Defenses

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1705.07204 v5 pith:KZG3KZYN submitted 2017-05-19 stat.ML cs.CRcs.LG

classification stat.MLcs.CRcs.LG
keywords adversarialtrainingattacksmodelsdataperturbationsblack-boxensemble
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Adversarial examples are perturbed inputs designed to fool machine learning models. Adversarial training injects such examples into training data to increase robustness. To scale this technique to large datasets, perturbations are crafted using fast single-step methods that maximize a linear approximation of the model's loss. We show that this form of adversarial training converges to a degenerate global minimum, wherein small curvature artifacts near the data points obfuscate a linear approximation of the loss. The model thus learns to generate weak perturbations, rather than defend against strong ones. As a result, we find that adversarial training remains vulnerable to black-box attacks, where we transfer perturbations computed on undefended models, as well as to a powerful novel single-step attack that escapes the non-smooth vicinity of the input data via a small random step. We further introduce Ensemble Adversarial Training, a technique that augments training data with perturbations transferred from other models. On ImageNet, Ensemble Adversarial Training yields models with strong robustness to black-box attacks. In particular, our most robust model won the first round of the NIPS 2017 competition on Defenses against Adversarial Attacks. However, subsequent work found that more elaborate black-box attacks could significantly enhance transferability and reduce the accuracy of our models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 35 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 1,109 citations worldwide. Full citation record

  1. Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks

    cs.LG 2026-08 conditional novelty 6.0 of 10

    BMAT couples initialization, perturbation, and surrogate adaptation in one bilevel-minimax optimization, markedly improving adversarial example transfer to unseen victims.

  2. Adversarially Robust Abductive Fusion of Pre-trained Transformer-based Perception Models

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A domain-knowledge-free error-detection layer, built from per-model label vector pools, is fused via consistency-based abduction to match hand-crafted rules on clean data and outperform majority voting under coordinat...

  3. Make me an Expert: Distilling from Generalist Black-Box Models into Specialized Models for Semantic Segmentation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    ATGC selects the best input scale for a black-box open-vocabulary segmentation API, using DINOv2 attention entropy, improving one-hot-label distillation on Cityscapes and ACDC.

  4. Exploring Visual Prompting: Robustness Inheritance and Beyond

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Visual prompts built on robust source models inherit adversarial robustness but lose standard accuracy; a max-pooling over logit blocks (PBL) improves accuracy while keeping most robustness.

  5. Adversarially Robust Spiking Neural Networks with Sparse Connectivity

    cs.NE 2025-05 conditional novelty 6.0 of 10

    A conversion pipeline turns robustly trained, pruned artificial networks into sparse spiking networks that keep adversarial robustness and cut stored weights by up to 100x.

  6. Towards Robust Stability Prediction in Smart Grids: GAN-based Approach under Data Constraints and Adversarial Challenges

    cs.CR 2025-01 conditional novelty 6.0 of 10

    A GAN trained only on stable grid data can flag unstable states and adversarial attacks, reaching 98.1% stability accuracy on the augmented UCI grid dataset.

  7. A Generative Victim Model for Segmentation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A diffusion model's conditional and unconditional scores can be combined to generate transferable adversarial perturbations for segmentation without a segmentation victim model.

  8. Spatiotemporally Constrained Action Space Attacks on Deep Reinforcement Learning Agents

    cs.LG 2019-09 conditional novelty 6.0 of 10

    A dynamics-aware, look-ahead attack on the action space of deep RL agents consistently beats a myopic action-space attack at equal total budget and also reveals which actuators are most vulnerable.

  9. Metric Learning for Adversarial Robustness

    cs.LG 2019-09 conditional novelty 6.0 of 10

    Adding a triplet loss with semi-hard negative sampling to adversarial training improves robustness and adversarial-example detection on MNIST, CIFAR-10, and Tiny ImageNet.

  10. Improving Adversarial Robustness via Attention and Adversarial Logit Pairing

    cs.LG 2019-08 reject novelty 6.0 of 10

    Aligning attention maps and logits between clean and adversarial images during training improves robustness accuracy over adversarial training on three small image datasets.

  11. Protecting Neural Networks with Hierarchical Random Switching: Towards Better Robustness-Accuracy Trade-off for Stochastic Defenses

    cs.LG 2019-08 conditional novelty 6.0 of 10

    A stochastic neural network defense using randomly switched parallel weight channels achieves a high defense-per-accuracy-drop ratio on MNIST and CIFAR-10 and is reported as the first defense against adversarial repro...

  12. Defending Against Adversarial Iris Examples Using Wavelet Decomposition

    cs.CV 2019-08 conditional novelty 6.0 of 10

    Three wavelet-based denoising strategies detect adversarial iris images by removing or denoising the wavelet sub-bands most affected by the attack.

  13. Adversarial Self-Defense for Cycle-Consistent GANs

    cs.CV 2019-08 conditional novelty 6.0 of 10

    Two defenses for cycle-consistent GANs, additive noise and a guess discriminator, reduce hidden-information 'self-adversarial attacks' and improve robustness to high-frequency perturbations.

  14. Foveation-Guided Dynamic Token Selection for Robust and Efficient Vision Transformers

    cs.CV 2026-07 conditional novelty 5.0 of 10

    FDT adds foveation and binary fixation modules to DeiT so multi-scale tokens are selected dynamically in one pass, improving ImageNet100 accuracy, MACs, and robustness without adversarial training.

  15. Generating Transferrable Adversarial Examples via Local Mixing and Logits Optimization for Remote Sensing Object Recognition

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A local-mixing and logit-optimization attack improves transferability of adversarial examples for remote sensing object recognition, outperforming 12 prior methods on two benchmarks.

  16. ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision Transformers

    cs.CV 2025-08 conditional novelty 5.0 of 10

    ViT-EnsembleAttack augments each ViT surrogate with three randomized strategies, tunes their parameters by Bayesian optimization, and ensembles them to substantially improve adversarial transferability.

  17. Boosting Adversarial Transferability Against Defenses via Multi-Scale Transformation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A Segmented Gaussian Pyramid transformation that averages gradients over three downsampled scales improves black-box adversarial transferability against defense models.

  18. DUMB and DUMBer: Is Adversarial Training Worth It in the Real World?

    cs.CR 2025-06 conditional novelty 5.0 of 10

    Across 240 model configurations and 13 attacks, adaptive and curriculum adversarial training give the largest robustness gains, but 20.53% of evaluations show negative gains, mostly under mismatched source-target mode...

  19. Boosting Adversarial Transferability via High-Frequency Augmentation and Hierarchical-Gradient Fusion

    cs.CV 2025-05 conditional novelty 5.0 of 10

    FSA combines Fourier high-frequency augmentation with Gaussian pyramid gradient fusion to boost adversarial transferability against defended black-box models.

  20. Towards Adaptive Meta-Gradient Adversarial Examples for Visual Tracking

    cs.CV 2025-05 conditional novelty 5.0 of 10

    The AMGA attack, built from an ensemble of image classifiers trained with momentum, Gaussian smoothing, and a meta-learning-style update, substantially reduces the accuracy of seven visual trackers on three benchmarks...

  21. Universal, transferable and targeted adversarial attacks

    cs.LG 2019-08 reject novelty 5.0 of 10

    A trained encoder-decoder network (FTN) transforms source images into targeted adversarial examples that reportedly transfer across VGG19, Inception-v3, ResNet variants, DenseNet, and a black-box commercial classifier...

  22. Once a MAN: Towards Multi-Target Attack via Learning Multi-Target Adversarial Network Once

    cs.CV 2019-08 conditional novelty 5.0 of 10

    By feeding a one-hot target label into an encoder-decoder, a single MAN model can attack any ImageNet or CIFAR10 class and outperforms single-target generators in attack rate and transferability.

  23. Enhancing Adversarial Transferability through Block Stretch and Shrink

    cs.LG 2025-11 reject novelty 4.0 of 10

    A block stretch-and-shrink input transformation improves black-box adversarial transferability in experiments on 1000 ImageNet images, but the submitted manuscript contains missing figures and an abstract describing a...

  24. DeepDefense: Robust Learning via Layer-Wise Gradient-Feature Alignment

    cs.LG 2025-11 reject novelty 4.0 of 10

    A layer-wise gradient-feature alignment regularizer is claimed to make neural networks robust to adversarial perturbations, with empirical gains over a PGD-based adversarial training baseline.

  25. Improving Adversarial Robustness Through Adaptive Learning-Driven Multi-Teacher Knowledge Distillation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A multi-teacher adversarial robustness distillation method (MTKD-AR) trains a clean-data student using cosine-similarity-weighted logits from adversarially trained teachers, reporting improved robustness on MNIST and ...

  26. Adversarial Semantic and Label Perturbation Attack for Pedestrian Attribute Recognition

    cs.CV 2025-05 conditional novelty 4.0 of 10

    ASL-PAR creates universal adversarial noise using label and semantic perturbation, dropping PromptPAR's mean accuracy by up to 40 points on standard PAR benchmarks, while a filter-and-prompt defense restores most of the drop.

  27. Are classical deep neural networks weakly adversarially robust?

    cs.CV 2025-05 reject novelty 4.0 of 10

    Using layer-wise feature paths and class-centered paths, the paper reports 44.17% and 46.1% adversarial accuracy on CIFAR-10 for ResNet-20 and ResNet-18, respectively, without adversarial training.

  28. Enhancing Adversarial Transferability via Component-Wise Transformation

    cs.CV 2025-01 conditional novelty 4.0 of 10

    A block-wise interpolation and selective rotation attack, CWT, improves adversarial transferability across CNN and transformer models on ImageNet.

  29. Face De-identification: State-of-the-art Methods and Comparative Studies

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A structured survey with new experimental comparisons showing identity-based semantic-level de-identification methods best preserve the privacy-utility trade-off.

  30. DAPAS : Denoising Autoencoder to Prevent Adversarial attack in Semantic Segmentation

    cs.CV 2019-08 reject novelty 4.0 of 10

    A denoising autoencoder placed before DeepLab V3 Plus partially restores segmentation accuracy after FGSM and I-FGSM attacks, but only against attacks that ignore the filter.

  31. AdvGAN++ : Harnessing latent layers for adversary generation

    cs.CV 2019-08 reject novelty 4.0 of 10

    Using a target model's latent features as the conditioning input to a GAN generator yields higher adversarial attack success rates on MNIST and CIFAR-10 than AdvGAN's image-conditioned generator.

  32. PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training

    cs.CR 2025-07 reject novelty 3.0 of 10

    A PRM-free alignment pipeline combining genetic algorithm red teaming and multi-objective adversarial training is claimed to beat PRM-based methods at 61% lower cost, but the experiments are unverifiable.

  33. Deep Neural Network Ensembles against Deception: Ensemble Diversity, Accuracy and Robustness

    cs.LG 2019-08 reject novelty 3.0 of 10

    Selecting DNN ensemble teams by low Kappa disagreement is presented as a defense against adversarial examples, but the evidence is preliminary and incomplete.

  34. Security and Privacy of Digital Twins for Advanced Manufacturing: A Survey

    eess.SY 2024-12 conditional novelty 2.0 of 10

    A survey of cybersecurity and privacy risks for manufacturing digital twins, grouping threats and defenses into data collection, data sharing, machine learning, and system-level security.

  35. A Review of the Duality of Adversarial Learning in Network Intrusion: Attacks and Countermeasures

    cs.CR 2024-12 conditional novelty 2.0 of 10

    A survey of adversarial learning attacks and defenses for network intrusion detection, organized around data poisoning, test-time evasion, and reverse engineering, that finds the NIDS-specific niche remains small and ...

Pith tools