Pith. sign in

REVIEW 8 cited by

Robustness May Be at Odds with Accuracy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1805.12152 v5 pith:7RGZ5QJU submitted 2018-05-30 stat.ML cs.CVcs.LGcs.NE

classification stat.MLcs.CVcs.LGcs.NE
keywords standardaccuracyrobustrobustnessadversarialclassifiersmodelsphenomenon
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We show that there may exist an inherent tension between the goal of adversarial robustness and that of standard generalization. Specifically, training robust models may not only be more resource-consuming, but also lead to a reduction of standard accuracy. We demonstrate that this trade-off between the standard accuracy of a model and its robustness to adversarial perturbations provably exists in a fairly simple and natural setting. These findings also corroborate a similar phenomenon observed empirically in more complex settings. Further, we argue that this phenomenon is a consequence of robust classifiers learning fundamentally different feature representations than standard classifiers. These differences, in particular, seem to result in unexpected benefits: the representations learned by robust models tend to align better with salient data characteristics and human perception.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 373 citations worldwide. Full citation record

  1. Hiding in Plain Sight: An Effective Physical Adversarial Patch Attack against Visual-Infrared Fused Face Detection

    cs.CR 2026-07 conditional novelty 6.0 of 10

    A jointly optimized gradient-mask plus band-aid patch reportedly bypasses visible-infrared fused face detectors with >90% attack success in both digital and physical settings.

  2. Adversarial Examples Are Not Bugs, They Are Superposition

    cs.LG 2025-08 unverdicted novelty 6.0 of 10

    The paper argues that adversarial examples arise from superposition, and shows that changing superposition changes robustness and vice versa in toy models and ResNet18.

  3. Adversarial Training Improves Generalization Under Distribution Shifts in Bioacoustics

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Output-space adversarial training improved clean-data performance and adversarial robustness of two bird sound classifiers across seven soundscape test sets, and stabilized prototype-based explanations.

  4. Exploring Visual Prompting: Robustness Inheritance and Beyond

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Visual prompts built on robust source models inherit adversarial robustness but lose standard accuracy; a max-pooling over logit blocks (PBL) improves accuracy while keeping most robustness.

  5. Evaluation of Adversarial Robustness in Arabic Language Models

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Arabic BERT-family sentiment models lose up to 92% accuracy under diacritics and 58% under conjunction attacks; paraphrase attacks cut accuracy by 76% on average, and adversarial training only partially helps.

  6. RobQFL: Robust Quantum Federated Learning in Adversarial Environment

    quant-ph 2025-09 conditional novelty 5.0 of 10

    Partial adversarial coverage in simulated quantum federated learning improves small-perturbation robustness with little clean-accuracy loss, but label-sorted non-IID data removes about half the robustness benefit.

  7. DUMB and DUMBer: Is Adversarial Training Worth It in the Real World?

    cs.CR 2025-06 conditional novelty 5.0 of 10

    Across 240 model configurations and 13 attacks, adaptive and curriculum adversarial training give the largest robustness gains, but 20.53% of evaluations show negative gains, mostly under mismatched source-target mode...

  8. PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training

    cs.CR 2025-07 reject novelty 3.0 of 10

    A PRM-free alignment pipeline combining genetic algorithm red teaming and multi-objective adversarial training is claimed to beat PRM-based methods at 61% lower cost, but the experiments are unverifiable.

Pith tools