REVIEW 19 cited by
Robustness May Be at Odds with Accuracy
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We show that there may exist an inherent tension between the goal of adversarial robustness and that of standard generalization. Specifically, training robust models may not only be more resource-consuming, but also lead to a reduction of standard accuracy. We demonstrate that this trade-off between the standard accuracy of a model and its robustness to adversarial perturbations provably exists in a fairly simple and natural setting. These findings also corroborate a similar phenomenon observed empirically in more complex settings. Further, we argue that this phenomenon is a consequence of robust classifiers learning fundamentally different feature representations than standard classifiers. These differences, in particular, seem to result in unexpected benefits: the representations learned by robust models tend to align better with salient data characteristics and human perception.
Forward citations
Cited by 19 Pith papers
-
Direct Search Methods for Online Nonconvex Optimization Under Inexact Bandit Feedback
A randomized two-point direct-search algorithm for nonconvex time-varying optimization with inexact bandit feedback reaches epsilon-stationarity in O(p/epsilon^2) iterations under constant probing, and O(p/epsilon^2 l...
-
Hiding in Plain Sight: An Effective Physical Adversarial Patch Attack against Visual-Infrared Fused Face Detection
A jointly optimized gradient-mask plus band-aid patch reportedly bypasses visible-infrared fused face detectors with >90% attack success in both digital and physical settings.
-
Adversarial Examples Are Not Bugs, They Are Superposition
The paper argues that adversarial examples arise from superposition, and shows that changing superposition changes robustness and vice versa in toy models and ResNet18.
-
Adversarial Training Improves Generalization Under Distribution Shifts in Bioacoustics
Output-space adversarial training improved clean-data performance and adversarial robustness of two bird sound classifiers across seven soundscape test sets, and stabilized prototype-based explanations.
-
Exploring Visual Prompting: Robustness Inheritance and Beyond
Visual prompts built on robust source models inherit adversarial robustness but lose standard accuracy; a max-pooling over logit blocks (PBL) improves accuracy while keeping most robustness.
-
When Maximum Entropy Misleads Policy Optimization
Maximum entropy RL can be formally steered into arbitrary suboptimal policies at convergence by adding entropy trap states, while standard RL is unaffected.
-
Grounding Functional Similarity by Invariance-Aware Model Stitching
FuLA, a task-agnostic stitching objective that aligns intermediate features through the frozen end network, is claimed to be a more reliable functional similarity metric than task-based stitching.
-
DROP: Poison Dilution via Knowledge Distillation for Federated Learning
DROP combines clustering, client reputation tracking, and GAN-guided knowledge distillation to suppress targeted backdoor attacks in federated learning, reporting under 2% attack success in most tested IID settings.
-
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
Adversarial training at both the CLIP pre-training stage and the LLaVA instruction-tuning stage produces vision-language models with state-of-the-art robustness and near-baseline clean performance.
-
Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
Adaptive attacks reduce the robust accuracy of the 'Ensemble Everything Everywhere' defense to 11% on CIFAR-10 and 14% on CIFAR-100 under an l-infinity bound of 8/255.
-
MEGL: Multimodal Explanation-Guided Learning
A multimodal explanation-guided learning framework that jointly uses visual saliency maps and textual rationales to train image classifiers, improving accuracy, visual explanation overlap, and text explanation scores ...
-
A principled approach for generating adversarial images under non-smooth dissimilarity metrics
ProxLogBarrier extends the LogBarrier adversarial attack to non-smooth metrics via proximal gradient, achieving state-of-the-art ℓ0 perturbation results.
-
Evaluation of Adversarial Robustness in Arabic Language Models
Arabic BERT-family sentiment models lose up to 92% accuracy under diacritics and 58% under conjunction attacks; paraphrase attacks cut accuracy by 76% on average, and adversarial training only partially helps.
-
RobQFL: Robust Quantum Federated Learning in Adversarial Environment
Partial adversarial coverage in simulated quantum federated learning improves small-perturbation robustness with little clean-accuracy loss, but label-sorted non-IID data removes about half the robustness benefit.
-
DUMB and DUMBer: Is Adversarial Training Worth It in the Real World?
Across 240 model configurations and 13 attacks, adaptive and curriculum adversarial training give the largest robustness gains, but 20.53% of evaluations show negative gains, mostly under mismatched source-target mode...
-
Mitigating Spurious Negative Pairs for Robust Industrial Anomaly Detection
A contrastive anomaly detector trained on pseudo-anomalies and opposite-pair repulsion raises average robust AUROC under PGD-1000 from 39.7% (best prior) to 65.8%.
-
Cross-Entropy Attacks to Language Models via Rare Event Simulation
A cross-entropy optimization attack (CEA) with sememe- and MLM-based candidate words improves black-box adversarial attacks on classifiers and machine translation models.
-
Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles
Training traffic sign classifiers with two frozen text-prototype losses, built from VLM-generated descriptions and class names, improves accuracy under shadows, natural light, and printed patches, with no inference-ti...
-
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
A PRM-free alignment pipeline combining genetic algorithm red teaming and multi-objective adversarial training is claimed to beat PRM-based methods at 61% lower cost, but the experiments are unverifiable.
Discussion (0). Continue with ORCID to comment.