REVIEW 4 cited by
Certified Defenses against Adversarial Examples
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While neural networks have achieved high accuracy on standard image classification benchmarks, their accuracy drops to nearly zero in the presence of small adversarial perturbations to test inputs. Defenses based on regularization and adversarial training have been proposed, but often followed by new, stronger attacks that defeat these defenses. Can we somehow end this arms race? In this work, we study this problem for neural networks with one hidden layer. We first propose a method based on a semidefinite relaxation that outputs a certificate that for a given network and test input, no attack can force the error to exceed a certain value. Second, as this certificate is differentiable, we jointly optimize it with the network parameters, providing an adaptive regularizer that encourages robustness against all attacks. On MNIST, our approach produces a network and a certificate that no attack that perturbs each pixel by at most \epsilon = 0.1 can cause more than 35% test error.
Forward citations
Cited by 4 Pith papers
-
Robust Representation Consistency Model via Contrastive Denoising
rRCM, a contrastive denoising pre-training and fine-tuning scheme, gives a single-pass robust classifier that beats diffusion-based defenses on ImageNet and CIFAR-10 while reducing inference cost by up to 85x.
-
TorchLean: Formalizing Neural Networks in Lean
A Lean 4 framework gives neural networks one machine-checked semantics shared by training, autodiff, and CROWN-style verification, demonstrated on small robustness, PINN, and controller cases.
-
Enhancing Adversarial Transferability via Component-Wise Transformation
A block-wise interpolation and selective rotation attack, CWT, improves adversarial transferability across CNN and transformer models on ImageNet.
-
Unsupervised dense retrieval with conterfactual contrastive learning
A Shapley-value-based counterfactual regularization improves dense retrievers' robustness to adversarial attacks and enables key passage extraction without passage-level relevance annotations.
Discussion (0). Continue with ORCID to comment.