REVIEW 3 cited by
Provably Minimally-Distorted Adversarial Examples
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The ability to deploy neural networks in real-world, safety-critical systems is severely limited by the presence of adversarial examples: slightly perturbed inputs that are misclassified by the network. In recent years, several techniques have been proposed for increasing robustness to adversarial examples --- and yet most of these have been quickly shown to be vulnerable to future attacks. For example, over half of the defenses proposed by papers accepted at ICLR 2018 have already been broken. We propose to address this difficulty through formal verification techniques. We show how to construct provably minimally distorted adversarial examples: given an arbitrary neural network and input sample, we can construct adversarial examples which we prove are of minimal distortion. Using this approach, we demonstrate that one of the recent ICLR defense proposals, adversarial retraining, provably succeeds at increasing the distortion required to construct adversarial examples by a factor of 4.2.
Forward citations
Cited by 3 Pith papers
-
Achieving Verified Robustness to Symbol Substitutions via Interval Bound Propagation
Modeling word or character substitutions as a simplex and using interval bound propagation yields text classifiers whose robustness can be certified in two forward passes, at a small nominal accuracy cost.
-
Position: Adversarial ML for LLMs Is Not Making Any Progress
The authors argue that LLM-era adversarial machine learning is less well-defined, harder to solve, and harder to evaluate, so meaningful progress may not be achievable or trackable in the current paradigm.
-
Statistical Runtime Verification for LLMs via Robustness Estimation
RoMA, a statistical robustness estimator, is adapted to black-box language models and is shown to approximate exact verification within 1% on small networks while scaling to BERT sentiment analysis.
Discussion (0). Continue with ORCID to comment.