REVIEW 11 cited by
On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent work has shown that it is possible to train deep neural networks that are provably robust to norm-bounded adversarial perturbations. Most of these methods are based on minimizing an upper bound on the worst-case loss over all possible adversarial perturbations. While these techniques show promise, they often result in difficult optimization procedures that remain hard to scale to larger networks. Through a comprehensive analysis, we show how a simple bounding technique, interval bound propagation (IBP), can be exploited to train large provably robust neural networks that beat the state-of-the-art in verified accuracy. While the upper bound computed by IBP can be quite weak for general networks, we demonstrate that an appropriate loss and clever hyper-parameter schedule allow the network to adapt such that the IBP bound is tight. This results in a fast and stable learning algorithm that outperforms more sophisticated methods and achieves state-of-the-art results on MNIST, CIFAR-10 and SVHN. It also allows us to train the largest model to be verified beyond vacuous bounds on a downscaled version of ImageNet.
Forward citations
Cited by 11 Pith papers
-
No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason
An AI may assert or deny high-stakes claims only when it can exhibit a publicly contestable certificate; otherwise it is obligated to return Undetermined.
-
IoUCert: Robustness Verification for Anchor-based Object Detectors
IoUCert derives exact IoU bounds over anchor-offset boxes via a coordinate transformation and uses them to formally verify single-object SSD, YOLOv2, and YOLOv3 models under brightness, contrast, and motion-blur pertu...
-
How Context Attribution Handles What the Model Already Knows
Context attribution methods cannot disentangle in-context from in-weight knowledge and assign unfaithful scores under overlap; new metrics and WMDP-Cyber++ quantify the failure.
-
Certified Training for Convolutional Perturbations
A certified-training method using parameterized blur kernels and symbolic bound propagation gives provable robustness to motion blur and related convolutional perturbations, reaching over 80% verified accuracy on CIFAR10.
-
Training with Hard Constraints: Learning Neural Certificates and Controllers for SDEs
Neural reach-avoid certificates for SDEs can be trained with hard guarantees via a bound-based loss, or with PAC guarantees via scenario optimization on the last layer.
-
Sample Efficient Certification of Discrete-Time Control Barrier Functions
A Lipschitz-based verification method for discrete-time control barrier functions that reduces required samples by allowing coarser sampling away from the safety boundary, with sample complexity bounds and a numerical...
-
Adversarial Examples Are Not Bugs, They Are Superposition
The paper argues that adversarial examples arise from superposition, and shows that changing superposition changes robustness and vice versa in toy models and ResNet18.
-
Verification of Visual Controllers via Compositional Geometric Transformations
The paper combines DeepG pixel bounds with CROWN bound propagation to compute outer approximations of reachable sets for vision-based controllers under entity-specific geometric perturbations.
-
Evaluation of Adversarial Robustness in Arabic Language Models
Arabic BERT-family sentiment models lose up to 92% accuracy under diacritics and 58% under conjunction attacks; paraphrase attacks cut accuracy by 76% on average, and adversarial training only partially helps.
-
No Certificate for Alignment: Two Independent Impossibilities and the Pareto Frontier of Achievable Safety Guarantees
No sound, complete, and tractable procedure can certify AI alignment over full open-ended domains; only pairwise combinations of those three properties are achievable.
-
SAIL: Sound Abstract Interpreters with LLMs
SAIL synthesizes globally sound abstract transformers for neural-network operators by combining LLM generation with syntactic validation, SMT-based soundness checking, and cost-guided iterative refinement.
Discussion (0). Sign in to comment.