Pith. sign in

REVIEW 11 cited by

On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.12715 v4 pith:3DGIZWS7 submitted 2018-10-30 cs.LG cs.CRstat.ML

classification cs.LGcs.CRstat.ML
keywords boundnetworksrobusttrainadversarialintervallossmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work has shown that it is possible to train deep neural networks that are provably robust to norm-bounded adversarial perturbations. Most of these methods are based on minimizing an upper bound on the worst-case loss over all possible adversarial perturbations. While these techniques show promise, they often result in difficult optimization procedures that remain hard to scale to larger networks. Through a comprehensive analysis, we show how a simple bounding technique, interval bound propagation (IBP), can be exploited to train large provably robust neural networks that beat the state-of-the-art in verified accuracy. While the upper bound computed by IBP can be quite weak for general networks, we demonstrate that an appropriate loss and clever hyper-parameter schedule allow the network to adapt such that the IBP bound is tight. This results in a fast and stable learning algorithm that outperforms more sophisticated methods and achieves state-of-the-art results on MNIST, CIFAR-10 and SVHN. It also allows us to train the largest model to be verified beyond vacuous bounds on a downscaled version of ImageNet.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason

    cs.CY 2026-03 unverdicted novelty 7.0 of 10

    An AI may assert or deny high-stakes claims only when it can exhibit a publicly contestable certificate; otherwise it is obligated to return Undetermined.

  2. IoUCert: Robustness Verification for Anchor-based Object Detectors

    cs.LG 2026-03 conditional novelty 7.0 of 10

    IoUCert derives exact IoU bounds over anchor-offset boxes via a coordinate transformation and uses them to formally verify single-object SSD, YOLOv2, and YOLOv3 models under brightness, contrast, and motion-blur pertu...

  3. How Context Attribution Handles What the Model Already Knows

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Context attribution methods cannot disentangle in-context from in-weight knowledge and assign unfaithful scores under overlap; new metrics and WMDP-Cyber++ quantify the failure.

  4. Certified Training for Convolutional Perturbations

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A certified-training method using parameterized blur kernels and symbolic bound propagation gives provable robustness to motion blur and related convolutional perturbations, reaching over 80% verified accuracy on CIFAR10.

  5. Training with Hard Constraints: Learning Neural Certificates and Controllers for SDEs

    eess.SY 2026-02 conditional novelty 6.0 of 10

    Neural reach-avoid certificates for SDEs can be trained with hard guarantees via a bound-based loss, or with PAC guarantees via scenario optimization on the last layer.

  6. Sample Efficient Certification of Discrete-Time Control Barrier Functions

    eess.SY 2025-09 conditional novelty 6.0 of 10

    A Lipschitz-based verification method for discrete-time control barrier functions that reduces required samples by allowing coarser sampling away from the safety boundary, with sample complexity bounds and a numerical...

  7. Adversarial Examples Are Not Bugs, They Are Superposition

    cs.LG 2025-08 unverdicted novelty 6.0 of 10

    The paper argues that adversarial examples arise from superposition, and shows that changing superposition changes robustness and vice versa in toy models and ResNet18.

  8. Verification of Visual Controllers via Compositional Geometric Transformations

    cs.RO 2025-07 reject novelty 6.0 of 10

    The paper combines DeepG pixel bounds with CROWN bound propagation to compute outer approximations of reachable sets for vision-based controllers under entity-specific geometric perturbations.

  9. Evaluation of Adversarial Robustness in Arabic Language Models

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Arabic BERT-family sentiment models lose up to 92% accuracy under diacritics and 58% under conjunction attacks; paraphrase attacks cut accuracy by 76% on average, and adversarial training only partially helps.

  10. No Certificate for Alignment: Two Independent Impossibilities and the Pareto Frontier of Achievable Safety Guarantees

    stat.ML 2026-03 accept novelty 5.0 of 10

    No sound, complete, and tractable procedure can certify AI alignment over full open-ended domains; only pairwise combinations of those three properties are achievable.

  11. SAIL: Sound Abstract Interpreters with LLMs

    cs.PL 2025-11 reject novelty 5.0 of 10

    SAIL synthesizes globally sound abstract transformers for neural-network operators by combining LLM generation with syntactic validation, SMT-based soundness checking, and cost-guided iterative refinement.

Pith tools