Pith. sign in

REVIEW 22 cited by

RobustBench: a standardized adversarial robustness benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.09670 v3 pith:Q2YEQM3L submitted 2020-10-19 cs.LG cs.CRcs.CVstat.ML

RobustBench: a standardized adversarial robustness benchmark

classification cs.LG cs.CRcs.CVstat.ML
keywords robustnessmodelsadversarialrobustbenchautoattackevaluationsattacksbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

As a research community, we are still lacking a systematic understanding of the progress on adversarial robustness which often makes it hard to identify the most promising ideas in training robust models. A key challenge in benchmarking robustness is that its evaluation is often error-prone leading to robustness overestimation. Our goal is to establish a standardized benchmark of adversarial robustness, which as accurately as possible reflects the robustness of the considered models within a reasonable computational budget. To this end, we start by considering the image classification task and introduce restrictions (possibly loosened in the future) on the allowed models. We evaluate adversarial robustness with AutoAttack, an ensemble of white- and black-box attacks, which was recently shown in a large-scale study to improve almost all robustness evaluations compared to the original publications. To prevent overadaptation of new defenses to AutoAttack, we welcome external evaluations based on adaptive attacks, especially where AutoAttack flags a potential overestimation of robustness. Our leaderboard, hosted at https://robustbench.github.io/, contains evaluations of 120+ models and aims at reflecting the current state of the art in image classification on a set of well-defined tasks in $\ell_\infty$- and $\ell_2$-threat models and on common corruptions, with possible extensions in the future. Additionally, we open-source the library https://github.com/RobustBench/robustbench that provides unified access to 80+ robust models to facilitate their downstream applications. Finally, based on the collected models, we analyze the impact of robustness on the performance on distribution shifts, calibration, out-of-distribution detection, fairness, privacy leakage, smoothness, and transferability.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Toward Calibrated, Fair, and accurate Deepfake Detection

    cs.LG 2026-06 unverdicted novelty 7.0

    Face-Feature Tuning is a label-free logit remapping method that reduces FPR/TPR gaps across groups in deepfake detection while preserving overall accuracy.

  2. Optimality of Sub-network Laplace Approximations: New Results and Methods

    stat.ML 2026-05 conditional novelty 7.0

    Sub-network Laplace approximations always underestimate full-model predictive variance, and two new gradient-based and greedy selection rules provide theoretically grounded improvements.

  3. Agent-Centric Observation Adaptation for Robust Visual Control under Dynamic Perturbations

    cs.RO 2026-04 unverdicted novelty 7.0

    ACO-MoE recovers 95.3% of clean-input performance in visual control tasks under Markov-switching corruptions by routing restoration experts and anchoring representations to clean foreground masks.

  4. Agent-Centric Observation Adaptation for Robust Visual Control under Dynamic Perturbations

    cs.RO 2026-04 unverdicted novelty 7.0

    ACO-MoE employs agent-centric mixture-of-experts to decouple task-relevant features from dynamic visual perturbations in RL, recovering 95.3% of clean performance on the new VDCS benchmark.

  5. Learning Robustness at Test-Time from a Non-Robust Teacher

    cs.CV 2026-04 unverdicted novelty 7.0

    A test-time adaptation framework anchors adversarial training to a non-robust teacher's predictions, yielding more stable optimization and better robustness-accuracy trade-offs than standard self-consistency methods.

  6. Contrastive Residual Energy Test-time Adaptation

    cs.LG 2025-05 unverdicted novelty 7.0

    CreTTA reformulates test-time adaptation of marginal distributions as residual energy learning, producing a contrastive objective that cancels the partition function and uses relative energy differences for adaptive g...

  7. RoAd-RL: A Unified Library and Benchmark for Robust Adversarial Reinforcement Learning

    cs.LG 2026-06 conditional novelty 6.0

    RoAd-RL is a new benchmarking library for adversarial reinforcement learning that evaluates DQN, PPO, and SAC agents across 192 attack-defense configurations and finds substantial robustness variations plus cases wher...

  8. Sensitivity as a Double-Edged Sword: A Trade-off Between Discriminability and Adversarial Robustness

    cs.CV 2026-06 unverdicted novelty 6.0

    Identifies sensitivity as the source of both discriminability and vulnerability in FC classifiers versus robustness in l2 classifiers, and introduces HPM prototype fusion plus MSA evaluation to improve adversarial robustness.

  9. Optimality of Sub-network Laplace Approximations: New Results and Methods

    stat.ML 2026-05 conditional novelty 6.0

    Sub-network Laplace approximations always underestimate an idealized predictive variance, and the proposed gradient- and greedy-based parameter selection rules provably close that gap better than existing heuristics.

  10. Detecting Adversarial Data via Provable Adversarial Noise Amplification

    cs.LG 2026-05 unverdicted novelty 6.0

    A provable adversarial noise amplification theorem under sufficient conditions enables a custom-trained detector that identifies adversarial examples at inference time using enhanced layer-wise noise signals.

  11. Adversarial Label Invariant Graph Data Augmentations for Out-of-Distribution Generalization

    cs.LG 2026-04 unverdicted novelty 6.0

    RIA uses adversarial exploration of counterfactual graph environments via label-invariant augmentations to improve OoD generalization in graph classification tasks.

  12. Compression as an Adversarial Amplifier Through Decision Space Reduction

    cs.CV 2026-04 unverdicted novelty 6.0

    Compression acts as an adversarial amplifier by reducing the decision space of image classifiers, making attacks in compressed representations substantially more effective than pixel-space attacks under the same pertu...

  13. Sparse Autoencoders are Capable LLM Jailbreak Mitigators

    cs.CR 2026-02 conditional novelty 6.0

    CC-Delta defends LLMs against jailbreaks by statistically selecting and steering sparse-SAEs features that change when harmful prompts are embedded in jailbreak contexts, outperforming dense activation steering across...

  14. How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks

    cs.CV 2025-07 unverdicted novelty 6.0

    Multimodal foundation models achieve respectable but sub-specialist performance on semantic vision tasks and weaker results on geometric tasks when evaluated through prompt chaining on established benchmarks.

  15. SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

    cs.LG 2023-10 accept novelty 6.0

    SmoothLLM mitigates jailbreaking attacks on LLMs by randomly perturbing multiple copies of a prompt at the character level and aggregating the outputs to detect adversarial inputs.

  16. Language Guided Adversarial Purification

    cs.LG 2023-09 unverdicted novelty 6.0

    LGAP generates captions from images to guide diffusion-based purification, outperforming other adversarial defenses without specialized training.

  17. Unsolved Problems in ML Safety

    cs.LG 2021-09 accept novelty 6.0

    The paper presents a roadmap that identifies four unsolved problems in ML safety: robustness against hazards, monitoring for hazards, alignment of model goals with human intent, and systemic safety.

  18. Breaking TinyML: Why Quantized Neural Networks Need Domain-Specific Security Analysis

    cs.CR 2026-06 conditional novelty 5.5

    Surrogate extraction followed by FGSM/PGD reduces int-8 TinyML accuracy by up to 47% on CIFAR-10 with 50k queries, outperforming gray-box baselines and exposing hardware-specific QNN vulnerabilities.

  19. SoK: Adversarial Robustness of the Variational Quantum Eigensolver via Red-Teaming

    quant-ph 2026-07 conditional novelty 5.0

    On a unified VQE benchmark, QNBAD noise-induced attacks amplify energy error up to 8.84x, QTrojan up to 7.52x, and QDoor at most 1.37x.

  20. A combination of noise and bilateral filters achieve supralinear and scalable adversarial robustness in CNNs

    cs.LG 2026-06 unverdicted novelty 5.0

    A preprocessor of Gaussian noise plus bilateral filtering yields supralinear adversarial robustness in CNNs and, when paired with adversarial training, ranks near the top of RobustBench while using far less compute, p...

  21. Beyond Attack Success Rate: A Multi-Metric Evaluation of Adversarial Transferability in Medical Imaging Models

    cs.CV 2026-04 unverdicted novelty 4.0

    Perceptual quality metrics correlate strongly with each other but show minimal correlation with attack success rate across medical imaging models and datasets, making ASR alone inadequate for assessing adversarial robustness.

  22. LLM-Safety Evaluations Lack Robustness

    cs.CR 2025-03 unverdicted novelty 4.0

    LLM safety evaluations are hindered by noise in dataset curation, automated red-teaming, response generation, and LLM-judge evaluation, making fair comparisons difficult and slowing progress.