Pith. sign in

REVIEW 3 cited by

Patch-Fool: Are Vision Transformers Always Robust Against Adversarial Perturbations?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.08392 v3 pith:2ZIQELZT submitted 2022-03-16 cs.CV

classification cs.CV
keywords vitscnnspatch-fooladversarialrobustnessattacksperturbationsvision
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision transformers (ViTs) have recently set off a new wave in neural architecture design thanks to their record-breaking performance in various vision tasks. In parallel, to fulfill the goal of deploying ViTs into real-world vision applications, their robustness against potential malicious attacks has gained increasing attention. In particular, recent works show that ViTs are more robust against adversarial attacks as compared with convolutional neural networks (CNNs), and conjecture that this is because ViTs focus more on capturing global interactions among different input/feature patches, leading to their improved robustness to local perturbations imposed by adversarial attacks. In this work, we ask an intriguing question: "Under what kinds of perturbations do ViTs become more vulnerable learners compared to CNNs?" Driven by this question, we first conduct a comprehensive experiment regarding the robustness of both ViTs and CNNs under various existing adversarial attacks to understand the underlying reason favoring their robustness. Based on the drawn insights, we then propose a dedicated attack framework, dubbed Patch-Fool, that fools the self-attention mechanism by attacking its basic component (i.e., a single patch) with a series of attention-aware optimization techniques. Interestingly, our Patch-Fool framework shows for the first time that ViTs are not necessarily more robust than CNNs against adversarial perturbations. In particular, we find that ViTs are more vulnerable learners compared with CNNs against our Patch-Fool attack which is consistent across extensive experiments, and the observations from Sparse/Mild Patch-Fool, two variants of Patch-Fool, indicate an intriguing insight that the perturbation density and strength on each patch seem to be the key factors that influence the robustness ranking between ViTs and CNNs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Preventing Adversarial AI Attacks Against Autonomous Situational Awareness: A Maritime Case Study

    cs.CR 2025-05 conditional novelty 6.0 of 10

    DFCR combines AIS, radar, and optical object detection with validation components to lower AI confidence on adversarial contacts, reporting up to 100% loss reduction on patch and spoofing attacks.

  2. Attacking Attention of Foundation Models Disrupts Downstream Tasks

    cs.CR 2025-06 conditional novelty 5.0 of 10

    A task-agnostic attack that perturbs attention and embeddings of CLIP/ViT backbones degrades classification, retrieval, captioning, segmentation, and depth estimation without using labels or text.

  3. Vision Transformer with Adversarial Indicator Token against Adversarial Attacks in Radio Signal Classifications

    cs.LG 2025-06 reject novelty 4.0 of 10

    A vision transformer with an additional adversarial indicator token detects and withstands white-box adversarial attacks on radio signal modulation classification better than several prior defenses.

Pith tools