REVIEW 9 cited by
Rethinking Model Ensemble in Transfer-based Adversarial Attacks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
It is widely recognized that deep learning models lack robustness to adversarial examples. An intriguing property of adversarial examples is that they can transfer across different models, which enables black-box attacks without any knowledge of the victim model. An effective strategy to improve the transferability is attacking an ensemble of models. However, previous works simply average the outputs of different models, lacking an in-depth analysis on how and why model ensemble methods can strongly improve the transferability. In this paper, we rethink the ensemble in adversarial attacks and define the common weakness of model ensemble with two properties: 1) the flatness of loss landscape; and 2) the closeness to the local optimum of each model. We empirically and theoretically show that both properties are strongly correlated with the transferability and propose a Common Weakness Attack (CWA) to generate more transferable adversarial examples by promoting these two properties. Experimental results on both image classification and object detection tasks validate the effectiveness of our approach to improving the adversarial transferability, especially when attacking adversarially trained models. We also successfully apply our method to attack a black-box large vision-language model -- Google's Bard, showing the practical effectiveness. Code is available at \url{https://github.com/huanranchen/AdversarialAttacks}.
Forward citations
Cited by 9 Pith papers
-
Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks
BMAT couples initialization, perturbation, and surrogate adaptation in one bilevel-minimax optimization, markedly improving adversarial example transfer to unseen victims.
-
Mitigating Object Hallucination via Robust Local Perception Search
A training-free decoding method that uses an MLLM's own local object descriptions as a reward prior, combined with CLIP similarity, to cut object hallucination, especially under adversarial image noise.
-
Transferable Adversarial Attacks on Black-Box Vision-Language Models
Targeted, barely visible image perturbations transfer from open-source surrogate models to proprietary black-box VLLMs like GPT-4o, Claude, and Gemini, achieving high attack success on captioning, VQA, and receipt tex...
-
ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow
ActionSink claims to improve robot manipulation precision by representing actions as optical flow and integrating retrieved historical flows, with a 7.9% LIBERO success-rate gain over prior SOTA.
-
Disrupting Semantic and Abstract Features for Better Adversarial Transferability
A transfer-based adversarial attack that mixes image blocks and rotated frequency spectra improves transferability against unseen models.
-
Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment
FOA-Attack aligns global and clustered local features via optimal transport with dynamic ensemble weighting to create targeted adversarial images that transfer to closed-source multimodal LLMs.
-
FSPGD: Rethinking Black-box Attacks on Semantic Segmentation
FSPGD uses two feature-similarity losses to craft segmentation attacks that transfer across CNN and transformer models, and reports large mIoU drops on Pascal VOC and Cityscapes.
-
Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks
A step-by-step multimodal 'chain of attack' improves the transferability of targeted adversarial images against open vision-language models, with a new LLM-judged success metric.
-
Blockchain Network Analysis using Quantum Inspired Graph Neural Networks & Ensemble Models
The submission's abstract claims a quantum-inspired GNN with a CP-decomposition layer reaches 74.8% F2 on blockchain fraud detection, but the uploaded full text is an unrelated paper on VLM agent security.
Discussion (0). Continue with ORCID to comment.