Pith. sign in

REVIEW 4 cited by

Rethinking Model Ensemble in Transfer-based Adversarial Attacks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.09105 v2 pith:JBNN4YGY submitted 2023-03-16 cs.CV

classification cs.CV
keywords adversarialmodelensemblemodelstransferabilityattacksexamplesproperties
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

It is widely recognized that deep learning models lack robustness to adversarial examples. An intriguing property of adversarial examples is that they can transfer across different models, which enables black-box attacks without any knowledge of the victim model. An effective strategy to improve the transferability is attacking an ensemble of models. However, previous works simply average the outputs of different models, lacking an in-depth analysis on how and why model ensemble methods can strongly improve the transferability. In this paper, we rethink the ensemble in adversarial attacks and define the common weakness of model ensemble with two properties: 1) the flatness of loss landscape; and 2) the closeness to the local optimum of each model. We empirically and theoretically show that both properties are strongly correlated with the transferability and propose a Common Weakness Attack (CWA) to generate more transferable adversarial examples by promoting these two properties. Experimental results on both image classification and object detection tasks validate the effectiveness of our approach to improving the adversarial transferability, especially when attacking adversarially trained models. We also successfully apply our method to attack a black-box large vision-language model -- Google's Bard, showing the practical effectiveness. Code is available at \url{https://github.com/huanranchen/AdversarialAttacks}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mitigating Object Hallucination via Robust Local Perception Search

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A training-free decoding method that uses an MLLM's own local object descriptions as a reward prior, combined with CLIP similarity, to cut object hallucination, especially under adversarial image noise.

  2. ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    ActionSink claims to improve robot manipulation precision by representing actions as optical flow and integrating retrieved historical flows, with a 7.9% LIBERO success-rate gain over prior SOTA.

  3. Disrupting Semantic and Abstract Features for Better Adversarial Transferability

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A transfer-based adversarial attack that mixes image blocks and rotated frequency spectra improves transferability against unseen models.

  4. Blockchain Network Analysis using Quantum Inspired Graph Neural Networks & Ensemble Models

    cs.LG 2025-08 unverdicted novelty 4.0 of 10

    The submission's abstract claims a quantum-inspired GNN with a CP-decomposition layer reaches 74.8% F2 on blockchain fraud detection, but the uploaded full text is an unrelated paper on VLM agent security.

Pith tools