Pith. sign in

REVIEW

GNP Attack: Transferable Adversarial Examples via Gradient Norm Penalty

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.04099 v1 pith:XUEUCEZR submitted 2023-07-09 cs.LG cs.CRcs.CV

classification cs.LGcs.CRcs.CV
keywords modelstransferabilitygradientmethodstargetveryadversarialattacks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adversarial examples (AE) with good transferability enable practical black-box attacks on diverse target models, where insider knowledge about the target models is not required. Previous methods often generate AE with no or very limited transferability; that is, they easily overfit to the particular architecture and feature representation of the source, white-box model and the generated AE barely work for target, black-box models. In this paper, we propose a novel approach to enhance AE transferability using Gradient Norm Penalty (GNP). It drives the loss function optimization procedure to converge to a flat region of local optima in the loss landscape. By attacking 11 state-of-the-art (SOTA) deep learning models and 6 advanced defense methods, we empirically show that GNP is very effective in generating AE with high transferability. We also demonstrate that it is very flexible in that it can be easily integrated with other gradient based methods for stronger transfer-based attacks.

Discussion (0). Sign in to comment.

Pith tools