Pith. sign in

REVIEW 1 cited by

A New Kind of Adversarial Example

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.02430 v2 pith:35DIITJW submitted 2022-08-04 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords adversarialmodelattackattacksdecideefficientexamplesfool
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Almost all adversarial attacks are formulated to add an imperceptible perturbation to an image in order to fool a model. Here, we consider the opposite which is adversarial examples that can fool a human but not a model. A large enough and perceptible perturbation is added to an image such that a model maintains its original decision, whereas a human will most likely make a mistake if forced to decide (or opt not to decide at all). Existing targeted attacks can be reformulated to synthesize such adversarial examples. Our proposed attack, dubbed NKE, is similar in essence to the fooling images, but is more efficient since it uses gradient descent instead of evolutionary algorithms. It also offers a new and unified perspective into the problem of adversarial vulnerability. Experimental results over MNIST and CIFAR-10 datasets show that our attack is quite efficient in fooling deep neural networks. Code is available at https://github.com/aliborji/NKE.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A New Kind of Adversarial Example: Measuring the Human-Model Gap, and Its Relationship to OOD Detection

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Large, clearly visible perturbations that keep a network's correct label at ~1.0 confidence leave humans (and human proxies) unable to recognize the image, are invisible to confidence/energy OOD detectors, can evade f...

Pith tools