Pith. sign in

REVIEW 2 cited by

How to Manipulate CNNs to Make Them Lie: the GradCAM Case

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.10901 v2 pith:DV6OVDV5 submitted 2019-07-25 cs.CV cs.CRcs.LG

classification cs.CVcs.CRcs.LG
keywords explanationinputgradcammodelbeendecisionsexplainmake
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recently many methods have been introduced to explain CNN decisions. However, it has been shown that some methods can be sensitive to manipulation of the input. We continue this line of work and investigate the explanation method GradCAM. Instead of manipulating the input, we consider an adversary that manipulates the model itself to attack the explanation. By changing weights and architecture, we demonstrate that it is possible to generate any desired explanation, while leaving the model's accuracy essentially unchanged. This illustrates that GradCAM cannot explain the decision of every CNN and provides a proof of concept showing that it is possible to obfuscate the inner workings of a CNN. Finally, we combine input and model manipulation. To this end we put a backdoor in the network: the explanation is correct unless there is a specific pattern present in the input, which triggers a malicious explanation. Our work raises new security concerns, especially in settings where explanations of models may be used to make decisions, such as in the medical domain.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Why Do You Say It Like That? A Phoneme-Level Framework for Explainable Speech Deepfake Detection

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Phoneme-aligned Grad-CAM on a WavLM-CNN detector reveals significant attack- and speaker-dependent importance of vowels, fricatives and pauses for spoof vs bona-fide decisions on ASVspoof 5.

  2. Impact of Adversarial Attacks on Deep Learning Model Explainability

    cs.LG 2024-12 conditional novelty 4.0 of 10

    Under FGSM and BIM attacks, EfficientNetV2B0 accuracy drops from 89.94% to 58.73% and 45.50%, while explanation IoU and RMSE scores against SAM masks change only slightly.

Pith tools