Pith. sign in

REVIEW 2 cited by

Certifiably Robust Interpretation in Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.12105 v3 pith:3RZASWEF submitted 2019-05-28 cs.LG stat.ML

classification cs.LGstat.ML
keywords interpretationdeeplearningcertifiablyrobustadversarialmapsmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep learning interpretation is essential to explain the reasoning behind model predictions. Understanding the robustness of interpretation methods is important especially in sensitive domains such as medical applications since interpretation results are often used in downstream tasks. Although gradient-based saliency maps are popular methods for deep learning interpretation, recent works show that they can be vulnerable to adversarial attacks. In this paper, we address this problem and provide a certifiable defense method for deep learning interpretation. We show that a sparsified version of the popular SmoothGrad method, which computes the average saliency maps over random perturbations of the input, is certifiably robust against adversarial perturbations. We obtain this result by extending recent bounds for certifiably robust smooth classifiers to the interpretation setting. Experiments on ImageNet samples validate our theory.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Provably Robust Explainable Graph Neural Networks against Graph Perturbation Attacks

    cs.CR 2025-02 conditional novelty 6.0 of 10

    XGNNCert certifies that a GNN explanation keeps at least lambda edges under any graph perturbation of bounded size, using majority voting over hash-partitioned hybrid subgraphs.

  2. A Super-pixel-based Approach to the Stable Interpretation of Neural Networks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Averaging saliency values within super-pixel groups reduces the variance and improves the stability and generalizability of gradient-based interpretation maps.

Pith tools