Pith. sign in

REVIEW 1 cited by

Robust Models Are More Interpretable Because Attributions Look Normal

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.11257 v3 pith:P3CWHSIR submitted 2021-03-20 cs.LG cs.CV

classification cs.LGcs.CV
keywords boundariesattributionsrobustboundarydecisioninterpretabilitymodelsnormal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work has found that adversarially-robust deep networks used for image classification are more interpretable: their feature attributions tend to be sharper, and are more concentrated on the objects associated with the image's ground-truth class. We show that smooth decision boundaries play an important role in this enhanced interpretability, as the model's input gradients around data points will more closely align with boundaries' normal vectors when they are smooth. Thus, because robust models have smoother boundaries, the results of gradient-based attribution methods, like Integrated Gradients and DeepLift, will capture more accurate information about nearby decision boundaries. This understanding of robust interpretability leads to our second contribution: \emph{boundary attributions}, which aggregate information about the normal vectors of local decision boundaries to explain a classification outcome. We show that by leveraging the key factors underpinning robust interpretability, boundary attributions produce sharper, more concentrated visual explanations -- even on non-robust models. Any example implementation can be found at \url{https://github.com/zifanw/boundary}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Frequency-Aware Model Parameter Explorer: A new attribution method for improving explainability

    cs.LG 2025-09 unverdicted novelty 5.0 of 10

    FAMPE is a new attribution method that applies FFT-based frequency-selective perturbations integrated with model parameter exploration to produce fine-grained feature importance maps, showing gains over AttEXplore on ...

Pith tools