Pith. sign in

REVIEW 2 cited by

How explainable are adversarially-robust CNNs?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.13042 v2 pith:XZ4PL7QJ submitted 2022-05-25 cs.CV cs.AIcs.HC

classification cs.CVcs.AIcs.HC
keywords cnnsmethodscriteriaexplainabilitymodelsrobustthreevanilla
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Three important criteria of existing convolutional neural networks (CNNs) are (1) test-set accuracy; (2) out-of-distribution accuracy; and (3) explainability. While these criteria have been studied independently, their relationship is unknown. For example, do CNNs that have a stronger out-of-distribution performance have also stronger explainability? Furthermore, most prior feature-importance studies only evaluate methods on 2-3 common vanilla ImageNet-trained CNNs, leaving it unknown how these methods generalize to CNNs of other architectures and training algorithms. Here, we perform the first, large-scale evaluation of the relations of the three criteria using 9 feature-importance methods and 12 ImageNet-trained CNNs that are of 3 training algorithms and 5 CNN architectures. We find several important insights and recommendations for ML practitioners. First, adversarially robust CNNs have a higher explainability score on gradient-based attribution methods (but not CAM-based or perturbation-based methods). Second, AdvProp models, despite being highly accurate more than both vanilla and robust models alone, are not superior in explainability. Third, among 9 feature attribution methods tested, GradCAM and RISE are consistently the best methods. Fourth, Insertion and Deletion are biased towards vanilla and robust models respectively, due to their strong correlation with the confidence score distributions of a CNN. Fifth, we did not find a single CNN to be the best in all three criteria, which interestingly suggests that CNNs are harder to interpret as they become more accurate.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Assessing the Noise Robustness of Class Activation Maps: A Framework for Reliable Model Interpretability

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A consistency-times-responsiveness robustness metric, based on rank-biased overlap of segment rankings, ranks GradCAM++ as most noise-robust and EigenCAM and AblationCAM as least.

  2. Explainable AI in Genomics: Transcription Factor Binding Site Prediction with Mixture of Experts

    cs.LG 2025-07 reject novelty 4.0 of 10

    A Mixture of Experts ensemble of three CNN experts improves out-of-distribution transcription factor binding site prediction, and a shift-averaged gradient method called ShiftSmooth gives more stable motif attributions.

Pith tools