REVIEW 3 cited by
Rethinking Stability for Attribution-based Explanations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As attribution-based explanation methods are increasingly used to establish model trustworthiness in high-stakes situations, it is critical to ensure that these explanations are stable, e.g., robust to infinitesimal perturbations to an input. However, previous works have shown that state-of-the-art explanation methods generate unstable explanations. Here, we introduce metrics to quantify the stability of an explanation and show that several popular explanation methods are unstable. In particular, we propose new Relative Stability metrics that measure the change in output explanation with respect to change in input, model representation, or output of the underlying predictor. Finally, our experimental evaluation with three real-world datasets demonstrates interesting insights for seven explanation methods and different stability metrics.
Forward citations
Cited by 3 Pith papers
-
Value bounds and Convergence Analysis for Averages of LRP attributions
Averaged LRP-beta attributions have Hoeffding convergence bounds independent of weight norms, unlike gradient-based explanations.
-
Assessing the Noise Robustness of Class Activation Maps: A Framework for Reliable Model Interpretability
A consistency-times-responsiveness robustness metric, based on rank-biased overlap of segment rankings, ranks GradCAM++ as most noise-robust and EigenCAM and AblationCAM as least.
-
VARSHAP: Addressing Global Dependency Problems in Explainable AI with Variance-Based Local Feature Attribution
VARSHAP defines local feature attribution as the Shapley value of a variance-reduction game and claims greater stability than SHAP and LIME.
Discussion (0). Continue with ORCID to comment.