REVIEW 7 cited by
A Consistent and Efficient Evaluation Strategy for Attribution Methods
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
With a variety of local feature attribution methods being proposed in recent years, follow-up work suggested several evaluation strategies. To assess the attribution quality across different attribution techniques, the most popular among these evaluation strategies in the image domain use pixel perturbations. However, recent advances discovered that different evaluation strategies produce conflicting rankings of attribution methods and can be prohibitively expensive to compute. In this work, we present an information-theoretic analysis of evaluation strategies based on pixel perturbations. Our findings reveal that the results are strongly affected by information leakage through the shape of the removed pixels as opposed to their actual values. Using our theoretical insights, we propose a novel evaluation framework termed Remove and Debias (ROAD) which offers two contributions: First, it mitigates the impact of the confounders, which entails higher consistency among evaluation strategies. Second, ROAD does not require the computationally expensive retraining step and saves up to 99% in computational costs compared to the state-of-the-art. We release our source code at https://github.com/tleemann/road_evaluation.
Forward citations
Cited by 7 Pith papers
-
Saliency Methods are Encoders: Analysing Logical Relations Towards Interpretation
On synthetic AND/OR/XOR tasks, saliency methods frequently rank irrelevant baseline inputs above logically relevant ones, and retrained models can recover class information from masked inputs, indicating that score or...
-
On Spectral Properties of Gradient-based Explanation Methods
Gradient-based explanations behave like frequency-band selectors: the gradient acts as a high-pass filter, perturbation as a low-pass filter, and their combination creates explanations that shift with the perturbation scale.
-
A Super-pixel-based Approach to the Stable Interpretation of Neural Networks
Averaging saliency values within super-pixel groups reduces the variance and improves the stability and generalizability of gradient-based interpretation maps.
-
Derivative-Free Diffusion Manifold-Constrained Gradient for Unified XAI
FreeMCG estimates on-manifold classifier gradients from black-box outputs using diffusion particles and an ensemble Kalman filter, and uses the same estimate for feature attribution and counterfactual explanation.
-
Advancing Attribution-Based Neural Network Explainability through Relative Absolute Magnitude Layer-Wise Relevance Propagation and Multi-Component Evaluation
A new LRP rule that scales attributions by absolute activation magnitude, plus a unified evaluation metric, is tested across three architectures and two datasets.
-
Saliency Maps are Ambiguous: Analysis of Logical Relations on First and Second Order Attributions
On synthetic AND/OR/XOR datasets with perfectly accurate models, every tested saliency method sometimes ranks a truly irrelevant input above a necessary one, so the scores cannot be trusted as relevance rankings.
-
Explainability for Vision Foundation Models: A Survey
A structured review of 122 papers on explainability for vision foundation models, with a taxonomy and the finding that quantitative evaluation is rare (36%).
Discussion (0). Continue with ORCID to comment.