Pith. sign in

REVIEW 6 cited by

A Consistent and Efficient Evaluation Strategy for Attribution Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.00449 v2 pith:HMYWV4JQ submitted 2022-02-01 cs.CV cs.LG

classification cs.CVcs.LG
keywords evaluationattributionstrategiesmethodsroaddifferentexpensiveperturbations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With a variety of local feature attribution methods being proposed in recent years, follow-up work suggested several evaluation strategies. To assess the attribution quality across different attribution techniques, the most popular among these evaluation strategies in the image domain use pixel perturbations. However, recent advances discovered that different evaluation strategies produce conflicting rankings of attribution methods and can be prohibitively expensive to compute. In this work, we present an information-theoretic analysis of evaluation strategies based on pixel perturbations. Our findings reveal that the results are strongly affected by information leakage through the shape of the removed pixels as opposed to their actual values. Using our theoretical insights, we propose a novel evaluation framework termed Remove and Debias (ROAD) which offers two contributions: First, it mitigates the impact of the confounders, which entails higher consistency among evaluation strategies. Second, ROAD does not require the computationally expensive retraining step and saves up to 99% in computational costs compared to the state-of-the-art. We release our source code at https://github.com/tleemann/road_evaluation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AIM: Adversarial Information Masking for Faithfulness Evaluation of Saliency Maps

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    AIM is a new saliency-guided adversarial feature replacement method to evaluate faithfulness of saliency maps and reliability of masking operators on image, audio, and EEG tasks.

  2. GRALIS: A Unified Canonical Framework for Linear Attribution Methods via Riesz Representation

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    GRALIS unifies linear XAI attribution methods via a Riesz Representation Theorem-derived canonical form (Q, w, Delta), delivering seven theorems on completeness, convergence, interactions, and multi-scale extensions.

  3. GRALIS: A Unified Canonical Framework for Linear Attribution Methods via Riesz Representation

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    GRALIS supplies a canonical representation (Q, w, Delta) for every additive linear continuous attribution functional on L^2 via the Riesz Representation Theorem, unifying SHAP, IG, LIME and linearized GradCAM while pr...

  4. H-Sets: Hessian-Guided Discovery of Set-Level Feature Interactions in Image Classifiers

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    H-Sets detects higher-order feature interactions in image classifiers via Hessian-guided pair merging and attributes them with IDG-Vis to generate more interpretable saliency maps than existing marginal or coarse methods.

  5. On Spectral Properties of Gradient-based Explanation Methods

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Gradient-based explanations behave like frequency-band selectors: the gradient acts as a high-pass filter, perturbation as a low-pass filter, and their combination creates explanations that shift with the perturbation scale.

  6. Explainable AI needs formalization

    cs.LG 2024-09 unverdicted novelty 5.0 of 10

    Current XAI methods fail because they lack well-defined problems and correctness criteria, causing them to attribute importance to features independent of the prediction target.

Pith tools