Pith. sign in

REVIEW 2 cited by

Challenging common interpretability assumptions in feature attribution explanations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.02748 v1 pith:RKBJXMEF submitted 2020-12-04 cs.LG cs.CYcs.HC

classification cs.LGcs.CYcs.HC
keywords interpretabilityexperimentassumptionsattributioncommondecisionevaluationexplanations
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As machine learning and algorithmic decision making systems are increasingly being leveraged in high-stakes human-in-the-loop settings, there is a pressing need to understand the rationale of their predictions. Researchers have responded to this need with explainable AI (XAI), but often proclaim interpretability axiomatically without evaluation. When these systems are evaluated, they are often tested through offline simulations with proxy metrics of interpretability (such as model complexity). We empirically evaluate the veracity of three common interpretability assumptions through a large scale human-subjects experiment with a simple "placebo explanation" control. We find that feature attribution explanations provide marginal utility in our task for a human decision maker and in certain cases result in worse decisions due to cognitive and contextual confounders. This result challenges the assumed universal benefit of applying these methods and we hope this work will underscore the importance of human evaluation in XAI research. Supplemental materials -- including anonymized data from the experiment, code to replicate the study, an interactive demo of the experiment, and the models used in the analysis -- can be found at: https://doi.pizza/challenging-xai.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Probabilistic Stability Guarantees for Feature Attributions

    cs.LG 2025-04 conditional novelty 6.0 of 10

    Soft stability measures the probability that an explanation's prediction survives additive feature perturbations, and a sampling algorithm certifies this rate with statistical guarantees.

  2. DLBacktrace: A Model Agnostic Explainability for any Deep Learning Models

    cs.LG 2024-11 reject novelty 3.0 of 10

    DLBacktrace is a backward relevance-propagation technique that resembles existing layer-wise relevance propagation (LRP) but is presented as novel, with benchmarks that omit the closest competitor.

Pith tools