REVIEW 2 cited by
Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
As machine learning black boxes are increasingly being deployed in domains such as healthcare and criminal justice, there is growing emphasis on building tools and techniques for explaining these black boxes in an interpretable manner. Such explanations are being leveraged by domain experts to diagnose systematic errors and underlying biases of black boxes. In this paper, we demonstrate that post hoc explanations techniques that rely on input perturbations, such as LIME and SHAP, are not reliable. Specifically, we propose a novel scaffolding technique that effectively hides the biases of any given classifier by allowing an adversarial entity to craft an arbitrary desired explanation. Our approach can be used to scaffold any biased classifier in such a way that its predictions on the input data distribution still remain biased, but the post hoc explanations of the scaffolded classifier look innocuous. Using extensive evaluation with multiple real-world datasets (including COMPAS), we demonstrate how extremely biased (racist) classifiers crafted by our framework can easily fool popular explanation techniques such as LIME and SHAP into generating innocuous explanations which do not reflect the underlying biases.
Forward citations
Cited by 2 Pith papers
-
Evaluating Model Explanations without Ground Truth
AXE evaluates local explanations by k-NN predictiveness of top features, requiring neither ground-truth explanations nor model sensitivity.
-
Enhanced Photonic Chip Design via Interpretable Machine Learning Techniques
Applying LIME explanations to a CNN classifier of photonic multiplexers shows that connected etched regions favor bandwidth, and a V-shaped initial condition informed by this insight yields all 20 redesigned devices w...
Discussion (0). Continue with ORCID to comment.