Pith. sign in

REVIEW 1 cited by

Why You Should Not Trust Interpretations in Machine Learning: Adversarial Attacks on Partial Dependence Plots

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.18702 v2 pith:6FESK3UR submitted 2024-04-29 cs.LG cs.CRstat.APstat.ML

classification cs.LGcs.CRstat.APstat.ML
keywords modelplotsadversarialblack-boxframeworkinterpretationoriginalpredictions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The adoption of artificial intelligence (AI) across industries has led to the widespread use of complex black-box models and interpretation tools for decision making. This paper proposes an adversarial framework to uncover the vulnerability of permutation-based interpretation methods for machine learning tasks, with a particular focus on partial dependence (PD) plots. This adversarial framework modifies the original black box model to manipulate its predictions for instances in the extrapolation domain. As a result, it produces deceptive PD plots that can conceal discriminatory behaviors while preserving most of the original model's predictions. This framework can produce multiple fooled PD plots via a single model. By using real-world datasets including an auto insurance claims dataset and COMPAS (Correctional Offender Management Profiling for Alternative Sanctions) dataset, our results show that it is possible to intentionally hide the discriminatory behavior of a predictor and make the black-box model appear neutral through interpretation tools like PD plots while retaining almost all the predictions of the original black-box model. Managerial insights for regulators and practitioners are provided based on the findings.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Shapley Values: Cooperative Games for the Interpretation of Machine Learning Models

    stat.ML 2025-06 conditional novelty 3.0 of 10

    A review-style position paper that separates the choice of value function from the choice of an efficient allocation, presenting Weber and Harsanyi sets as generalizations of Shapley values.

Pith tools