Pith. sign in

REVIEW 2 cited by

Model extraction from counterfactual explanations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.01884 v1 pith:SYKRN4TR submitted 2020-09-03 cs.LG cs.CRstat.ML

classification cs.LGcs.CRstat.ML
keywords explanationsmodelcounterfactualblack-boxextractiontheyachieveadversary
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Post-hoc explanation techniques refer to a posteriori methods that can be used to explain how black-box machine learning models produce their outcomes. Among post-hoc explanation techniques, counterfactual explanations are becoming one of the most popular methods to achieve this objective. In particular, in addition to highlighting the most important features used by the black-box model, they provide users with actionable explanations in the form of data instances that would have received a different outcome. Nonetheless, by doing so, they also leak non-trivial information about the model itself, which raises privacy issues. In this work, we demonstrate how an adversary can leverage the information provided by counterfactual explanations to build high-fidelity and high-accuracy model extraction attacks. More precisely, our attack enables the adversary to build a faithful copy of a target model by accessing its counterfactual explanations. The empirical evaluation of the proposed attack on black-box models trained on real-world datasets demonstrates that they can achieve high-fidelity and high-accuracy extraction even under low query budgets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do Explanations Increase the Risk of Decision Logic Leakage? Explanation-Guided Stealing of Graph Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    An attacker can exploit explanation heatmaps from deployed graph neural networks to train a surrogate that matches both predictions and highlighted decision logic, outperforming prior stealing attacks.

  2. A Systematic Survey of Model Extraction Attacks and Defenses: State-of-the-Art and Perspectives

    cs.CR 2025-08 conditional novelty 4.0 of 10

    The paper classifies model extraction attacks and defenses into attack, defense, and computing environment categories and surveys their current state.

Pith tools