Pith. sign in

REVIEW 3 cited by

Prediction Poisoning: Towards Defenses Against DNN Model Stealing Attacks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.10908 v2 pith:WCEE7YYF submitted 2019-06-26 cs.LG cs.CRcs.CVstat.ML

classification cs.LGcs.CRcs.CVstat.ML
keywords attacksstealingmodeldefensesdefenseapplicationsattackerexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

High-performance Deep Neural Networks (DNNs) are increasingly deployed in many real-world applications e.g., cloud prediction APIs. Recent advances in model functionality stealing attacks via black-box access (i.e., inputs in, predictions out) threaten the business model of such applications, which require a lot of time, money, and effort to develop. Existing defenses take a passive role against stealing attacks, such as by truncating predicted information. We find such passive defenses ineffective against DNN stealing attacks. In this paper, we propose the first defense which actively perturbs predictions targeted at poisoning the training objective of the attacker. We find our defense effective across a wide range of challenging datasets and DNN model stealing attacks, and additionally outperforms existing defenses. Our defense is the first that can withstand highly accurate model stealing attacks for tens of thousands of queries, amplifying the attacker's error rate up to a factor of 85$\times$ with minimal impact on the utility for benign users.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ADS-C: Antidistillation Sampling for Classification

    cs.LG 2026-07 accept novelty 7.0 of 10

    ADS-C perturbs served classification probabilities under a per-input margin budget, preserving every top-1 prediction while degrading distilled students by 13–30 percentage points.

  2. MISLEADER: Defending against Model Extraction with Ensembles of Distilled Models

    cs.CR 2025-06 reject novelty 5.0 of 10

    MISLEADER trains an ensemble of distilled models that stay accurate for benign users but output misleading predictions on augmented inputs, reducing clone model accuracy in model extraction attacks.

  3. A Systematic Survey of Model Extraction Attacks and Defenses: State-of-the-Art and Perspectives

    cs.CR 2025-08 conditional novelty 4.0 of 10

    The paper classifies model extraction attacks and defenses into attack, defense, and computing environment categories and surveys their current state.

Pith tools