Pith. sign in

REVIEW 4 cited by

STRIP: A Defence Against Trojan Attacks on Deep Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1902.06531 v2 pith:YMSKOUFN submitted 2019-02-18 cs.CR

classification cs.CR
keywords trojanattacksattackinputsmodelstripattackerbenign
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

A recent trojan attack on deep neural network (DNN) models is one insidious variant of data poisoning attacks. Trojan attacks exploit an effective backdoor created in a DNN model by leveraging the difficulty in interpretability of the learned model to misclassify any inputs signed with the attacker's chosen trojan trigger. Since the trojan trigger is a secret guarded and exploited by the attacker, detecting such trojan inputs is a challenge, especially at run-time when models are in active operation. This work builds STRong Intentional Perturbation (STRIP) based run-time trojan attack detection system and focuses on vision system. We intentionally perturb the incoming input, for instance by superimposing various image patterns, and observe the randomness of predicted classes for perturbed inputs from a given deployed model---malicious or benign. A low entropy in predicted classes violates the input-dependence property of a benign model and implies the presence of a malicious input---a characteristic of a trojaned input. The high efficacy of our method is validated through case studies on three popular and contrasting datasets: MNIST, CIFAR10 and GTSRB. We achieve an overall false acceptance rate (FAR) of less than 1%, given a preset false rejection rate (FRR) of 1%, for different types of triggers. Using CIFAR10 and GTSRB, we have empirically achieved result of 0% for both FRR and FAR. We have also evaluated STRIP robustness against a number of trojan attack variants and adaptive attacks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Invisible Backdoor Attacks on Deep Neural Networks via Steganography and Regularization

    cs.CR 2019-09 conditional novelty 6.0 of 10

    Backdoor triggers hidden via LSB steganography or Lp-regularized noise achieve high attack success while looking nearly identical to clean images in perceptual metrics, and single-target variants evade Neural Cleanse.

  2. TABOR: A Highly Accurate Approach to Inspecting and Restoring Trojan Backdoors in AI Systems

    cs.CR 2019-08 conditional novelty 6.0 of 10

    Adding explainability- and heuristic-based regularizers plus a new trigger-quality metric to trigger-recovery optimization detects trojan backdoors more accurately and restores them more faithfully than Neural Cleanse...

  3. From Detection to Correction: Backdoor-Resilient Face Recognition via Vision-Language Trigger Detection and Noise-Based Neutralization

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    A majority vote of large vision-language models is claimed to detect backdoor triggers in face images, with calibrated noise correcting poisoned samples at 100% accuracy.

  4. Identifying Physically Realizable Triggers for Backdoored Face Recognition Networks

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A trigger-inversion plus object-retrieval pipeline finds physically realizable backdoor triggers in face recognition networks without poisoned examples.

Pith tools