Pith. sign in

hub

Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering

14 Pith papers cite this work. Polarity classification is still indexing.

14 Pith papers citing it
abstract

While machine learning (ML) models are being increasingly trusted to make decisions in different and varying areas, the safety of systems using such models has become an increasing concern. In particular, ML models are often trained on data from potentially untrustworthy sources, providing adversaries with the opportunity to manipulate them by inserting carefully crafted samples into the training set. Recent work has shown that this type of attack, called a poisoning attack, allows adversaries to insert backdoors or trojans into the model, enabling malicious behavior with simple external backdoor triggers at inference time and only a blackbox perspective of the model itself. Detecting this type of attack is challenging because the unexpected behavior occurs only when a backdoor trigger, which is known only to the adversary, is present. Model users, either direct users of training data or users of pre-trained model from a catalog, may not guarantee the safe operation of their ML-based system. In this paper, we propose a novel approach to backdoor detection and removal for neural networks. Through extensive experimental results, we demonstrate its effectiveness for neural networks classifying text and images. To the best of our knowledge, this is the first methodology capable of detecting poisonous data crafted to insert backdoors and repairing the model that does not require a verified and trusted dataset.

hub tools

citation-role summary

baseline 1

citation-polarity summary

years

2026 14

roles

baseline 1

polarities

baseline 1

representative citing papers

SCRUB-FL: Sanitizing and Cleansing Representations via Unlearning of Backdoors

cs.LG · 2026-06-21 · unverdicted · novelty 6.0

SCRUB-FL uses client spectral analysis and WGAN-GP to model suspicious patterns during FL training, aggregates generators server-side, then synthesizes triggers and applies unlearning to reduce backdoor success rates to 3.88% on CIFAR-10/GTSRB while retaining >91% clean accuracy.

DETOUR: A Practical Backdoor Attack against Object Detection

cs.CR · 2026-04-27 · unverdicted · novelty 6.0

DETOUR enables practical backdoor attacks on object detectors by training with rescaled semantic triggers from real-world objects placed at multiple locations to exploit the trigger radiating effect for reliable activation under varying fields of view and spatial configurations.

CSC: Turning the Adversary's Poison against Itself

cs.CR · 2026-04-23 · unverdicted · novelty 6.0

CSC identifies backdoored samples via early-epoch latent clustering and conceals them by relabeling to a virtual class, driving attack success rates near zero on benchmarks with little clean accuracy loss.

citing papers explorer

Showing 14 of 14 citing papers.