A clean-label backdoor hidden across multiple classes is activated and amplified by unlearning clean samples, reaching attack success above 90 percent after forgetting.
Silent Killer: A Stealthy, Clean-Label, Black-Box Backdoor Attack
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Backdoor poisoning attacks pose a well-known risk to neural networks. However, most studies have focused on lenient threat models. We introduce Silent Killer, a novel attack that operates in clean-label, black-box settings, uses a stealthy poison and trigger and outperforms existing methods. We investigate the use of universal adversarial perturbations as triggers in clean-label attacks, following the success of such approaches under poison-label settings. We analyze the success of a naive adaptation and find that gradient alignment for crafting the poison is required to ensure high success rates. We conduct thorough experiments on MNIST, CIFAR10, and a reduced version of ImageNet and achieve state-of-the-art results.
citation-role summary
citation-polarity summary
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
When Forgetting Triggers Backdoors: A Clean Unlearning Attack
A clean-label backdoor hidden across multiple classes is activated and amplified by unlearning clean samples, reaching attack success above 90 percent after forgetting.