A retraining-free backdoor attack that prunes the least important attention head and injects a pre-trained malicious head achieves over 99.5% attack success in the paper's experiments while evading four defenses.
https://cdn.openai.com/research-covers/language-unsupervised/ language_understanding_paper.pdf
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models
A retraining-free backdoor attack that prunes the least important attention head and injects a pre-trained malicious head achieves over 99.5% attack success in the paper's experiments while evading four defenses.