Backdoored language models distort predictions on unrelated fine-tuned tasks, collapsing triggered inputs to a single class, and a multi-task correction reduces this distortion without hurting attack success.
Spinning Lan- guage Models: Risks of Propaganda-As-A-Service and Coun- termeasures
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
Backdoored language models distort predictions on unrelated fine-tuned tasks, collapsing triggered inputs to a single class, and a multi-task correction reduces this distortion without hurting attack success.