REVIEW 14 cited by
Fine-Tuning Is All You Need to Mitigate Backdoor Attacks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Backdoor attacks represent one of the major threats to machine learning models. Various efforts have been made to mitigate backdoors. However, existing defenses have become increasingly complex and often require high computational resources or may also jeopardize models' utility. In this work, we show that fine-tuning, one of the most common and easy-to-adopt machine learning training operations, can effectively remove backdoors from machine learning models while maintaining high model utility. Extensive experiments over three machine learning paradigms show that fine-tuning and our newly proposed super-fine-tuning achieve strong defense performance. Furthermore, we coin a new term, namely backdoor sequela, to measure the changes in model vulnerabilities to other attacks before and after the backdoor has been removed. Empirical evaluation shows that, compared to other defense methods, super-fine-tuning leaves limited backdoor sequela. We hope our results can help machine learning model owners better protect their models from backdoor threats. Also, it calls for the design of more advanced attacks in order to comprehensively assess machine learning models' backdoor vulnerabilities.
Forward citations
Cited by 14 Pith papers
-
Follow My Eyes: Backdoor Attacks on Goal-Directed Scanpath Prediction
Scene-conditioned spatial-misdirection and duration-inflation backdoors succeed at 2.5–10% poison ratios on multimodal scanpath predictors and resist five adapted defenses.
-
Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation
BadSem shows that semantic mismatches between images and text can serve as stealthy backdoor triggers for VLMs, achieving near-perfect attack success with low poisoning rates.
-
Backdoor Decontamination Dynamics in LLM Agents
Defensive poisoning followed by unlearning erases most unknown backdoors in LLM tool-calling agents, while trigger recognition can outlive malicious execution.
-
Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models
Backdoors propagate through SLM components with persistence or erasure depending on the targeted part, and poisoned samples are not directly separable from benign ones in shared multitask embeddings.
-
Mitigating Data Exfiltration Attacks through Layer-Wise Learning Rate Decay Fine-Tuning
A layer-wise learning rate decay fine-tuning protocol corrupts steganographically embedded training data in exported medical models while preserving classification utility.
-
BadSR: Stealthy Label Backdoor Attacks on Image Super-Resolution
BadSR creates stealthy poisoned high-resolution labels for super-resolution backdoors, achieving above 80% attack success across five SR models while keeping labels visually close to clean images.
-
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
A cross-lingual paragraph structure, a fixed language-order sequence of segments, can serve as a stealthy backdoor trigger in fine-tuned LLMs, achieving high attack success at 3-5% poisoning.
-
MADE: Graph Backdoor Defense with Masked Unlearning
MADE is a training-set-only graph backdoor defense combining homophily-based poisoned-sample isolation with masked unlearning to drive attack success rate to near zero while keeping accuracy high.
-
PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
PEFTGuard detects backdoored PEFT adapters by classifying transformed low-rank weight matrices, with near-perfect accuracy on the authors' 13,300-adapter PADBench.
-
Bounding-box Watermarking: Defense against Model Extraction Attacks on Object Detectors
A backdoor watermarking scheme for object detectors that poisons bounding-box coordinates in API responses, enabling near-perfect detection of extracted models in several settings.
-
BackdoorMBTI: A Backdoor Learning Multimodal Benchmark Tool Kit for Backdoor Defense Evaluation
BackdoorMBTI is the first backdoor security benchmark and toolkit that covers image, text, and audio modalities with a unified evaluation pipeline.
-
Combining Machine Learning Defenses without Conflicts
A stage-and-risk-based decision rule predicts whether pairs of ML defenses conflict, with reported balanced accuracy of 90% on eight prior combinations and 81-86% on 30 new ones.
-
A Robust Attack: Displacement Backdoor Attack
Displacement Backdoor Attack blends shifted self-copies of an image into the original as a backdoor trigger and reportedly maintains high attack success under data augmentation.
-
A Survey on Privacy Risks and Protection in Large Language Models
The paper surveys LLM privacy leaks and attacks, organizes them into a taxonomy, and reviews defenses without adding new empirical results.
Discussion (0). Continue with ORCID to comment.