Pith. sign in

REVIEW 14 cited by

Fine-Tuning Is All You Need to Mitigate Backdoor Attacks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.09067 v1 pith:53MGPZZZ submitted 2022-12-18 cs.CR cs.CVcs.LG

classification cs.CRcs.CVcs.LG
keywords backdoorlearningmachinemodelsattacksfine-tuningmodelbackdoors
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Backdoor attacks represent one of the major threats to machine learning models. Various efforts have been made to mitigate backdoors. However, existing defenses have become increasingly complex and often require high computational resources or may also jeopardize models' utility. In this work, we show that fine-tuning, one of the most common and easy-to-adopt machine learning training operations, can effectively remove backdoors from machine learning models while maintaining high model utility. Extensive experiments over three machine learning paradigms show that fine-tuning and our newly proposed super-fine-tuning achieve strong defense performance. Furthermore, we coin a new term, namely backdoor sequela, to measure the changes in model vulnerabilities to other attacks before and after the backdoor has been removed. Empirical evaluation shows that, compared to other defense methods, super-fine-tuning leaves limited backdoor sequela. We hope our results can help machine learning model owners better protect their models from backdoor threats. Also, it calls for the design of more advanced attacks in order to comprehensively assess machine learning models' backdoor vulnerabilities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Follow My Eyes: Backdoor Attacks on Goal-Directed Scanpath Prediction

    cs.CR 2026-04 conditional novelty 7.5 of 10

    Scene-conditioned spatial-misdirection and duration-inflation backdoors succeed at 2.5–10% poison ratios on multimodal scanpath predictors and resist five adapted defenses.

  2. Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    BadSem shows that semantic mismatches between images and text can serve as stealthy backdoor triggers for VLMs, achieving near-perfect attack success with low poisoning rates.

  3. Backdoor Decontamination Dynamics in LLM Agents

    cs.CR 2026-08 conditional novelty 6.0 of 10

    Defensive poisoning followed by unlearning erases most unknown backdoors in LLM tool-calling agents, while trigger recognition can outlive malicious execution.

  4. Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models

    cs.CL 2025-10 unverdicted novelty 6.0 of 10

    Backdoors propagate through SLM components with persistence or erasure depending on the targeted part, and poisoned samples are not directly separable from benign ones in shared multitask embeddings.

  5. Mitigating Data Exfiltration Attacks through Layer-Wise Learning Rate Decay Fine-Tuning

    cs.LG 2025-08 conditional novelty 6.0 of 10

    A layer-wise learning rate decay fine-tuning protocol corrupts steganographically embedded training data in exported medical models while preserving classification utility.

  6. BadSR: Stealthy Label Backdoor Attacks on Image Super-Resolution

    cs.CV 2025-05 conditional novelty 6.0 of 10

    BadSR creates stealthy poisoned high-resolution labels for super-resolution backdoors, achieving above 80% attack success across five SR models while keeping labels visually close to clean images.

  7. CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers

    cs.CR 2024-12 conditional novelty 6.0 of 10

    A cross-lingual paragraph structure, a fixed language-order sequence of segments, can serve as a stealthy backdoor trigger in fine-tuned LLMs, achieving high attack success at 3-5% poisoning.

  8. MADE: Graph Backdoor Defense with Masked Unlearning

    cs.CR 2024-11 conditional novelty 6.0 of 10

    MADE is a training-set-only graph backdoor defense combining homophily-based poisoned-sample isolation with masked unlearning to drive attack success rate to near zero while keeping accuracy high.

  9. PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning

    cs.CR 2024-11 conditional novelty 6.0 of 10

    PEFTGuard detects backdoored PEFT adapters by classifying transformed low-rank weight matrices, with near-perfect accuracy on the authors' 13,300-adapter PADBench.

  10. Bounding-box Watermarking: Defense against Model Extraction Attacks on Object Detectors

    cs.CR 2024-11 conditional novelty 6.0 of 10

    A backdoor watermarking scheme for object detectors that poisons bounding-box coordinates in API responses, enabling near-perfect detection of extracted models in several settings.

  11. BackdoorMBTI: A Backdoor Learning Multimodal Benchmark Tool Kit for Backdoor Defense Evaluation

    cs.CR 2024-11 conditional novelty 6.0 of 10

    BackdoorMBTI is the first backdoor security benchmark and toolkit that covers image, text, and audio modalities with a unified evaluation pipeline.

  12. Combining Machine Learning Defenses without Conflicts

    cs.CR 2024-11 conditional novelty 5.0 of 10

    A stage-and-risk-based decision rule predicts whether pairs of ML defenses conflict, with reported balanced accuracy of 90% on eight prior combinations and 81-86% on 30 new ones.

  13. A Robust Attack: Displacement Backdoor Attack

    cs.CR 2025-02 conditional novelty 4.0 of 10

    Displacement Backdoor Attack blends shifted self-copies of an image into the original as a backdoor trigger and reportedly maintains high attack success under data augmentation.

  14. A Survey on Privacy Risks and Protection in Large Language Models

    cs.CR 2025-05 conditional novelty 2.0 of 10

    The paper surveys LLM privacy leaks and attacks, organizes them into a taxonomy, and reviews defenses without adding new empirical results.

Pith tools