Pith. sign in

REVIEW 9 cited by

Fine-Tuning Is All You Need to Mitigate Backdoor Attacks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.09067 v1 pith:53MGPZZZ submitted 2022-12-18 cs.CR cs.CVcs.LG

classification cs.CRcs.CVcs.LG
keywords backdoorlearningmachinemodelsattacksfine-tuningmodelbackdoors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Backdoor attacks represent one of the major threats to machine learning models. Various efforts have been made to mitigate backdoors. However, existing defenses have become increasingly complex and often require high computational resources or may also jeopardize models' utility. In this work, we show that fine-tuning, one of the most common and easy-to-adopt machine learning training operations, can effectively remove backdoors from machine learning models while maintaining high model utility. Extensive experiments over three machine learning paradigms show that fine-tuning and our newly proposed super-fine-tuning achieve strong defense performance. Furthermore, we coin a new term, namely backdoor sequela, to measure the changes in model vulnerabilities to other attacks before and after the backdoor has been removed. Empirical evaluation shows that, compared to other defense methods, super-fine-tuning leaves limited backdoor sequela. We hope our results can help machine learning model owners better protect their models from backdoor threats. Also, it calls for the design of more advanced attacks in order to comprehensively assess machine learning models' backdoor vulnerabilities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Follow My Eyes: Backdoor Attacks on Goal-Directed Scanpath Prediction

    cs.CR 2026-04 conditional novelty 7.5 of 10

    Scene-conditioned spatial-misdirection and duration-inflation backdoors succeed at 2.5–10% poison ratios on multimodal scanpath predictors and resist five adapted defenses.

  2. RogueMerge: Robust and Unified Attacks against LLM Model Merging

    cs.CR 2026-06 unverdicted novelty 7.0 of 10

    RogueMerge is a unified attack method that jointly optimizes task vectors to succeed after merging, using stochastic min-max simulation for unknown merging settings and a Taylor-approximated DRO for prompt generalizat...

  3. Follow My Eyes: Backdoor Attacks on Goal-Directed Scanpath Prediction

    cs.CR 2026-04 conditional novelty 7.0 of 10

    Backdoor attacks on VLM-based scanpath predictors can redirect fixations toward chosen objects or inflate durations using input-conditioned triggers that evade cluster detection, and no tested defense blocks them with...

  4. Immunizing 3D Gaussian Generative Models Against Unauthorized Fine-Tuning via Attribute-Space Traps

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    GaussLock embeds traps targeting position, scale, rotation, opacity, and color in 3D Gaussian models to degrade unauthorized fine-tunes while preserving authorized performance.

  5. Hammer and Anvil: Toward a Theory of Backdoors in Federated Learning

    cs.LG 2025-09 conditional novelty 7.0 of 10

    Hammer and Anvil framework categorizes backdoors by update deviation δ and shows that principled combinations of Type-1 outlier/robust and Type-2 removal defenses resist full-information adaptive adversaries.

  6. Quantization as a Malicious Task: Removing Quantization-Conditioned Backdoors via Task Arithmetic

    cs.CR 2026-06 unverdicted novelty 6.0 of 10

    QVec removes quantization-conditioned backdoors by subtracting a structured malicious direction from model weights estimated via one quantization pass and task arithmetic, without retraining or trigger samples.

  7. Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Backdoor attacks on a speech-encoder-plus-LLM pipeline succeed across four tasks and four encoders, and the audio encoder is the main component that carries and propagates the poison.

  8. Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models

    cs.CL 2025-10 unverdicted novelty 6.0 of 10

    Backdoors propagate through SLM components with persistence or erasure depending on the targeted part, and poisoned samples are not directly separable from benign ones in shared multitask embeddings.

  9. Mitigating Data Exfiltration Attacks through Layer-Wise Learning Rate Decay Fine-Tuning

    cs.LG 2025-08 conditional novelty 6.0 of 10

    A layer-wise learning rate decay fine-tuning protocol corrupts steganographically embedded training data in exported medical models while preserving classification utility.

Pith tools