REVIEW 68 references
Secure Transfer Learning: Training Clean Models Against Backdoor in (Both) Pre-trained Encoders and Downstream Datasets
T0 review · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read T-Core, a bootstrapping defense that sifts clean data and filters trusted encoder channels, reduces backdoor attack success rates below 10% across encoder and dataset poisoning threats in transfer learning.
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
T-Core flips the strategy. Instead of hunting for poison, it first picks out a small set of high-confidence clean examples. It does this by checking whether a sample stays in the same majority group across several layers of the network, which poisoned samples usually fail to do. It then expands this clean seed set, filters the encoder's neurons to keep only those that support clean recognition, and gradually trains a model on the growing clean pool, leaving suspicious samples out.
The paper evaluates T-Core against five encoder poisoning attacks, seven dataset poisoning attacks, and three combined threat scenarios on five image datasets. In most tested cases it keeps attack success rates under 10% while preserving accuracy. The main open points are that the authors do not release code or error bars, and an adaptive attacker with a large perturbation budget can still break the first sifting step.
Extended reading notes
Core claim
T-Core defends against all three backdoor threats in transfer learning (encoder poisoning, dataset poisoning, and adaptive poisoning), keeping attack success rates below 10% while maintaining accuracy, even when defenders only fine-tune the classification head and a small fraction of encoder parameters. The paper states: 'T-Core effectively defends against all considered backdoor threats, as shown in Tables 9 to 11. Specifically, the attack success rates for Threat-1, Threat-2, and Threat-3 are all below 10%.'
Load-bearing premise
The initial seed sifting (TIS) assumes poisoned samples are a minority in each class and are topologically inconsistent across network layers (Majority Rule and Consistency Rule, Sec 4.1.1). If an attacker crafts triggers that place poisoned samples in the majority clusters with consistent neighbors, the clean seed set is contaminated and the whole bootstrapping pipeline inherits the backdoor. The adaptive attack in Sec 5.6.1 demonstrates this: at l-infinity perturbation budget 16, TIS selects 38 poisoned samples as clean and ASR jumps to 92.5%.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
free parameters (9)
- alpha (seed data proportion) =
1%
- m (number of nearest neighbors) =
50
- rexpand (seed expansion ratio) =
5%
- Dsub size before encoder filtering =
10% of D
- ACCmin (unlearning stopping accuracy) =
20%
- channel filtering threshold (keep ratio) =
90%
- gamma1, gamma2, gamma3 (bootstrapping selection rates) =
2%, 2%, 5%
- rho (bootstrapping halt ratio) =
90%
- L (number of layers considered in TIS) =
3
assumptions (4)
- domain assumption Majority Rule: poisoned samples are a minority within each class and form smaller clusters across all considered layers.
- domain assumption Consistency Rule: clean samples have consistent nearest neighbors from the same class across different DNN layers, whereas poisoned samples do not.
- domain assumption After confusion training with a small clean seed, the largest-loss region of the training set is populated almost exclusively by clean samples.
- domain assumption Clean hard examples are harder to learn than easily inserted backdoor-poisoned examples, and any mistakenly selected poisoned samples require significantly more data and training to become effective.
Cite this review
Pith. "Pith review of Secure Transfer Learning: Training Clean Models Against Backdoor in (Both) Pre-trained Encoders and Downstream Datasets." pith.science (2026). https://pith.science/paper/CIJJ3O5R
@misc{pith2026250411990,
author = {Pith},
title = {Pith review of: Secure Transfer Learning: Training Clean Models Against Backdoor in (Both) Pre-trained Encoders and Downstream Datasets},
year = {2026},
howpublished = {\url{https://pith.science/paper/CIJJ3O5R}},
note = {Machine review of arXiv:2504.11990}
}
read the original abstract
Transfer learning from pre-trained encoders has become essential in modern machine learning, enabling efficient model adaptation across diverse tasks. However, this combination of pre-training and downstream adaptation creates an expanded attack surface, exposing models to sophisticated backdoor embeddings at both the encoder and dataset levels--an area often overlooked in prior research. Additionally, the limited computational resources typically available to users of pre-trained encoders constrain the effectiveness of generic backdoor defenses compared to end-to-end training from scratch. In this work, we investigate how to mitigate potential backdoor risks in resource-constrained transfer learning scenarios. Specifically, we conduct an exhaustive analysis of existing defense strategies, revealing that many follow a reactive workflow based on assumptions that do not scale to unknown threats, novel attack types, or different training paradigms. In response, we introduce a proactive mindset focused on identifying clean elements and propose the Trusted Core (T-Core) Bootstrapping framework, which emphasizes the importance of pinpointing trustworthy data and neurons to enhance model security. Our empirical evaluations demonstrate the effectiveness and superiority of T-Core, specifically assessing 5 encoder poisoning attacks, 7 dataset poisoning attacks, and 14 baseline defenses across five benchmark datasets, addressing four scenarios of 3 potential backdoor threats.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Badnets: Evaluating backdooring attacks on deep neural networks,
T. Gu, K. Liu, B. Dolan-Gavitt, and S. Garg, “Badnets: Evaluating backdooring attacks on deep neural networks,” IEEE Access, vol. 7, pp. 47230–47244, 2019
work page 2019
-
[2]
Targeted backdoor attacks on deep learning systems using data poisoning,
X. Chen, C. Liu, B. Li, K. Lu, and D. Song, “Targeted backdoor attacks on deep learning systems using data poisoning,”arXiv preprint arXiv:1712.05526, 2017
arXiv 2017
-
[3]
Badclip: Dual-embedding guided backdoor attack on multimodal contrastive learning,
S. Liang, M. Zhu, A. Liu, B. Wu, X. Cao, and E.-C. Chang, “Badclip: Dual-embedding guided backdoor attack on multimodal contrastive learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024
work page 2024
-
[4]
BadEncoder: Backdoor Attacks to Pre-trained Encoders in Self-Supervised Learning,
J. Jia, Y . Liu, and N. Z. Gong, “BadEncoder: Backdoor Attacks to Pre-trained Encoders in Self-Supervised Learning,” in 2022 IEEE Symposium on Security and Privacy (SP 22), pp. 2043–2059, 2022
work page 2022
-
[5]
Backdoor Attacks Against Transfer Learning With Pre-Trained Deep Learning Models,
S. Wang, S. Nepal, C. Rudolph, M. Grobler, S. Chen, and T. Chen, “Backdoor Attacks Against Transfer Learning With Pre-Trained Deep Learning Models,” IEEE Transactions on Services Computing, vol. 15, pp. 1526–1539, May 2022
work page 2022
-
[6]
An em- barrassingly simple backdoor attack on self-supervised learning,
C. Li, R. Pang, Z. Xi, T. Du, S. Ji, Y . Yao, and T. Wang, “An em- barrassingly simple backdoor attack on self-supervised learning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 4367–4378, October 2023
work page 2023
-
[7]
Dis- tribution Preserving Backdoor Attack in Self-supervised Learning,
G. Tao, Z. Wang, S. Feng, G. Shen, S. Ma, and X. Zhang, “Dis- tribution Preserving Backdoor Attack in Self-supervised Learning,” in 2024 IEEE Symposium on Security and Privacy (SP 24), IEEE Computer Society, 2024. ISSN: 2375-1207
work page 2024
-
[8]
Back- door attacks on self-supervised learning,
A. Saha, A. Tejankar, S. A. Koohpayegani, and H. Pirsiavash, “Back- door attacks on self-supervised learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 22), pp. 13337–13346, 2022
work page 2022
Show all 68 references
-
[9]
CorruptEncoder: Data poi- soning based backdoor attacks to contrastive learning,
J. Zhang, H. Liu, J. Jia, and N. Z. Gong, “CorruptEncoder: Data poi- soning based backdoor attacks to contrastive learning,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 24), 2024
2024
-
[10]
A new backdoor attack in cnns by training set corruption without label poisoning,
M. Barni, K. Kallas, and B. Tondi, “A new backdoor attack in cnns by training set corruption without label poisoning,” in 2019 IEEE International Conference on Image Processing (ICIP 19), pp. 101– 105, IEEE, 2019
2019
-
[11]
Demon in the variant: Statistical analysis of DNNs for robust backdoor contamination de- tection,
D. Tang, X. Wang, H. Tang, and K. Zhang, “Demon in the variant: Statistical analysis of DNNs for robust backdoor contamination de- tection,” in 30th USENIX Security Symposium (USENIX Security 21), pp. 1541–1558, 2021
2021
-
[12]
Revisiting the assumption of latent separability for backdoor defenses,
X. Qi, T. Xie, Y . Li, S. Mahloujifar, and P. Mittal, “Revisiting the assumption of latent separability for backdoor defenses,” in The eleventh International Conference on Learning Representations (ICLR 23), 2023
2023
-
[13]
WaNet - imperceptible warping- based backdoor attack,
T. A. Nguyen and A. T. Tran, “WaNet - imperceptible warping- based backdoor attack,” in International Conference on Learning Representations (ICLR 21), 2021
2021
-
[14]
Backdoor defense via deconfounded representation learning,
Z. Zhang, Q. Liu, Z. Wang, Z. Lu, and Q. Hu, “Backdoor defense via deconfounded representation learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 23), pp. 12228–12238, 2023
2023
-
[15]
Anti-Backdoor Learning: Training Clean Models on Poisoned Data,
Y . Li, X. Lyu, N. Koren, L. Lyu, B. Li, and X. Ma, “Anti-Backdoor Learning: Training Clean Models on Poisoned Data,” in Advances in Neural Information Processing Systems (NeurIPS 21), vol. 34, pp. 14900–14912, Curran Associates, Inc., 2021
2021
-
[16]
SPECTRE: defending against backdoor attacks using robust statistics,
J. Hayase, W. Kong, R. Somani, and S. Oh, “SPECTRE: defending against backdoor attacks using robust statistics,” in Proceedings of the 38th International Conference on Machine Learning (ICML 21), 2021
2021
-
[17]
Data-free backdoor re- moval based on channel lipschitzness,
R. Zheng, R. Tang, J. Li, and L. Liu, “Data-free backdoor re- moval based on channel lipschitzness,” in European Conference on Computer Vision (ECCV 22), pp. 175–191, Springer, 2022
2022
-
[18]
Detecting backdoor attacks on deep neural networks by activation clustering,
B. Chen, W. Carvalho, N. Baracaldo, H. Ludwig, B. Edwards, T. Lee, I. Molloy, and B. Srivastava, “Detecting backdoor attacks on deep neural networks by activation clustering,” arXiv preprint arXiv:1811.03728, 2018
2018 arXiv
-
[19]
Spectral signatures in backdoor attacks,
B. Tran, J. Li, and A. Madry, “Spectral signatures in backdoor attacks,” in Advances in Neural Information Processing Systems (NeurIPS 18), pp. 8000–8010, 2018
2018
-
[20]
ASSET: Robust back- door data detection across a multiplicity of deep learning paradigms,
M. Pan, Y . Zeng, L. Lyu, X. Lin, and R. Jia, “ASSET: Robust back- door data detection across a multiplicity of deep learning paradigms,” in 32nd USENIX Security Symposium (USENIX Security 23), pp. 2725–2742, 2023
2023
-
[21]
Towards a proactive ML approach for detecting backdoor poison samples,
X. Qi, T. Xie, J. T. Wang, T. Wu, S. Mahloujifar, and P. Mittal, “Towards a proactive ML approach for detecting backdoor poison samples,” in 32nd USENIX Security Symposium (USENIX Security 23), pp. 1685–1702, 2023
2023
-
[22]
STRIP: A defence against trojan attacks on deep neural networks,
Y . Gao, C. Xu, D. Wang, S. Chen, D. C. Ranasinghe, and S. Nepal, “STRIP: A defence against trojan attacks on deep neural networks,” in Proceedings of the 35th Annual Computer Security Applications Conference (ACSAC 19), pp. 113–125, 2019
2019
-
[23]
IBD-PSC: Input-level backdoor detection via parameter-oriented scaling consis- tency,
L. Hou, R. Feng, Z. Hua, W. Luo, L. Y . Zhang, and Y . Li, “IBD-PSC: Input-level backdoor detection via parameter-oriented scaling consis- tency,” in Forty-first International Conference on Machine Learning (ICML 24), 2024
2024
-
[24]
SCALE- UP: An efficient black-box input-level backdoor detection via ana- lyzing scaled prediction consistency,
J. Guo, Y . Li, X. Chen, H. Guo, L. Sun, and C. Liu, “SCALE- UP: An efficient black-box input-level backdoor detection via ana- lyzing scaled prediction consistency,” in The Eleventh International Conference on Learning Representations (ICLR 23), 2023
2023
-
[25]
Adversarial unlearning of backdoors via implicit hypergradient,
Y . Zeng, S. Chen, W. Park, Z. Mao, M. Jin, and R. Jia, “Adversarial unlearning of backdoors via implicit hypergradient,” in International Conference on Learning Representations (ICLR 22), 2022
2022
-
[26]
Enhancing fine-tuning based backdoor defense with sharpness-aware minimization,
M. Zhu, S. Wei, L. Shen, Y . Fan, and B. Wu, “Enhancing fine-tuning based backdoor defense with sharpness-aware minimization,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 23), 2023
2023
-
[27]
On the effectiveness of distil- lation in mitigating backdoors in pre-trained encoder,
T. Han, S. Huang, Z. Ding, W. Sun, Y . Feng, C. Fang, J. Li, H. Qian, C. Wu, Q. Zhang, et al., “On the effectiveness of distil- lation in mitigating backdoors in pre-trained encoder,” arXiv preprint arXiv:2403.03846, 2024
2024
-
[28]
Februus: Input purification defense against trojan attacks on deep neural network systems,
B. G. Doan, E. Abbasnejad, and D. C. Ranasinghe, “Februus: Input purification defense against trojan attacks on deep neural network systems,” in Proceedings of the 36th Annual Computer Security Applications Conference (ACSAC 20), pp. 897–912, 2020
2020
-
[29]
Backdoor attack in the physical world,
Y . Li, T. Zhai, Y . Jiang, Z. Li, and S.-T. Xia, “Backdoor attack in the physical world,” arXiv preprint arXiv:2104.02361, 2021
2021 arXiv
-
[30]
Deepsweep: An evaluation framework for mitigating dnn back- door attacks using data augmentation,
H. Qiu, Y . Zeng, S. Guo, T. Zhang, M. Qiu, and B. Thuraisingham, “Deepsweep: An evaluation framework for mitigating dnn back- door attacks using data augmentation,” in Proceedings of the 2021 ACM Asia Conference on Computer and Communications Security (AsiaCCS 21), pp. 363–377, 2021
2021
-
[31]
Detecting backdoors during the inference stage based on corruption robustness consistency,
X. Liu, M. Li, H. Wang, S. Hu, D. Ye, H. Jin, L. Wu, and C. Xiao, “Detecting backdoors during the inference stage based on corruption robustness consistency,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 23), pp. 16363– 16372, 2023
2023
-
[32]
Robust backdoor detection for deep learning via topo- logical evolution dynamics,
X. Mo, Y . Zhang, L. Zhang, W. Luo, N. Sun, S. Hu, S. Gao, and Y . Xiang, “Robust backdoor detection for deep learning via topo- logical evolution dynamics,” in 2024 IEEE Symposium on Security and Privacy (SP 24), (Los Alamitos, CA, USA), pp. 174–174, IEEE Computer Society, may 2024
2024
-
[33]
The ”Beat- rix
W. Ma, D. Wang, R. Sun, M. Xue, S. Wen, and Y . Xiang, “The ”Beat- rix” Resurrections: Robust backdoor detection via gram matrices,” in NDSS, 2023
2023
-
[34]
Backdoor defense via decoupling the training process,
K. Huang, Y . Li, B. Wu, Z. Qin, and K. Ren, “Backdoor defense via decoupling the training process,” in Proceedings of the International Conference on Learning Representations (ICLR 22), 2022
2022
-
[35]
Training with more confi- dence: Mitigating injected and natural backdoors during training,
Z. Wang, H. Ding, J. Zhai, and S. Ma, “Training with more confi- dence: Mitigating injected and natural backdoors during training,” in Advances in Neural Information Processing Systems (NeurIPS 22) (A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, eds.), 2022
2022
-
[36]
Progressive poisoned data isolation for training-time backdoor defense,
Y . Chen, H. Wu, and J. Zhou, “Progressive poisoned data isolation for training-time backdoor defense,” in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 24), pp. 11425–11433, 2024
2024
-
[37]
Fine-Pruning: Defending against backdooring attacks on deep neural networks,
K. Liu, B. Dolan-Gavitt, and S. Garg, “Fine-Pruning: Defending against backdooring attacks on deep neural networks,” inInternational Symposium on Research in Attacks, Intrusions and Defenses (RAID 18), pp. 273–294, Springer, 2018
2018
-
[38]
Towards inspecting and eliminating trojan backdoors in deep neural networks,
W. Guo, L. Wang, Y . Xu, X. Xing, M. Du, and D. Song, “Towards inspecting and eliminating trojan backdoors in deep neural networks,” in 2020 IEEE International Conference on Data Mining (ICDM 20), IEEE, 2020
2020
-
[39]
Backdoor scanning for deep neural networks through k- arm optimization,
G. Shen, Y . Liu, G. Tao, S. An, Q. Xu, S. Cheng, S. Ma, and X. Zhang, “Backdoor scanning for deep neural networks through k- arm optimization,” in International Conference on Machine Learning (ICML 21), pp. 9525–9536, PMLR, 2021
2021
-
[40]
Neural Cleanse: Identifying and mitigating backdoor attacks in neural networks,
B. Wang, Y . Yao, S. Shan, H. Li, B. Viswanath, H. Zheng, and B. Y . Zhao, “Neural Cleanse: Identifying and mitigating backdoor attacks in neural networks,” in 2019 IEEE Symposium on Security and Privacy (SP 19), pp. 707–723, May 2019
2019
-
[41]
Detecting backdoors in pre-trained encoders,
S. Feng, G. Tao, S. Cheng, G. Shen, X. Xu, Y . Liu, K. Zhang, S. Ma, and X. Zhang, “Detecting backdoors in pre-trained encoders,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 23), pp. 16352–16362, 2023
2023
-
[42]
Backdoor defense via adaptively splitting poisoned dataset,
K. Gao, Y . Bai, J. Gu, Y . Yang, and S.-T. Xia, “Backdoor defense via adaptively splitting poisoned dataset,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 23), 2023
2023
-
[43]
Model-agnostic meta-learning for fast adaptation of deep networks,
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in International Conference on Machine Learning (ICML 17), pp. 1126–1135, PMLR, 2017
2017
-
[44]
Learning multiple layers of features from tiny im- ages,
A. Krizhevsky, “Learning multiple layers of features from tiny im- ages,” Tech Report, 2009
2009
-
[45]
Man vs. computer: Benchmarking machine learning algorithms for traffic sign recognition,
J. Stallkamp, M. Schlipsing, J. Salmen, and C. Igel, “Man vs. computer: Benchmarking machine learning algorithms for traffic sign recognition,” Neural networks, vol. 32, pp. 323–332, 2012
2012
-
[46]
Reading digits in natural images with unsupervised feature learning,
Y . Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y . Ng, “Reading digits in natural images with unsupervised feature learning,” in NIPS Workshop on Deep Learning and Unsupervised Feature Learning, 2011
2011
-
[47]
An analysis of single-layer networks in unsupervised feature learning,
A. Coates, A. Ng, and H. Lee, “An analysis of single-layer networks in unsupervised feature learning,” in AISTATS, 2011
2011
-
[48]
ImageNet: A large-scale hierarchical image database,
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 09), pp. 248–255, IEEE, 2009
2009
-
[49]
META- SIFT : How to sift out a clean subset in the presence of data poi- soning?,
Y . Zeng, M. Pan, H. Jahagirdar, M. Jin, L. Lyu, and R. Jia, “META- SIFT : How to sift out a clean subset in the presence of data poi- soning?,” in 32nd USENIX Security Symposium (USENIX Security 23), pp. 1667–1684, 2023
2023
-
[50]
Mutual information guided backdoor mitigation for pre- trained encoders,
T. Han, W. Sun, Z. Ding, C. Fang, H. Qian, J. Li, Z. Chen, and X. Zhang, “Mutual information guided backdoor mitigation for pre- trained encoders,” arXiv preprint arXiv:2406.03508, 2024
2024
-
[51]
Mitigat- ing Backdoor Attacks in Pre-Trained Encoders via Self-Supervised Knowledge Distillation ,
R. Bie, J. Jiang, H. Xie, Y . Guo, Y . Miao, and X. Jia, “ Mitigat- ing Backdoor Attacks in Pre-Trained Encoders via Self-Supervised Knowledge Distillation ,” IEEE Transactions on Services Computing, vol. 17, pp. 2613–2625, Sept. 2024
2024
-
[52]
A density-based algorithm for discovering clusters in large spatial databases with noise,
M. Ester, H.-P. Kriegel, J. Sander, X. Xu, et al., “A density-based algorithm for discovering clusters in large spatial databases with noise,” in KDD, pp. 226–231, 1996. Appendix A. Implementation Details of T-Core In Sec. 4, we present the overall algorithmic details of our T...
1996
-
[53]
https://github.com/leftthomas/SimCLR
-
[54]
https://github.com/UMBCvision/SSL-Backdoor
-
[55]
https://github.com/jinyuan-jia/BadEncoder
-
[56]
https://github.com/Gwinhen/DRUPE
-
[57]
https://github.com/jzhang538/CorruptEncoder
-
[58]
Car” for STL-10, “Bird
https://github.com/vtu81/backdoor-toolbox experimental setups for certain attack methods on some datasets. For instance, TaCT attack on ImageNet is originally not supported. We conducted experiments by adapting the settings from other datasets. Trigger and Target Class: In the...
-
[59]
https://github.com/Unispac/Fight-Poison-With-Poison
-
[60]
https://github.com/THUYimingLi/BackdoorBox
-
[61]
https://github.com/reds-lab/ASSET
-
[62]
https://github.com/bboylyg/ABL
-
[63]
https://github.com/zaixizhang/CBD
-
[64]
https://github.com/rkteddy/channel-Lipschitzness-based-pruning
-
[65]
https://github.com/wssun/SSLBackdoorMitigation
-
[66]
https://github.com/ruoxi-jia-group/Meta-Sift
-
[67]
https://github.com/SCLBD/BackdoorBench
-
[68]
https://github.com/YiZeng623/I-BAU TABLE 19: Performance of the state-of-the-art inference-time defenses SCALE-UP and IBD-PSC under different threat scenarios, using the CIFAR-10 dataset. SCALE-UP IBD-PSCThreat Type Poisoning Method TPR(%) ↑ FPR(%) ↓ AUC↑ F1↑ TPR(%) ↑ FPR(%) ↓...
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.