REVIEW 3 major objections 7 minor 67 references
A single poison trigger can be planted in training so that an entire family of unseen inference-time triggers — none of which appear in the victim's training data — reliably fires the backdoor, while clean accuracy barely drops.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 02:44 UTC pith:QWFOSFIX
load-bearing objection A genuinely new backdoor setting with broad evidence, but the theory is a scaffold: the key transfer assumption is never measured, and the screening condition in the implementation is weaker than the theorem requires. the 3 major comments →
Lilith: Backdoor Generalization under Training-Inference Trigger Shift
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms: a backdoor can be made to generalize from a single training-time trigger to an inference-only family of unseen triggers, as long as the family lives inside the representation basin created by the anchor. The authors model trigger variation as a shift in attack support and characterize the mechanism geometrically: a training anchor must establish a compact target-oriented region with positive decision-boundary clearance (anchor clearance), and valid inference variants must keep their representation reach below that clearance (family reach). Under bounded surrogate–victim geometric discrepancy, surrogate screening of proposed variants suffices to guarantee victim-side
What carries the argument
Two coupled modules. Anchor Vulnerability Induction maximizes a geometric objective — target margin over twice the Lipschitz constant minus the triggered representation's concentration radius — to create a compact anchor basin on the target side of the surrogate decision boundary. Representation-Aligned Family Construction proposes candidates in a bounded subspace (orthonormal projection around the anchor), then retains only variants whose representation reach stays inside the surrogate anchor margin. The analytical backbone is a set of local Lipschitz and regularity assumptions leading to Proposition 4.2 (parameter-to-representation stability, bounding representation deviation by C_m·η), Th
Load-bearing premise
The pivotal premise is Assumption 4 (Bounded Geometry Transfer): for every retained family member, the victim's normalized geometric risk ζ_v is at most the surrogate's ζ_s plus a constant κ_tr — that is, the representation geometry the attacker measures on the surrogate transfers to the unknown victim within a bounded gap, and this gap is never measured in the paper.
What would settle it
Train a victim model with the same anchor trigger on a held-out dataset, then for each surrogate-screened family member compute ζ_s on the surrogate and ζ_v on the victim (using the definitions in Eqs. (1)–(2)); if even a few percent of retained variants satisfy ζ_s < 1 but ζ_v ≥ 1, the bounded-transfer assumption fails and family-wise ASR should drop accordingly — a direct quantitative check that the paper does not perform.
If this is right
- Backdoor evaluation should measure family-wise attack success and the trigger generalization gap, not just exact-trigger or training-exposed-trigger success, because exact-trigger protocols miss behavior that generalizes across trigger supports.
- A single poisoned training example (or a small poisoning budget, e.g. 0.5% of the training set) can yield family-wise attack success above 90% across diverse datasets and victim architectures, implying a broader post-deployment attack surface than previously assumed.
- The representation alignment result implies that defenses must address the geometry of the anchor-induced region rather than any one trigger pattern: input-level purification leaves high residual family activation, while representation-shaping defenses (e.g. feature shift tuning) suppress it more strongly.
- The trigger generalization gap provides a quantitative yardstick for comparing backdoor attacks and defenses under shift, and the paper's sufficient conditions give concrete design criteria for when an inference-only family will preserve a training-time backdoor.
- The finding that off-subspace variants fail while aligned in-subspace variants succeed suggests that trigger families are only dangerous when they share the representation basin, which could inform both attack construction and defense screening.
Where Pith is reading between the lines
- A direct testable extension is to measure the surrogate and victim geometric risks ζ_s and ζ_v empirically for each retained variant across multiple victim training runs; if the inequality ζ_v ≤ ζ_s + κ_tr fails for a nontrivial fraction of variants, family-wise attack success should collapse, providing a way to falsify the bounded-transfer assumption before deployment.
- The framework suggests a new defense objective: instead of detecting or removing a specific trigger, flatten the target-side representation region so that no bounded family can preserve clearance — for example, by regularizing the margin around triggered representations during training.
- If the geometric-transfer assumption holds for clean-label poisoning as well (the paper notes the analysis applies to clean-label realizations satisfying the anchor condition), the same anchor-to-family effect could be achievable without label flipping, making the attack harder to detect by label-based filtering.
- The trigger generalization gap could be adopted as a standard evaluation axis for backdoor robustness, analogous to how accuracy under distribution shift is reported for benign models; this would shift the field's focus from memorization of artifacts to generalization of malicious behavior.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a new backdoor attack setting: training–inference trigger shift, where a single anchor trigger is used to poison a victim model during training, and the attacker later uses a family of unseen triggers at inference. The proposed framework, Lilith, uses a surrogate model to (i) induce a compact target-side region with the anchor and (ii) construct an inference-time family of trigger variants that preserve the anchor's representation geometry. The authors prove sufficient conditions for family-wise activation based on anchor clearance, family reach, and a bounded geometry-transfer assumption. They evaluate the attack across five datasets, five architectures, three poisoning rates, multiple proposal mechanisms, and several defenses, reporting high family ASR and low trigger generalization gap.
Significance. If the empirical results are reproducible, the paper identifies a practically important blind spot: exact-trigger evaluations may miss a large class of inference-time triggers that activate a backdoor trained with one anchor. The formal results are logically sound but rely on unmeasured quantities, in particular Assumption 4. The breadth of the empirical study (75 dataset–architecture–poisoning configurations, plus defenses) is a strength. However, the paper's theoretical characterization is not yet connected to the algorithm actually run; the experiments demonstrate the phenomenon but do not verify the sufficient conditions.
major comments (3)
- [§4.3, Eq. (6) vs. Corollary 4.6, Eq. (13)] The screening constraint in Eq. (6) imposes R_s(Q,δ_F) < γ_{s,A}/(2L_s), i.e., ess sup ζ_s < 1. Corollary 4.6 requires ess sup ζ_s < 1−κ_tr. Since κ_tr>0, Eq. (6) permits variants with ζ_s ∈ [1−κ_tr, 1), for which the corollary's guarantee is void. No estimate of κ_tr or of the maximal ζ_s among retained variants is reported in §5. The sentence after Corollary 4.6 claiming that the premise is 'evaluated through model, dataset, and proposal-mechanism mismatch' does not substitute for a measurement. This is load-bearing because the screening actually implemented is not the screening the theorem requires.
- [§3, Assumption 4] Assumption 4 (Bounded Geometry Transfer) asserts ζ_v(θ,δ) ≤ ζ_s(θ,δ)+κ_tr for each retained family member. This is the key link between surrogate screening and victim activation in Corollary 4.6. However, the paper provides no measurement of ζ_s, ζ_v, or κ_tr for any experiment. The assumption is thus unverified; if it fails, family-wise ASR can collapse. The statement that the premise is 'evaluated through model, dataset, and proposal-mechanism mismatch' is not operationalized. The theory's central bridge is therefore an unverified assumption rather than an independently grounded prediction.
- [§4.2, §4.3 (Eqs. 2, 5, 12)] The sufficient conditions of Proposition 4.1 (Eq. 5), Corollary 4.4 (Eq. 12), and the definition of ζ_m in Eq. (2) involve L_v, L_{φ,v}, r_v(θ_A,δ_A), γ_{v,A}, and C_v, none of which are measured on the victim models used in §5. The text in §4.2 says L_s is used 'only for geometric interpretation' and no global Lipschitz constant is estimated. As a result, the paper does not instantiate its own sufficient conditions. The empirical attack may still work, but the claimed theoretical 'characterization' is not validated by the experiments.
minor comments (7)
- [§5.1 / Figure 3] Specify whether, in the matched trigger-shift comparison, all baselines are evaluated on the same family distribution or on each baseline's own natural family under the same budget. The phrase 'under the same transformation set and perceptual budget' is ambiguous.
- [§4.3, Eq. (6)] The objective V(Q) is left unspecified. Define it (e.g., entropy, pairwise distance) or state informally what diversity measure is used; as written, Eq. (6) is not fully specified.
- [§3, Eq. (2)] Clarify the role of δ in ζ_m(θ,δ) and its relation to δ_A and δ_F used later. The notation suggests a per-member failure probability, but the connection is not explicit.
- [Table 3] The 'Train-visible' column mixes module names and boolean values; align the table so each row has one value (e.g., Yes/No/Module I).
- [§5.5, Figure 5(c)] The PCA projection of spectral statistics is qualitative; consider reporting AUROC or a quantitative separation measure for the spectral detector.
- [General] The paper uses the 'Conference’17' template with a 2026 copyright. Update the template and venue information before submission. Several references are arXiv-only; add peer-reviewed versions where available.
- [§4.2] The phrase 'The implementation uses L_s only for geometric interpretation and does not estimate a global Lipschitz constant' is confusing given that Eq. (6) depends on L_s. Explain how the screening constraint is implemented in practice.
Circularity Check
No circularity: the derivation is a set of valid conditionals; the unmeasured transfer assumption and screening-condition gap are premise-verification issues, not reductions.
full rationale
Lilith's derivation chain is conditional and self-contained in the formal sense. Proposition 4.1, Proposition 4.2, Theorem 4.3, and Corollaries 4.4-4.6 are each proved from the definitions in Section 3 and the explicitly stated assumptions (local score regularity, learned anchor basin, regular trigger manifold, bounded geometry transfer); no conclusion is used to define its own premise. Assumption 4 (zeta_v <= zeta_s + kappa_tr) is the central transfer premise, but it is not measured or derived, and the text after Corollary 4.6 only says the premise is 'evaluated through model, dataset, and proposal-mechanism mismatch' without reporting zeta_s, zeta_v, or kappa_tr. Moreover, the implemented screening in Eq. (6) enforces R_s < gamma_s,A/(2L_s), i.e. zeta_s < 1, not the stricter zeta_s < 1 - kappa_tr required by Eq. (13). These are genuine soundness/premise-validation gaps, but they are not circular reductions: the victim-side success condition is not the same equation as the assumption, and the empirical RQ1-RQ5 measurements of ASR_F and TGG provide independent evidence. The self-citations present (e.g., [3], [37]) are contextual and not load-bearing. I therefore find no significant circularity.
Axiom & Free-Parameter Ledger
free parameters (3)
- Family subspace radius η
- Anchor optimization loss weights
- Transfer constant κ_tr
axioms (6)
- domain assumption Assumption 1 (Local Score Regularity): each score s_{m,j} is L_m-Lipschitz on the needed neighborhoods.
- domain assumption Assumption 2 (Learned Anchor Basin): the victim anchor has finite concentration radius r_v(θ_A,δ_A) and positive margin γ_v,A.
- domain assumption Assumption 3 (Regular Trigger Manifold): pattern map g is continuously differentiable with bounded Jacobian, B is Lipschitz, and φ_m is locally Lipschitz.
- domain assumption Assumption 4 (Bounded Geometry Transfer): ζ_v(θ,δ) ≤ ζ_s(θ,δ)+κ_tr for each retained family member.
- domain assumption Poison-only black-box victim trains on the poisoned dataset with an unknown algorithm A_v.
- standard math Standard analysis tools: mean-value theorem, Jensen's inequality, triangle inequality, Lipschitz composition.
read the original abstract
Machine-learning services increasingly rely on public data, third-party providers, and outsourced training, creating opportunities for data-poisoning attacks that implant persistent malicious behavior while preserving benign utility. However, existing backdoor studies largely evaluate exact trigger reuse, training-exposed trigger diversity, or variations along predefined transformation axes. They therefore leave a critical blind spot: whether a backdoor learned from one training-time trigger can generalize to an inference-time trigger family absent from victim training. We formulate this problem as backdoor generalization under training--inference trigger shift and introduce Lilith, a black-box anchor-to-family framework. Using only disjoint surrogate resources, Lilith first induces a compact target-side vulnerability with a single training anchor, then constructs a bounded inference-only family that preserves the anchor-induced representation geometry. We characterize this mechanism through anchor clearance and family reach, deriving sufficient conditions for family-wise target preservation under local regularity and bounded surrogate--victim discrepancy. Experiments across datasets, architectures, poisoning rates, and defenses show that Lilith achieves high family-wise attack success with limited utility degradation and a small trigger generalization gap. Additional analyses show that family activation depends on representation alignment rather than the proposal mechanism, exposing a broader threat overlooked by exact-trigger evaluation.
Figures
Reference graph
Works this paper leans on
-
[1]
Amazon Web Services. 2026. Amazon SageMaker Model Monitor and Clarify
2026
-
[2]
Lucas Beyer, Xiaohua Zhai, and Alexander Kolesnikov. 2022. Better Plain ViT Baselines for ImageNet-1k.arXiv preprint arXiv:2205.01580abs/2205.01580 (2022). arXiv:2205.01580
Pith/arXiv arXiv 2022
-
[3]
Jiahao Chen, Zhiqiang Shen, Yuwen Pu, Chunyi Zhou, Changjiang Li, Jiliang Li, Ting Wang, and Shouling Ji. 2024. Rethinking the Vulnerabilities of Face Recog- nition Systems: From a Practical Perspective.arXiv preprint arXiv:2405.12786 abs/2405.12786 (2024). arXiv:2405.12786
Pith/arXiv arXiv 2024
-
[4]
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song. 2017. Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning.arXiv preprint arXiv:1712.05526(2017)
Pith/arXiv arXiv 2017
-
[5]
Ambra Demontis, Marco Melis, Maura Pintor, Matthew Jagielski, Battista Biggio, Alina Oprea, Cristina Nita-Rotaru, and Fabio Roli. 2019. Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning Attacks. InUSENIX Sec. Symp.321–338
2019
-
[6]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Im- ageNet: A large-scale hierarchical image database. InIEEE Conf. Comput. Vis. Pattern Recog.248–255
2009
-
[7]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. InInt. Conf. Learn. Represent.OpenReview.net
2021
-
[8]
Ranasinghe, and Surya Nepal
Yansong Gao, Change Xu, Derui Wang, Shiping Chen, Damith C. Ranasinghe, and Surya Nepal. 2019. STRIP: A Defence Against Trojan Attacks on Deep Neural Networks. InAnnu. Comput. Secur. Appl. Conf. (ACSAC). 113–125
2019
-
[9]
Huaizhi Ge, Yiming Li, Qifan Wang, Yongfeng Zhang, and Ruixiang Tang. 2025. When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model- Generated Explanations. InAnnu. Meet. Assoc. Comput. Linguist.2278–2296
2025
-
[10]
Goodfellow, Yaroslav Bulatov, Julian Ibarz, Sacha Arnoud, and Vinay D
Ian J. Goodfellow, Yaroslav Bulatov, Julian Ibarz, Sacha Arnoud, and Vinay D. Shet. 2014. Multi-digit Number Recognition from Street View Imagery using Deep Convolutional Neural Networks. InInt. Conf. Learn. Represent
2014
-
[11]
Google Cloud. 2024. Google Vertex AI Model Monitoring
2024
-
[12]
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg. 2017. BadNets: Identify- ing Vulnerabilities in the Machine Learning Model Supply Chain.arXiv preprint arXiv:1708.06733(2017)
Pith/arXiv arXiv 2017
-
[13]
Frank, and Kilian Q
Chuan Guo, Jared S. Frank, and Kilian Q. Weinberger. 2019. Low Frequency Adversarial Perturbation. InConf. Uncertain. Artif. Intell.1127–1137
2019
-
[14]
Junfeng Guo, Yiming Li, Xun Chen, Hanqing Guo, Lichao Sun, and Cong Liu
-
[15]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. InIEEE Conf. Comput. Vis. Pattern Recog.IEEE, 770–778
2016
-
[16]
Matthias Hein and Maksym Andriushchenko. 2017. Formal Guarantees on the Robustness of a Classifier against Adversarial Manipulation. InAdv. Neural Inform. Process. Syst.2266–2276
2017
-
[17]
Linshan Hou, Ruili Feng, Zhongyun Hua, Wei Luo, Leo Yu Zhang, and Yiming Li
-
[18]
Kunzhe Huang, Yiming Li, Baoyuan Wu, Zhan Qin, and Kui Ren. 2022. Backdoor Defense via Decoupling the Training Process. InInt. Conf. Learn. Represent
2022
-
[19]
2009.Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky and Geoffrey Hinton. 2009.Learning Multiple Layers of Features from Tiny Images. Technical Report. University of Toronto. Technical Report
2009
-
[20]
Yann Le and Xuan Yang. 2015. Tiny ImageNet Visual Recognition Challenge. CS231n Course Project, Stanford University. Accessed: 2025-09-16
2015
-
[21]
Yiming Li, Yong Jiang, Zhifeng Li, and Shu-Tao Xia. 2024. Backdoor Learning: A Survey.IEEE Trans. Neural Netw. Learn. Syst.35, 1 (2024), 5–22
2024
-
[22]
Yuezun Li, Yiming Li, Baoyuan Wu, Longkang Li, Ran He, and Siwei Lyu. 2021. Invisible Backdoor Attack with Sample-Specific Triggers. InInt. Conf. Comput. Vis.16443–16452
2021
-
[23]
Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma. 2021. Anti-Backdoor Learning: Training Clean Models on Poisoned Data. InAdv. Neural Inform. Process. Syst.14900–14912
2021
-
[24]
Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma. 2021. Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks. InInt. Conf. Learn. Represent
2021
-
[25]
Yiming Li, Tongqing Zhai, Yong Jiang, Zhifeng Li, and Shu-Tao Xia. 2021. Back- door Attack in the Physical World.arXiv preprint arXiv:2104.02361(2021)
Pith/arXiv arXiv 2021
-
[26]
Chenhao Lin, Chenyang Zhao, Shiwei Wang, Longtian Wang, Chao Shen, and Zhengyu Zhao. 2025. Revisiting Training-Inference Trigger Intensity in Backdoor Attacks. InUSENIX Sec. Symp.6359–6378
2025
-
[27]
Xiaogeng Liu, Minghui Li, Haoyu Wang, Shengshan Hu, Dengpan Ye, Hai Jin, Libing Wu, and Chaowei Xiao. 2023. Detecting Backdoors During the Inference Stage Based on Corruption Robustness Consistency. InIEEE Conf. Comput. Vis. Pattern Recog.16363–16372
2023
-
[28]
Yunfei Liu, Xingjun Ma, James Bailey, and Feng Lu. 2020. Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks. InEur. Conf. Comput. Vis. 182–199
2020
-
[29]
Wanlun Ma, Derui Wang, Ruoxi Sun, Minhui Xue, Sheng Wen, and Yang Xiang
-
[30]
Microsoft. 2026. Azure Machine Learning Model Monitoring
2026
-
[31]
Rui Min, Zeyu Qin, Li Shen, and Minhao Cheng. 2023. Towards Stable Backdoor Purification through Feature Shift Tuning. InAdv. Neural Inform. Process. Syst. Conference’17, July 2017, Washington, DC, USA Zhou Feng et al
2023
-
[32]
The "Beatrix" Resurrections: Robust Backdoor Detection via Gram Matrices. InNetw. Distrib. Syst. Secur. Symp. (NDSS)
-
[33]
Ali Bou Nassif, Manar Abu Talib, Qassim Nasir, and Fatima Mohamad Dakalbab
-
[34]
Anh Nguyen and Anh Tuan Tran. 2020. Input-Aware Dynamic Backdoor Attack. InAdv. Neural Inform. Process. Syst
2020
-
[35]
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, and Pascal Frossard. 2017. Universal Adversarial Perturbations. InIEEE Conf. Comput. Vis. Pattern Recog.86–94
2017
-
[36]
Soumyadeep Pal, Yuguang Yao, Ren Wang, Bingquan Shen, and Sijia Liu. 2024. Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction Consistency. InInt. Conf. Learn. Represent
2024
-
[37]
Yuwen Pu, Jiahao Chen, Chunyi Zhou, Zhou Feng, Qingming Li, Chunqiang Hu, and Shouling Ji. 2024. Mellivora Capensis: A Backdoor-Free Training Framework on the Poisoned Dataset without Auxiliary Data.arXiv preprint arXiv:2405.12719 (2024)
arXiv 2024
-
[38]
Xiangyu Qi, Tinghao Xie, Yiming Li, Saeed Mahloujifar, and Prateek Mittal. 2023. Revisiting the Assumption of Latent Separability for Backdoor Defenses. InInt. Conf. Learn. Represent
2023
-
[39]
Tuan Anh Nguyen and Anh Tuan Tran. 2021. WaNet – Imperceptible Warping- based Backdoor Attack. InInt. Conf. Learn. Represent
2021
-
[40]
Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedan- tam, Devi Parikh, and Dhruv Batra
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedan- tam, Devi Parikh, and Dhruv Batra. 2017. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. InInt. Conf. Comput. Vis.618– 626
2017
-
[41]
Yucheng Shi, Mengnan Du, Xuansheng Wu, Zihan Guan, Jin Sun, and Ninghao Liu. 2023. Black-box Backdoor Defense via Zero-shot Image Purification. InAdv. Neural Inform. Process. Syst
2023
-
[42]
Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. InInt. Conf. Learn. Represent. arXiv:1409.1556
Pith/arXiv arXiv 2015
-
[43]
Wang, Tong Wu, Saeed Mahloujifar, and Prateek Mittal
Xiangyu Qi, Tinghao Xie, Jiachen T. Wang, Tong Wu, Saeed Mahloujifar, and Prateek Mittal. 2023. Towards a Proactive ML Approach for Detecting Backdoor Poison Samples. InUSENIX Sec. Symp.1685–1702
2023
-
[44]
Hossein Souri, Liam Fowl, Rama Chellappa, Micah Goldblum, and Tom Goldstein
-
[45]
Hoover, Aaron Isaksen, Andy Nealen, and Julian Togelius
Adam Summerville, Sam Snodgrass, Matthew Guzdial, Christoffer Holmgaård, Amy K. Hoover, Aaron Isaksen, Andy Nealen, and Julian Togelius. 2018. Proce- dural Content Generation via Machine Learning (PCGML).IEEE Trans. Games 10, 3 (2018), 257–270
2018
-
[46]
Brandon Tran, Jerry Li, and Aleksander Madry. 2018. Spectral Signatures in Backdoor Attacks. InAdv. Neural Inform. Process. Syst.8011–8021
2018
-
[47]
Jure Sokolic, Raja Giryes, Guillermo Sapiro, and Miguel R. D. Rodrigues. 2017. Robust Large Margin Deep Neural Networks.IEEE Trans. Signal Process.65, 16 (2017), 4265–4280
2017
-
[48]
Alexander Turner, Dimitris Tsipras, and Aleksander Madry. 2019. Label- Consistent Backdoor Attacks.arXiv preprint arXiv:1912.02771(2019)
Pith/arXiv arXiv 2019
-
[49]
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y. Zhao. 2019. Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural Networks. InIEEE Symp. Secur. Privacy (S&P). 707– 723
2019
-
[50]
Tong Wang, Yuan Yao, Feng Xu, Shengwei An, Hanghang Tong, and Ting Wang
-
[51]
Jun Xia, Zhihao Yue, Yingbo Zhou, Zhiwei Ling, Yiyu Shi, Xian Wei, and Mingsong Chen. 2024. WaveAttack: Asymmetric Frequency Obfuscation-based Backdoor Attacks Against Deep Neural Networks. InAdv. Neural Inform. Process. Syst
2024
-
[52]
Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama. 2018. Lipschitz-Margin Train- ing: Scalable Certification of Perturbation Invariance for Deep Neural Networks. InAdv. Neural Inform. Process. Syst.6542–6551
2018
-
[53]
Yanxin Yang, Chentao Jia, Dengke Yan, Ming Hu, Tianlin Li, Xiaofei Xie, Xian Wei, and Mingsong Chen. 2024. SampDetox: Black-Box Backdoor Defense via Perturbation-Based Sample Detoxification. InAdv. Neural Inform. Process. Syst
2024
-
[54]
Yuanshun Yao, Huiying Li, Haitao Zheng, and Ben Y. Zhao. 2019. Latent Backdoor Attacks on Deep Neural Networks. InACM SIGSAC Conf. Comput. Commun. Secur. 2041–2055
2019
-
[55]
Yi Zeng, Si Chen, Won Park, Zhuoqing Mao, Ming Jin, and Ruoxi Jia. 2022. Adversarial Unlearning of Backdoors via Implicit Hypergradient. InInt. Conf. Learn. Represent
2022
-
[56]
An Invisible Black-Box Backdoor Attack through Frequency Domain. In Eur. Conf. Comput. Vis.396–413
-
[57]
Efros, Eli Shechtman, and Oliver Wang
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang
-
[58]
Le Yang, Dongmei Jiang, Wenjing Han, and Hichem Sahli. 2017. DCNN and DNN Based Multi-Modal Depression Recognition. InInt. Conf. Affective Comput. Intell. Interact. (ACII). 484–489
2017
-
[59]
Mengxin Zheng, Jiaqi Xue, Zihao Wang, Xun Chen, Qian Lou, Lei Jiang, and Xiaofeng Wang. 2024. SSL-Cleanse: Trojan Detection and Mitigation in Self- Supervised Learning. InEur. Conf. Comput. Vis.405–421
2024
-
[60]
Mingli Zhu, Shaokui Wei, Li Shen, Yanbo Fan, and Baoyuan Wu. 2023. Enhancing Fine-Tuning Based Backdoor Defense with Sharpness-Aware Minimization. In Int. Conf. Comput. Vis.4443–4454. A Proofs We first establish a basic property of the target margin. For model 𝑚, define 𝑀𝑚,𝑡(𝑧)=𝑠 𝑚,𝑡(𝑧)−max 𝑗≠𝑡 𝑠𝑚,𝑗(𝑧). Proposition A.1 (Lipschitz Target Margin).If every s...
2023
-
[62]
Yi Zeng, Minzhou Pan, Hoang Anh Just, Lingjuan Lyu, Meikang Qiu, and Ruoxi Jia. 2023. Narcissus: A Practical Clean-Label Backdoor Attack with Limited Information. InACM Conf. Comput. Commun. Secur. (CCS). 771–785
2023
-
[65]
Shihao Zhao, Xingjun Ma, Xiang Zheng, James Bailey, Jingjing Chen, and Yu- Gang Jiang. 2020. Clean-Label Backdoor Attacks on Video Recognition Models. InIEEE Conf. Comput. Vis. Pattern Recog.14431–14440
2020
-
[2018]
InIEEE Conf
The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. InIEEE Conf. Comput. Vis. Pattern Recog.586–595
-
[2021]
Machine Learning for Anomaly Detection: A Systematic Review.IEEE Access9 (2021), 78658–78700
2021
-
[2022]
Sleeper Agent: Scalable Hidden Trigger Backdoors for Neural Networks Trained from Scratch. InAdv. Neural Inform. Process. Syst
-
[2023]
SCALE-UP: An Efficient Black-box Input-level Backdoor Detection via Analyzing Scaled Prediction Consistency. InInt. Conf. Learn. Represent
-
[2024]
IBD-PSC: Input-level Backdoor Detection via Parameter-oriented Scaling Consistency. InInt. Conf. Mach. Learn.764
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.