Pith. sign in

REVIEW 3 major objections 7 minor 67 references

A single poison trigger can be planted in training so that an entire family of unseen inference-time triggers — none of which appear in the victim's training data — reliably fires the backdoor, while clean accuracy barely drops.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 02:44 UTC pith:QWFOSFIX

load-bearing objection A genuinely new backdoor setting with broad evidence, but the theory is a scaffold: the key transfer assumption is never measured, and the screening condition in the implementation is weaker than the theorem requires. the 3 major comments →

arxiv 2607.26099 v1 pith:QWFOSFIX submitted 2026-07-28 cs.CR cs.LG

Lilith: Backdoor Generalization under Training-Inference Trigger Shift

classification cs.CR cs.LG
keywords backdoor attackdata poisoningtraining–inference trigger shifttrigger generalizationrepresentation geometryblack-box attackanchor-to-familymachine learning security
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper introduces 'backdoor generalization under training–inference trigger shift': instead of asking whether a backdoor tolerates small edits to a known trigger, it asks whether malicious behavior learned from one training-time anchor can be activated by a family of inference-time triggers never shown during victim training. The authors propose Lilith, a black-box framework that uses only a disjoint surrogate dataset and model to (1) induce a compact target-side vulnerability with one anchor trigger, and (2) construct a bounded inference-only family that preserves the anchor's representation geometry. They derive sufficient conditions — based on anchor clearance and family reach — under which unseen family members inherit the malicious decision, and they report experiments across five datasets, five architectures, three poisoning rates, multiple proposal mechanisms, and several defenses showing high family-wise attack success with a small trigger generalization gap. The central claim is that representation alignment, not trigger similarity or the proposal mechanism, governs whether an unseen trigger activates the backdoor.

Core claim

On the paper's own terms: a backdoor can be made to generalize from a single training-time trigger to an inference-only family of unseen triggers, as long as the family lives inside the representation basin created by the anchor. The authors model trigger variation as a shift in attack support and characterize the mechanism geometrically: a training anchor must establish a compact target-oriented region with positive decision-boundary clearance (anchor clearance), and valid inference variants must keep their representation reach below that clearance (family reach). Under bounded surrogate–victim geometric discrepancy, surrogate screening of proposed variants suffices to guarantee victim-side

What carries the argument

Two coupled modules. Anchor Vulnerability Induction maximizes a geometric objective — target margin over twice the Lipschitz constant minus the triggered representation's concentration radius — to create a compact anchor basin on the target side of the surrogate decision boundary. Representation-Aligned Family Construction proposes candidates in a bounded subspace (orthonormal projection around the anchor), then retains only variants whose representation reach stays inside the surrogate anchor margin. The analytical backbone is a set of local Lipschitz and regularity assumptions leading to Proposition 4.2 (parameter-to-representation stability, bounding representation deviation by C_m·η), Th

Load-bearing premise

The pivotal premise is Assumption 4 (Bounded Geometry Transfer): for every retained family member, the victim's normalized geometric risk ζ_v is at most the surrogate's ζ_s plus a constant κ_tr — that is, the representation geometry the attacker measures on the surrogate transfers to the unknown victim within a bounded gap, and this gap is never measured in the paper.

What would settle it

Train a victim model with the same anchor trigger on a held-out dataset, then for each surrogate-screened family member compute ζ_s on the surrogate and ζ_v on the victim (using the definitions in Eqs. (1)–(2)); if even a few percent of retained variants satisfy ζ_s < 1 but ζ_v ≥ 1, the bounded-transfer assumption fails and family-wise ASR should drop accordingly — a direct quantitative check that the paper does not perform.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Backdoor evaluation should measure family-wise attack success and the trigger generalization gap, not just exact-trigger or training-exposed-trigger success, because exact-trigger protocols miss behavior that generalizes across trigger supports.
  • A single poisoned training example (or a small poisoning budget, e.g. 0.5% of the training set) can yield family-wise attack success above 90% across diverse datasets and victim architectures, implying a broader post-deployment attack surface than previously assumed.
  • The representation alignment result implies that defenses must address the geometry of the anchor-induced region rather than any one trigger pattern: input-level purification leaves high residual family activation, while representation-shaping defenses (e.g. feature shift tuning) suppress it more strongly.
  • The trigger generalization gap provides a quantitative yardstick for comparing backdoor attacks and defenses under shift, and the paper's sufficient conditions give concrete design criteria for when an inference-only family will preserve a training-time backdoor.
  • The finding that off-subspace variants fail while aligned in-subspace variants succeed suggests that trigger families are only dangerous when they share the representation basin, which could inform both attack construction and defense screening.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A direct testable extension is to measure the surrogate and victim geometric risks ζ_s and ζ_v empirically for each retained variant across multiple victim training runs; if the inequality ζ_v ≤ ζ_s + κ_tr fails for a nontrivial fraction of variants, family-wise attack success should collapse, providing a way to falsify the bounded-transfer assumption before deployment.
  • The framework suggests a new defense objective: instead of detecting or removing a specific trigger, flatten the target-side representation region so that no bounded family can preserve clearance — for example, by regularizing the margin around triggered representations during training.
  • If the geometric-transfer assumption holds for clean-label poisoning as well (the paper notes the analysis applies to clean-label realizations satisfying the anchor condition), the same anchor-to-family effect could be achievable without label flipping, making the attack harder to detect by label-based filtering.
  • The trigger generalization gap could be adopted as a standard evaluation axis for backdoor robustness, analogous to how accuracy under distribution shift is reported for benign models; this would shift the field's focus from memorization of artifacts to generalization of malicious behavior.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper introduces a new backdoor attack setting: training–inference trigger shift, where a single anchor trigger is used to poison a victim model during training, and the attacker later uses a family of unseen triggers at inference. The proposed framework, Lilith, uses a surrogate model to (i) induce a compact target-side region with the anchor and (ii) construct an inference-time family of trigger variants that preserve the anchor's representation geometry. The authors prove sufficient conditions for family-wise activation based on anchor clearance, family reach, and a bounded geometry-transfer assumption. They evaluate the attack across five datasets, five architectures, three poisoning rates, multiple proposal mechanisms, and several defenses, reporting high family ASR and low trigger generalization gap.

Significance. If the empirical results are reproducible, the paper identifies a practically important blind spot: exact-trigger evaluations may miss a large class of inference-time triggers that activate a backdoor trained with one anchor. The formal results are logically sound but rely on unmeasured quantities, in particular Assumption 4. The breadth of the empirical study (75 dataset–architecture–poisoning configurations, plus defenses) is a strength. However, the paper's theoretical characterization is not yet connected to the algorithm actually run; the experiments demonstrate the phenomenon but do not verify the sufficient conditions.

major comments (3)
  1. [§4.3, Eq. (6) vs. Corollary 4.6, Eq. (13)] The screening constraint in Eq. (6) imposes R_s(Q,δ_F) < γ_{s,A}/(2L_s), i.e., ess sup ζ_s < 1. Corollary 4.6 requires ess sup ζ_s < 1−κ_tr. Since κ_tr>0, Eq. (6) permits variants with ζ_s ∈ [1−κ_tr, 1), for which the corollary's guarantee is void. No estimate of κ_tr or of the maximal ζ_s among retained variants is reported in §5. The sentence after Corollary 4.6 claiming that the premise is 'evaluated through model, dataset, and proposal-mechanism mismatch' does not substitute for a measurement. This is load-bearing because the screening actually implemented is not the screening the theorem requires.
  2. [§3, Assumption 4] Assumption 4 (Bounded Geometry Transfer) asserts ζ_v(θ,δ) ≤ ζ_s(θ,δ)+κ_tr for each retained family member. This is the key link between surrogate screening and victim activation in Corollary 4.6. However, the paper provides no measurement of ζ_s, ζ_v, or κ_tr for any experiment. The assumption is thus unverified; if it fails, family-wise ASR can collapse. The statement that the premise is 'evaluated through model, dataset, and proposal-mechanism mismatch' is not operationalized. The theory's central bridge is therefore an unverified assumption rather than an independently grounded prediction.
  3. [§4.2, §4.3 (Eqs. 2, 5, 12)] The sufficient conditions of Proposition 4.1 (Eq. 5), Corollary 4.4 (Eq. 12), and the definition of ζ_m in Eq. (2) involve L_v, L_{φ,v}, r_v(θ_A,δ_A), γ_{v,A}, and C_v, none of which are measured on the victim models used in §5. The text in §4.2 says L_s is used 'only for geometric interpretation' and no global Lipschitz constant is estimated. As a result, the paper does not instantiate its own sufficient conditions. The empirical attack may still work, but the claimed theoretical 'characterization' is not validated by the experiments.
minor comments (7)
  1. [§5.1 / Figure 3] Specify whether, in the matched trigger-shift comparison, all baselines are evaluated on the same family distribution or on each baseline's own natural family under the same budget. The phrase 'under the same transformation set and perceptual budget' is ambiguous.
  2. [§4.3, Eq. (6)] The objective V(Q) is left unspecified. Define it (e.g., entropy, pairwise distance) or state informally what diversity measure is used; as written, Eq. (6) is not fully specified.
  3. [§3, Eq. (2)] Clarify the role of δ in ζ_m(θ,δ) and its relation to δ_A and δ_F used later. The notation suggests a per-member failure probability, but the connection is not explicit.
  4. [Table 3] The 'Train-visible' column mixes module names and boolean values; align the table so each row has one value (e.g., Yes/No/Module I).
  5. [§5.5, Figure 5(c)] The PCA projection of spectral statistics is qualitative; consider reporting AUROC or a quantitative separation measure for the spectral detector.
  6. [General] The paper uses the 'Conference’17' template with a 2026 copyright. Update the template and venue information before submission. Several references are arXiv-only; add peer-reviewed versions where available.
  7. [§4.2] The phrase 'The implementation uses L_s only for geometric interpretation and does not estimate a global Lipschitz constant' is confusing given that Eq. (6) depends on L_s. Explain how the screening constraint is implemented in practice.

Circularity Check

0 steps flagged

No circularity: the derivation is a set of valid conditionals; the unmeasured transfer assumption and screening-condition gap are premise-verification issues, not reductions.

full rationale

Lilith's derivation chain is conditional and self-contained in the formal sense. Proposition 4.1, Proposition 4.2, Theorem 4.3, and Corollaries 4.4-4.6 are each proved from the definitions in Section 3 and the explicitly stated assumptions (local score regularity, learned anchor basin, regular trigger manifold, bounded geometry transfer); no conclusion is used to define its own premise. Assumption 4 (zeta_v <= zeta_s + kappa_tr) is the central transfer premise, but it is not measured or derived, and the text after Corollary 4.6 only says the premise is 'evaluated through model, dataset, and proposal-mechanism mismatch' without reporting zeta_s, zeta_v, or kappa_tr. Moreover, the implemented screening in Eq. (6) enforces R_s < gamma_s,A/(2L_s), i.e. zeta_s < 1, not the stricter zeta_s < 1 - kappa_tr required by Eq. (13). These are genuine soundness/premise-validation gaps, but they are not circular reductions: the victim-side success condition is not the same equation as the assumption, and the empirical RQ1-RQ5 measurements of ASR_F and TGG provide independent evidence. The self-citations present (e.g., [3], [37]) are contextual and not load-bearing. I therefore find no significant circularity.

Axiom & Free-Parameter Ledger

3 free parameters · 6 axioms · 0 invented entities

The central claim rests on a series of plausible but unmeasured domain assumptions. In particular, Assumption 4 is a restatement of the transfer property the attack requires. No free-parameter values are reported for the trigger optimization or family screening.

free parameters (3)
  • Family subspace radius η
    Bounds candidate parameters in Module II; controls family reach vs. diversity. No concrete value or selection rule is reported.
  • Anchor optimization loss weights
    Module I combines differentiable estimators for target consistency, contraction, and perceptual/frequency budgets; exact weights and budgets are not specified.
  • Transfer constant κ_tr
    Assumption 4's discrepancy bound; never estimated, but used in Corollary 4.6 as the critical threshold 1−κ_tr.
axioms (6)
  • domain assumption Assumption 1 (Local Score Regularity): each score s_{m,j} is L_m-Lipschitz on the needed neighborhoods.
    Standard margin-based assumption, not verified for the trained victim networks.
  • domain assumption Assumption 2 (Learned Anchor Basin): the victim anchor has finite concentration radius r_v(θ_A,δ_A) and positive margin γ_v,A.
    This is the key outcome the poisoning is supposed to produce; not a universal property and not proven for an unknown training algorithm.
  • domain assumption Assumption 3 (Regular Trigger Manifold): pattern map g is continuously differentiable with bounded Jacobian, B is Lipschitz, and φ_m is locally Lipschitz.
    Restricts trigger parameterizations to smooth families; plausible for patch/blending/warp/frequency triggers but not all designs.
  • domain assumption Assumption 4 (Bounded Geometry Transfer): ζ_v(θ,δ) ≤ ζ_s(θ,δ)+κ_tr for each retained family member.
    Assumes surrogate screening transfers to the victim; essentially the desired transfer property, not independently measured.
  • domain assumption Poison-only black-box victim trains on the poisoned dataset with an unknown algorithm A_v.
    Threat-model assumption defining the attacker's limited control.
  • standard math Standard analysis tools: mean-value theorem, Jensen's inequality, triangle inequality, Lipschitz composition.
    Used in the proofs of Propositions 4.1–4.2 and Theorem 4.3.

pith-pipeline@v1.3.0-alltime-deepseek · 18841 in / 12057 out tokens · 121099 ms · 2026-08-01T02:44:13.281647+00:00 · methodology

0 comments
read the original abstract

Machine-learning services increasingly rely on public data, third-party providers, and outsourced training, creating opportunities for data-poisoning attacks that implant persistent malicious behavior while preserving benign utility. However, existing backdoor studies largely evaluate exact trigger reuse, training-exposed trigger diversity, or variations along predefined transformation axes. They therefore leave a critical blind spot: whether a backdoor learned from one training-time trigger can generalize to an inference-time trigger family absent from victim training. We formulate this problem as backdoor generalization under training--inference trigger shift and introduce Lilith, a black-box anchor-to-family framework. Using only disjoint surrogate resources, Lilith first induces a compact target-side vulnerability with a single training anchor, then constructs a bounded inference-only family that preserves the anchor-induced representation geometry. We characterize this mechanism through anchor clearance and family reach, deriving sufficient conditions for family-wise target preservation under local regularity and bounded surrogate--victim discrepancy. Experiments across datasets, architectures, poisoning rates, and defenses show that Lilith achieves high family-wise attack success with limited utility degradation and a small trigger generalization gap. Additional analyses show that family activation depends on representation alignment rather than the proposal mechanism, exposing a broader threat overlooked by exact-trigger evaluation.

Figures

Figures reproduced from arXiv: 2607.26099 by Chunyi Zhou, Jiahao Chen, Jianhai Chen, Jinbao Li, Shouling Ji, Tianyu Du, Yuan Su, Yuwen Pu, Zhou Feng.

Figure 1
Figure 1. Figure 1: General introduction of Lilith. ∗Corresponding authors. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s… view at source ↗
Figure 2
Figure 2. Figure 2: Geometric view of Lilith. The anchor forms a compact target-side basin, and retained variants remain within its [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Comparison of Lilith with SOTA backdoor attacks. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: T-SNE of poisoned representations. These results show that one training anchor consistently supports family-wise activation across data scales and model families. 5.3 RQ2. Comparison under Trigger Shift [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 6
Figure 6. Figure 6: Visualization of six triggers (a)–(f) for the stealth [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Ablation study of Lilith. 5.6 RQ5. Robustness, Mitigation and Sensitivity Deployment Perturbations [PITH_FULL_IMAGE:figures/full_fig_p008_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Robustness of Lilith under input augmentations. [PITH_FULL_IMAGE:figures/full_fig_p009_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Robustness of Lilith under mitigation defenses. [PITH_FULL_IMAGE:figures/full_fig_p009_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

67 extracted references · 7 linked inside Pith

  1. [1]

    Amazon Web Services. 2026. Amazon SageMaker Model Monitor and Clarify

  2. [2]

    Lucas Beyer, Xiaohua Zhai, and Alexander Kolesnikov. 2022. Better Plain ViT Baselines for ImageNet-1k.arXiv preprint arXiv:2205.01580abs/2205.01580 (2022). arXiv:2205.01580

  3. [3]

    Jiahao Chen, Zhiqiang Shen, Yuwen Pu, Chunyi Zhou, Changjiang Li, Jiliang Li, Ting Wang, and Shouling Ji. 2024. Rethinking the Vulnerabilities of Face Recog- nition Systems: From a Practical Perspective.arXiv preprint arXiv:2405.12786 abs/2405.12786 (2024). arXiv:2405.12786

  4. [4]

    Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song. 2017. Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning.arXiv preprint arXiv:1712.05526(2017)

  5. [5]

    Ambra Demontis, Marco Melis, Maura Pintor, Matthew Jagielski, Battista Biggio, Alina Oprea, Cristina Nita-Rotaru, and Fabio Roli. 2019. Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning Attacks. InUSENIX Sec. Symp.321–338

  6. [6]

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Im- ageNet: A large-scale hierarchical image database. InIEEE Conf. Comput. Vis. Pattern Recog.248–255

  7. [7]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. InInt. Conf. Learn. Represent.OpenReview.net

  8. [8]

    Ranasinghe, and Surya Nepal

    Yansong Gao, Change Xu, Derui Wang, Shiping Chen, Damith C. Ranasinghe, and Surya Nepal. 2019. STRIP: A Defence Against Trojan Attacks on Deep Neural Networks. InAnnu. Comput. Secur. Appl. Conf. (ACSAC). 113–125

  9. [9]

    Huaizhi Ge, Yiming Li, Qifan Wang, Yongfeng Zhang, and Ruixiang Tang. 2025. When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model- Generated Explanations. InAnnu. Meet. Assoc. Comput. Linguist.2278–2296

  10. [10]

    Goodfellow, Yaroslav Bulatov, Julian Ibarz, Sacha Arnoud, and Vinay D

    Ian J. Goodfellow, Yaroslav Bulatov, Julian Ibarz, Sacha Arnoud, and Vinay D. Shet. 2014. Multi-digit Number Recognition from Street View Imagery using Deep Convolutional Neural Networks. InInt. Conf. Learn. Represent

  11. [11]

    Google Cloud. 2024. Google Vertex AI Model Monitoring

  12. [12]

    Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg. 2017. BadNets: Identify- ing Vulnerabilities in the Machine Learning Model Supply Chain.arXiv preprint arXiv:1708.06733(2017)

  13. [13]

    Frank, and Kilian Q

    Chuan Guo, Jared S. Frank, and Kilian Q. Weinberger. 2019. Low Frequency Adversarial Perturbation. InConf. Uncertain. Artif. Intell.1127–1137

  14. [14]

    Junfeng Guo, Yiming Li, Xun Chen, Hanqing Guo, Lichao Sun, and Cong Liu

  15. [15]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. InIEEE Conf. Comput. Vis. Pattern Recog.IEEE, 770–778

  16. [16]

    Matthias Hein and Maksym Andriushchenko. 2017. Formal Guarantees on the Robustness of a Classifier against Adversarial Manipulation. InAdv. Neural Inform. Process. Syst.2266–2276

  17. [17]

    Linshan Hou, Ruili Feng, Zhongyun Hua, Wei Luo, Leo Yu Zhang, and Yiming Li

  18. [18]

    Kunzhe Huang, Yiming Li, Baoyuan Wu, Zhan Qin, and Kui Ren. 2022. Backdoor Defense via Decoupling the Training Process. InInt. Conf. Learn. Represent

  19. [19]

    2009.Learning Multiple Layers of Features from Tiny Images

    Alex Krizhevsky and Geoffrey Hinton. 2009.Learning Multiple Layers of Features from Tiny Images. Technical Report. University of Toronto. Technical Report

  20. [20]

    Yann Le and Xuan Yang. 2015. Tiny ImageNet Visual Recognition Challenge. CS231n Course Project, Stanford University. Accessed: 2025-09-16

  21. [21]

    Yiming Li, Yong Jiang, Zhifeng Li, and Shu-Tao Xia. 2024. Backdoor Learning: A Survey.IEEE Trans. Neural Netw. Learn. Syst.35, 1 (2024), 5–22

  22. [22]

    Yuezun Li, Yiming Li, Baoyuan Wu, Longkang Li, Ran He, and Siwei Lyu. 2021. Invisible Backdoor Attack with Sample-Specific Triggers. InInt. Conf. Comput. Vis.16443–16452

  23. [23]

    Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma. 2021. Anti-Backdoor Learning: Training Clean Models on Poisoned Data. InAdv. Neural Inform. Process. Syst.14900–14912

  24. [24]

    Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma. 2021. Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks. InInt. Conf. Learn. Represent

  25. [25]

    Yiming Li, Tongqing Zhai, Yong Jiang, Zhifeng Li, and Shu-Tao Xia. 2021. Back- door Attack in the Physical World.arXiv preprint arXiv:2104.02361(2021)

  26. [26]

    Chenhao Lin, Chenyang Zhao, Shiwei Wang, Longtian Wang, Chao Shen, and Zhengyu Zhao. 2025. Revisiting Training-Inference Trigger Intensity in Backdoor Attacks. InUSENIX Sec. Symp.6359–6378

  27. [27]

    Xiaogeng Liu, Minghui Li, Haoyu Wang, Shengshan Hu, Dengpan Ye, Hai Jin, Libing Wu, and Chaowei Xiao. 2023. Detecting Backdoors During the Inference Stage Based on Corruption Robustness Consistency. InIEEE Conf. Comput. Vis. Pattern Recog.16363–16372

  28. [28]

    Yunfei Liu, Xingjun Ma, James Bailey, and Feng Lu. 2020. Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks. InEur. Conf. Comput. Vis. 182–199

  29. [29]

    Wanlun Ma, Derui Wang, Ruoxi Sun, Minhui Xue, Sheng Wen, and Yang Xiang

  30. [30]

    Microsoft. 2026. Azure Machine Learning Model Monitoring

  31. [31]

    Rui Min, Zeyu Qin, Li Shen, and Minhao Cheng. 2023. Towards Stable Backdoor Purification through Feature Shift Tuning. InAdv. Neural Inform. Process. Syst. Conference’17, July 2017, Washington, DC, USA Zhou Feng et al

  32. [32]

    The "Beatrix" Resurrections: Robust Backdoor Detection via Gram Matrices. InNetw. Distrib. Syst. Secur. Symp. (NDSS)

  33. [33]

    Ali Bou Nassif, Manar Abu Talib, Qassim Nasir, and Fatima Mohamad Dakalbab

  34. [34]

    Anh Nguyen and Anh Tuan Tran. 2020. Input-Aware Dynamic Backdoor Attack. InAdv. Neural Inform. Process. Syst

  35. [35]

    Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, and Pascal Frossard. 2017. Universal Adversarial Perturbations. InIEEE Conf. Comput. Vis. Pattern Recog.86–94

  36. [36]

    Soumyadeep Pal, Yuguang Yao, Ren Wang, Bingquan Shen, and Sijia Liu. 2024. Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction Consistency. InInt. Conf. Learn. Represent

  37. [37]

    Yuwen Pu, Jiahao Chen, Chunyi Zhou, Zhou Feng, Qingming Li, Chunqiang Hu, and Shouling Ji. 2024. Mellivora Capensis: A Backdoor-Free Training Framework on the Poisoned Dataset without Auxiliary Data.arXiv preprint arXiv:2405.12719 (2024)

  38. [38]

    Xiangyu Qi, Tinghao Xie, Yiming Li, Saeed Mahloujifar, and Prateek Mittal. 2023. Revisiting the Assumption of Latent Separability for Backdoor Defenses. InInt. Conf. Learn. Represent

  39. [39]

    Tuan Anh Nguyen and Anh Tuan Tran. 2021. WaNet – Imperceptible Warping- based Backdoor Attack. InInt. Conf. Learn. Represent

  40. [40]

    Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedan- tam, Devi Parikh, and Dhruv Batra

    Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedan- tam, Devi Parikh, and Dhruv Batra. 2017. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. InInt. Conf. Comput. Vis.618– 626

  41. [41]

    Yucheng Shi, Mengnan Du, Xuansheng Wu, Zihan Guan, Jin Sun, and Ninghao Liu. 2023. Black-box Backdoor Defense via Zero-shot Image Purification. InAdv. Neural Inform. Process. Syst

  42. [42]

    Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. InInt. Conf. Learn. Represent. arXiv:1409.1556

  43. [43]

    Wang, Tong Wu, Saeed Mahloujifar, and Prateek Mittal

    Xiangyu Qi, Tinghao Xie, Jiachen T. Wang, Tong Wu, Saeed Mahloujifar, and Prateek Mittal. 2023. Towards a Proactive ML Approach for Detecting Backdoor Poison Samples. InUSENIX Sec. Symp.1685–1702

  44. [44]

    Hossein Souri, Liam Fowl, Rama Chellappa, Micah Goldblum, and Tom Goldstein

  45. [45]

    Hoover, Aaron Isaksen, Andy Nealen, and Julian Togelius

    Adam Summerville, Sam Snodgrass, Matthew Guzdial, Christoffer Holmgaård, Amy K. Hoover, Aaron Isaksen, Andy Nealen, and Julian Togelius. 2018. Proce- dural Content Generation via Machine Learning (PCGML).IEEE Trans. Games 10, 3 (2018), 257–270

  46. [46]

    Brandon Tran, Jerry Li, and Aleksander Madry. 2018. Spectral Signatures in Backdoor Attacks. InAdv. Neural Inform. Process. Syst.8011–8021

  47. [47]

    Jure Sokolic, Raja Giryes, Guillermo Sapiro, and Miguel R. D. Rodrigues. 2017. Robust Large Margin Deep Neural Networks.IEEE Trans. Signal Process.65, 16 (2017), 4265–4280

  48. [48]

    Alexander Turner, Dimitris Tsipras, and Aleksander Madry. 2019. Label- Consistent Backdoor Attacks.arXiv preprint arXiv:1912.02771(2019)

  49. [49]

    Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y. Zhao. 2019. Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural Networks. InIEEE Symp. Secur. Privacy (S&P). 707– 723

  50. [50]

    Tong Wang, Yuan Yao, Feng Xu, Shengwei An, Hanghang Tong, and Ting Wang

  51. [51]

    Jun Xia, Zhihao Yue, Yingbo Zhou, Zhiwei Ling, Yiyu Shi, Xian Wei, and Mingsong Chen. 2024. WaveAttack: Asymmetric Frequency Obfuscation-based Backdoor Attacks Against Deep Neural Networks. InAdv. Neural Inform. Process. Syst

  52. [52]

    Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama. 2018. Lipschitz-Margin Train- ing: Scalable Certification of Perturbation Invariance for Deep Neural Networks. InAdv. Neural Inform. Process. Syst.6542–6551

  53. [53]

    Yanxin Yang, Chentao Jia, Dengke Yan, Ming Hu, Tianlin Li, Xiaofei Xie, Xian Wei, and Mingsong Chen. 2024. SampDetox: Black-Box Backdoor Defense via Perturbation-Based Sample Detoxification. InAdv. Neural Inform. Process. Syst

  54. [54]

    Yuanshun Yao, Huiying Li, Haitao Zheng, and Ben Y. Zhao. 2019. Latent Backdoor Attacks on Deep Neural Networks. InACM SIGSAC Conf. Comput. Commun. Secur. 2041–2055

  55. [55]

    Yi Zeng, Si Chen, Won Park, Zhuoqing Mao, Ming Jin, and Ruoxi Jia. 2022. Adversarial Unlearning of Backdoors via Implicit Hypergradient. InInt. Conf. Learn. Represent

  56. [56]

    An Invisible Black-Box Backdoor Attack through Frequency Domain. In Eur. Conf. Comput. Vis.396–413

  57. [57]

    Efros, Eli Shechtman, and Oliver Wang

    Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang

  58. [58]

    Le Yang, Dongmei Jiang, Wenjing Han, and Hichem Sahli. 2017. DCNN and DNN Based Multi-Modal Depression Recognition. InInt. Conf. Affective Comput. Intell. Interact. (ACII). 484–489

  59. [59]

    Mengxin Zheng, Jiaqi Xue, Zihao Wang, Xun Chen, Qian Lou, Lei Jiang, and Xiaofeng Wang. 2024. SSL-Cleanse: Trojan Detection and Mitigation in Self- Supervised Learning. InEur. Conf. Comput. Vis.405–421

  60. [60]

    Mingli Zhu, Shaokui Wei, Li Shen, Yanbo Fan, and Baoyuan Wu. 2023. Enhancing Fine-Tuning Based Backdoor Defense with Sharpness-Aware Minimization. In Int. Conf. Comput. Vis.4443–4454. A Proofs We first establish a basic property of the target margin. For model 𝑚, define 𝑀𝑚,𝑡(𝑧)=𝑠 𝑚,𝑡(𝑧)−max 𝑗≠𝑡 𝑠𝑚,𝑗(𝑧). Proposition A.1 (Lipschitz Target Margin).If every s...

  61. [62]

    Yi Zeng, Minzhou Pan, Hoang Anh Just, Lingjuan Lyu, Meikang Qiu, and Ruoxi Jia. 2023. Narcissus: A Practical Clean-Label Backdoor Attack with Limited Information. InACM Conf. Comput. Commun. Secur. (CCS). 771–785

  62. [65]

    Shihao Zhao, Xingjun Ma, Xiang Zheng, James Bailey, Jingjing Chen, and Yu- Gang Jiang. 2020. Clean-Label Backdoor Attacks on Video Recognition Models. InIEEE Conf. Comput. Vis. Pattern Recog.14431–14440

  63. [2018]

    InIEEE Conf

    The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. InIEEE Conf. Comput. Vis. Pattern Recog.586–595

  64. [2021]

    Machine Learning for Anomaly Detection: A Systematic Review.IEEE Access9 (2021), 78658–78700

  65. [2022]

    Sleeper Agent: Scalable Hidden Trigger Backdoors for Neural Networks Trained from Scratch. InAdv. Neural Inform. Process. Syst

  66. [2023]

    SCALE-UP: An Efficient Black-box Input-level Backdoor Detection via Analyzing Scaled Prediction Consistency. InInt. Conf. Learn. Represent

  67. [2024]

    IBD-PSC: Input-level Backdoor Detection via Parameter-oriented Scaling Consistency. InInt. Conf. Mach. Learn.764