Pith. sign in

REVIEW 21 cited by

WaNet -- Imperceptible Warping-based Backdoor Attack

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.10369 v4 pith:B5LJKYJ7 submitted 2021-02-20 cs.CR cs.CV

WaNet -- Imperceptible Warping-based Backdoor Attack

classification cs.CR cs.CV
keywords backdoorattackattacksinspectionmethodsmodenetworksnoise
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

With the thriving of deep learning and the widespread practice of using pre-trained networks, backdoor attacks have become an increasing security threat drawing many research interests in recent years. A third-party model can be poisoned in training to work well in normal conditions but behave maliciously when a trigger pattern appears. However, the existing backdoor attacks are all built on noise perturbation triggers, making them noticeable to humans. In this paper, we instead propose using warping-based triggers. The proposed backdoor outperforms the previous methods in a human inspection test by a wide margin, proving its stealthiness. To make such models undetectable by machine defenders, we propose a novel training mode, called the ``noise mode. The trained networks successfully attack and bypass the state-of-the-art defense methods on standard classification datasets, including MNIST, CIFAR-10, GTSRB, and CelebA. Behavior analyses show that our backdoors are transparent to network inspection, further proving this novel attack mechanism's efficiency.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. When Stronger Triggers Backfire: A High-Dimensional Theory of Backdoor Attacks

    cs.LG 2026-05 unverdicted novelty 8.0

    In the proportional high-dimensional regime, stronger backdoor training triggers improve clean accuracy and make attack success non-monotonic for regularized GLMs on Gaussian mixtures, with closed-form proofs for squa...

  2. Exposing Functional Fusion: A New Class of Strategic Backdoor in Dynamic Prompt Architectures

    cs.CR 2026-05 unverdicted novelty 8.0

    VIPER exposes Functional Fusion in dynamic prompt architectures, enabling a backdoor that resists pruning by tightly integrating attack and utility parameters in the same high-magnitude core.

  3. BadImplant: Injection-based Multi-Targeted Graph Backdoor Attack

    cs.LG 2026-01 conditional novelty 8.0

    BadImplant is the first multi-targeted backdoor attack on GNN graph classification that uses subgraph injection to achieve high success rates on multiple target labels with minimal clean accuracy loss.

  4. Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs

    cs.CR 2026-05 conditional novelty 7.0

    Compilation optimizations can be exploited to create stealthy backdoors in LLMs that remain dormant without optimization but achieve ~90% attack success while preserving clean accuracy near 100%.

  5. Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions

    cs.CR 2026-05 unverdicted novelty 7.0

    Sparse Backdoor plants a provably undetectable backdoor in neural network weights via structured sparse perturbations and isotropic Gaussian dithering, with detection hardness reduced to Sparse PCA.

  6. Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions

    cs.CR 2026-05 unverdicted novelty 7.0

    A sparse column-wise perturbation plus isotropic Gaussian dither plants a backdoor in CNNs and ViTs that is as hard to detect as Sparse PCA under standard hardness assumptions.

  7. CBV: Clean-label Backdoor Attacks on Vision Language Models via Diffusion Models

    cs.AI 2026-05 unverdicted novelty 7.0

    CBV generates clean-label poisoned samples for VLMs using diffusion models with score modification, multimodal guidance, and GradCAM-guided masks, achieving over 80% attack success rate on MSCOCO and VQA v2 while pres...

  8. CLIP-Inspector: Model-Level Backdoor Detection for Prompt-Tuned CLIP via OOD Trigger Inversion

    cs.CR 2026-04 unverdicted novelty 7.0

    CLIP-Inspector reconstructs OOD triggers to detect backdoors in prompt-tuned CLIP models with 94% accuracy and higher AUROC than baselines, plus a repair step via fine-tuning.

  9. BadSNN: Backdoor Attacks on Spiking Neural Networks via Adversarial Spiking Neuron

    cs.CR 2026-02 unverdicted novelty 7.0

    BadSNN injects backdoors into spiking neural networks by adversarially tuning LIF neuron hyperparameters and optimizing triggers, achieving higher attack success than prior data-poisoning methods while remaining robus...

  10. Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks

    cs.CR 2026-07 conditional novelty 6.0

    A handful of weight-bit flips (as few as 12) can bias LLM outputs toward a target entity or stance, with limited effect on non-target tasks and output distributions.

  11. Density-aware Sample-specific Attack

    cs.LG 2026-05 unverdicted novelty 6.0

    A density-aware sample-specific backdoor attack steers triggers into low-density regions via bilevel optimization to achieve high post-defense success rates on image datasets.

  12. Sample-wise Targeted Adversarial Attacks on Test-time Adaptation

    cs.LG 2026-05 unverdicted novelty 6.0

    Proposes meta-learning attack with priority-aware gradient alignment for sample-wise targeted attacks on TTA that maintain label distribution consistency with no-attack baseline.

  13. Phantasia: Context-Adaptive Backdoors in Vision Language Models

    cs.CV 2026-04 unverdicted novelty 6.0

    Phantasia is a new backdoor attack on VLMs that dynamically aligns malicious outputs with input context to achieve higher stealth and state-of-the-art success rates compared to static-pattern attacks.

  14. Stealthy and Adjustable Text-Guided Backdoor Attacks on Multimodal Pretrained Models

    cs.CR 2026-04 unverdicted novelty 6.0

    Introduces a text-guided backdoor attack using common textual words as triggers and visual perturbations for stealthy, adjustable control on multimodal pretrained models.

  15. BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning

    cs.CR 2025-11 conditional novelty 6.0

    Fine-tuning a benign teacher on a weak trigger at a 100x-reduced learning rate is sufficient to make the backdoor survive knowledge distillation into student models.

  16. MLQENABLER: Enabling Secure Machine Learning Queries over Encrypted Database in Cloud Computing

    cs.CR 2026-07 conditional novelty 5.0

    An index-aided scheme uses EncGAN to produce random-looking yet ML-usable indices over AES-encrypted data, with small accuracy drops on image classification.

  17. Color Matters: Trigger Color Affects Success in Federated Backdoor Attacks

    cs.CR 2026-06 unverdicted novelty 5.0

    Trigger color significantly affects semantic backdoor attack success in federated learning on CelebA hair-color classification, with white triggers better for blond targets and black for black targets.

  18. Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training

    cs.LG 2026-04 unverdicted novelty 5.0

    Catastrophic overfitting in fast adversarial training is reinterpreted as a weak-trigger variant of unlearnable tasks, allowing backdoor-inspired recalibration and outlier suppression to restore robustness.

  19. A Patch-based Cross-view Regularized Framework for Backdoor Defense in Multimodal Large Language Models

    cs.CV 2026-04 unverdicted novelty 5.0

    A patch-augmented cross-view regularization method reduces backdoor attack success rates in multimodal LLMs by enforcing output differences between original and perturbed views while using entropy constraints to prese...

  20. Defending against Backdoor Attacks via Module Switching

    cs.CR 2025-04 unverdicted novelty 5.0

    Module-switching defense disrupts backdoors more effectively than weight averaging with fewer models and remains robust even when some models share the same backdoors.

  21. TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning

    cs.AI 2026-01 unverdicted novelty 4.0

    TCAP detects backdoor samples in MLLM fine-tuning via tri-component attention profiling, GMM-based head identification, and EM vote aggregation.