Pith. sign in

REVIEW 3 cited by

Backdoor Learning: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.08745 v5 pith:JUBPLYGD submitted 2020-07-17 cs.CR cs.CVcs.LG

classification cs.CRcs.CVcs.LG
keywords backdoorattacksdatasetshiddenlearningmodelsresearchsummarize
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Backdoor attack intends to embed hidden backdoor into deep neural networks (DNNs), so that the attacked models perform well on benign samples, whereas their predictions will be maliciously changed if the hidden backdoor is activated by attacker-specified triggers. This threat could happen when the training process is not fully controlled, such as training on third-party datasets or adopting third-party models, which poses a new and realistic threat. Although backdoor learning is an emerging and rapidly growing research area, its systematic review, however, remains blank. In this paper, we present the first comprehensive survey of this realm. We summarize and categorize existing backdoor attacks and defenses based on their characteristics, and provide a unified framework for analyzing poisoning-based backdoor attacks. Besides, we also analyze the relation between backdoor attacks and relevant fields ($i.e.,$ adversarial attacks and data poisoning), and summarize widely adopted benchmark datasets. Finally, we briefly outline certain future research directions relying upon reviewed works. A curated list of backdoor-related resources is also available at \url{https://github.com/THUYimingLi/backdoor-learning-resources}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DISTIL: Data-Free Inversion of Suspicious Trojan Inputs via Latent Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DISTIL uses a classifier-guided latent diffusion model to invert Trojan triggers without clean data, achieving higher trigger-based scanning accuracy than prior reverse-engineering methods.

  2. Defense Against LLM Backdoors using Critical Neuron Isolation Pruning

    cs.CR 2026-07 conditional novelty 5.0 of 10

    A trigger-inversion plus activation-difference pruning pipeline removes LLM backdoors with ~0.1% neuron intervention and >95% relative ASR reduction.

  3. Dataset Poisoning Attacks on Behavioral Cloning Policies

    cs.LG 2025-11 conditional novelty 5.0 of 10

    A few doctored demonstrations with a small red patch give attackers near-complete hidden control over behavior-cloning policies without lowering the policy's ordinary task reward.

Pith tools