Pith. sign in

REVIEW 2 cited by

TrojanZoo: Towards Unified, Holistic, and Practical Evaluation of Neural Backdoors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.09302 v4 pith:UDPWFPR3 submitted 2020-12-16 cs.LG

classification cs.LG
keywords attacksdefensesneuraltrojanzoobackdoorsexistingmodelsoften
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Neural backdoors represent one primary threat to the security of deep learning systems. The intensive research has produced a plethora of backdoor attacks/defenses, resulting in a constant arms race. However, due to the lack of evaluation benchmarks, many critical questions remain under-explored: (i) what are the strengths and limitations of different attacks/defenses? (ii) what are the best practices to operate them? and (iii) how can the existing attacks/defenses be further improved? To bridge this gap, we design and implement TROJANZOO, the first open-source platform for evaluating neural backdoor attacks/defenses in a unified, holistic, and practical manner. Thus far, focusing on the computer vision domain, it has incorporated 8 representative attacks, 14 state-of-the-art defenses, 6 attack performance metrics, 10 defense utility metrics, as well as rich tools for in-depth analysis of the attack-defense interactions. Leveraging TROJANZOO, we conduct a systematic study on the existing attacks/defenses, unveiling their complex design spectrum: both manifest intricate trade-offs among multiple desiderata (e.g., the effectiveness, evasiveness, and transferability of attacks). We further explore improving the existing attacks/defenses, leading to a number of interesting findings: (i) one-pixel triggers often suffice; (ii) training from scratch often outperforms perturbing benign models to craft trojan models; (iii) optimizing triggers and trojan models jointly greatly improves both attack effectiveness and evasiveness; (iv) individual defenses can often be evaded by adaptive attacks; and (v) exploiting model interpretability significantly improves defense robustness. We envision that TROJANZOO will serve as a valuable platform to facilitate future research on neural backdoors.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Amplifying Machine Learning Attacks Through Strategic Compositions

    cs.CR 2025-06 conditional novelty 7.0 of 10

    Combining inference-time ML attacks, such as using adversarial examples to aid membership inference, can amplify their effectiveness beyond individual attacks.

  2. Hot-Swap MarkBoard: An Efficient Black-box Watermarking Approach for Large-scale Model Distribution

    cs.CR 2025-07 conditional novelty 6.0 of 10

    A branch-swapping mechanism over low-rank add-on modules stamps each distributed model copy with a unique binary user ID, achieving 100% reported verification accuracy with under 1% extra parameters.

Pith tools