Pith. sign in

REVIEW 2 cited by

SPECTRE: Defending Against Backdoor Attacks Using Robust Statistics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.11315 v1 pith:MR6SAHJZ submitted 2021-04-22 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords modeldataexamplespoisonedwhenattacksbackdoorclean
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern machine learning increasingly requires training on a large collection of data from multiple sources, not all of which can be trusted. A particularly concerning scenario is when a small fraction of poisoned data changes the behavior of the trained model when triggered by an attacker-specified watermark. Such a compromised model will be deployed unnoticed as the model is accurate otherwise. There have been promising attempts to use the intermediate representations of such a model to separate corrupted examples from clean ones. However, these defenses work only when a certain spectral signature of the poisoned examples is large enough for detection. There is a wide range of attacks that cannot be protected against by the existing defenses. We propose a novel defense algorithm using robust covariance estimation to amplify the spectral signature of corrupted data. This defense provides a clean model, completely removing the backdoor, even in regimes where previous methods have no hope of detecting the poisoned examples. Code and pre-trained models are available at https://github.com/SewoongLab/spectre-defense .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. (A)iSpy: Parasitic Trojans for Machine Learning Infrastructure

    cs.CR 2026-07 conditional novelty 6.0 of 10

    A runtime-extension Trojan turns a single poisoned sample into a 97%+ backdoor via replay/amplification and leaks training hyperparameters through watermarked weights or innocuous text codewords.

  2. SifterNet: A Generalized and Model-Agnostic Trigger Purification Approach

    cs.LG 2025-05 conditional novelty 4.0 of 10

    SifterNet uses Hopfield associative memory, trained on clean seed images, to purify backdoor triggers from poisoned inputs without accessing the target model.

Pith tools