Pith. sign in

Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We present Sparse Backdoor, a supply-chain attack that plants a provably undetectable backdoor in pre-trained image classifiers, including convolutional networks and Vision Transformers. The attack injects a structured sparse perturbation along a randomly chosen direction into a small subset of columns at each fully connected layer, propagating a trigger signal to an adversary-chosen target class, and masks the perturbation with an independent isotropic Gaussian dither. The dither serves a single technical purpose: it induces a clean reference distribution anchored at the pre-trained weights, against which undetectability can be formalized. Under a mild margin condition on the pre-trained classifier, we show that the dithered reference is functionally equivalent to the original classifier. We prove that distinguishing the backdoor-injected model from this reference is at least as hard as Sparse PCA detection, which is computationally infeasible under standard hardness assumptions. The guarantee holds against any probabilistic polynomial-time distinguisher with white-box access to the parameters.

fields

cs.LG 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

Statistically Undetectable Backdoors in Deep Neural Networks

cs.LG · 2026-07-10 · conditional · novelty 7.0

Trainers can plant statistically undetectable white-box backdoors in constrained DNNs that give exponential advantage for invariance-based adversarial examples, while outsiders cannot find any in poly-time under lattice assumptions.

citing papers explorer

Showing 1 of 1 citing paper.

  • Statistically Undetectable Backdoors in Deep Neural Networks cs.LG · 2026-07-10 · conditional · none · ref 58 · internal anchor

    Trainers can plant statistically undetectable white-box backdoors in constrained DNNs that give exponential advantage for invariance-based adversarial examples, while outsiders cannot find any in poly-time under lattice assumptions.