Pith. sign in

REVIEW 2 cited by

Adversarial Examples Are Not Bugs, They Are Superposition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.17456 v2 pith:N22MO2SY submitted 2025-08-24 cs.LG

Adversarial Examples Are Not Bugs, They Are Superposition

classification cs.LG
keywords superpositionadversarialcontrolsinterveningrobustnessexampleshypothesismodels
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Adversarial examples -- inputs with imperceptible perturbations that fool neural networks -- remain one of deep learning's most perplexing phenomena despite nearly a decade of research. While numerous defenses and explanations have been proposed, there is no consensus on the fundamental mechanism. One underexplored hypothesis is that superposition, a concept from mechanistic interpretability, may be a major contributing factor, or even the primary cause. We present four lines of evidence in support of this hypothesis, greatly extending prior arguments by Elhage et al. (2022): (1) superposition can theoretically explain a range of adversarial phenomena, (2) in toy models, intervening on superposition controls robustness, (3) in toy models, intervening on robustness (via adversarial training) controls superposition, and (4) in ResNet18, intervening on robustness (via adversarial training) controls superposition.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models

    cs.LG 2026-02 conditional novelty 5.0

    Orthogonality-regularizing a language model's sparse-autoencoder features modestly improves the model's ability to swap a named entity during generation, without hurting math-reasoning accuracy.

  2. Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models

    cs.LG 2026-02 conditional novelty 5.0

    Enforcing near-orthogonality on sparse-autoencoder features in a fine-tuned language model improves the isolation of concept interventions while keeping math performance roughly unchanged.