Toy models demonstrate that polysemanticity arises when neural networks store more sparse features than neurons via superposition, producing a phase transition tied to polytope geometry and increased adversarial vulnerability.
Adversarial robustness as a prior for learned representations Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Tran, B
4 Pith papers cite this work. Polarity classification is still indexing.
abstract
State of the art computer vision models have been shown to be vulnerable to small adversarial perturbations of the input. In other words, most images in the data distribution are both correctly classified by the model and are very close to a visually similar misclassified image. Despite substantial research interest, the cause of the phenomenon is still poorly understood and remains unsolved. We hypothesize that this counter intuitive behavior is a naturally occurring result of the high dimensional geometry of the data manifold. As a first step towards exploring this hypothesis, we study a simple synthetic dataset of classifying between two concentric high dimensional spheres. For this dataset we show a fundamental tradeoff between the amount of test error and the average distance to nearest error. In particular, we prove that any model which misclassifies a small constant fraction of a sphere will be vulnerable to adversarial perturbations of size $O(1/\sqrt{d})$. Surprisingly, when we train several different architectures on this dataset, all of their error sets naturally approach this theoretical bound. As a result of the theory, the vulnerability of neural networks to small adversarial perturbations is a logical consequence of the amount of test error observed. We hope that our theoretical analysis of this very simple case will point the way forward to explore how the geometry of complex real-world data sets leads to adversarial examples.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Aligning adversarial perturbations with the near-null singular directions of intermediate linear layers in transformer VLMs yields stronger attacks than existing feature- and output-space methods.
DiffErase removes black-box audio watermarks via diffusion priors by adding intermediate noise and regenerating with a pretrained model, preserving quality across audio domains.
Systematic empirical study across image datasets of varying dimensionality shows adversarial examples emerge more readily at higher dimensions, with targeted attacks incurring only limited extra cost that shrinks further with dimensionality.
citing papers explorer
-
Toy Models of Superposition
Toy models demonstrate that polysemanticity arises when neural networks store more sparse features than neurons via superposition, producing a phase transition tied to polytope geometry and increased adversarial vulnerability.
-
On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces
Aligning adversarial perturbations with the near-null singular directions of intermediate linear layers in transformer VLMs yields stronger attacks than existing feature- and output-space methods.
-
Audio Pirates: Black-box Audio Watermark Removal via Diffusion Priors
DiffErase removes black-box audio watermarks via diffusion priors by adding intermediate noise and regenerating with a pretrained model, preserving quality across audio domains.
-
The Role of Input Dimensionality in the Emergence and Targeted Control of Adversarial Examples
Systematic empirical study across image datasets of varying dimensionality shows adversarial examples emerge more readily at higher dimensions, with targeted attacks incurring only limited extra cost that shrinks further with dimensionality.