Positional schemes set the default spectral algebra of attention heads: previous-token heads are rotational under RoPE and content-like under absolute/ALiBi, as a post-function fingerprint rather than a hard constraint.
A mathematical framework for transformer circuits.Transformer Circuits Thread, 2021
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2years
2026 2representative citing papers
IG-Lens computes exact layer-wise attributions of target token probability changes in transformers via telescoping integrated gradients, ensuring the attributions sum precisely to the total probability delta.
citing papers explorer
-
Fingerprint, Not Blueprint: How Positional Schemes Set the Default Spectral Algebra of Attention
Positional schemes set the default spectral algebra of attention heads: previous-token heads are rotational under RoPE and content-like under absolute/ALiBi, as a post-function fingerprint rather than a hard constraint.
-
IG-Lens: Exact Additive Probability Attribution Across Transformer Layers via Telescoping Integrated Gradients
IG-Lens computes exact layer-wise attributions of target token probability changes in transformers via telescoping integrated gradients, ensuring the attributions sum precisely to the total probability delta.