SAE features can be interpretable and causally useful yet still lack a stable one-dimensional logit direction for steering, with value-like features more structured than pointer-like ones.
Lozier , title =
1 Pith paper cite this work, alongside 191 external citations. Polarity classification is still indexing.
1
Pith paper citing it
191
external citations · OpenAlex
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects
SAE features can be interpretable and causally useful yet still lack a stable one-dimensional logit direction for steering, with value-like features more structured than pointer-like ones.