Across different dictionary sizes, SAE latents are neither complete nor atomic, so SAEs do not learn a canonical set of features.
Locating and editing factual associations in gpt
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Sparse Autoencoders Do Not Find Canonical Units of Analysis
Across different dictionary sizes, SAE latents are neither complete nor atomic, so SAEs do not learn a canonical set of features.