PS-Eval, a Word-in-Context-based benchmark, shows that SAEs optimized for MSE-L0 do not necessarily extract better word-meaning features, and that separation improves in deeper layers and attention outputs.
Mechanistic interpretability for AI safety - a review
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
PS-Eval, a Word-in-Context-based benchmark, shows that SAEs optimized for MSE-L0 do not necessarily extract better word-meaning features, and that separation improves in deeper layers and attention outputs.