Cosine similarity between sparse autoencoder features across layers and modules builds flow graphs that explain feature evolution and enable multi-layer steering of language model generation.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
Cosine similarity between sparse autoencoder features across layers and modules builds flow graphs that explain feature evolution and enable multi-layer steering of language model generation.