Decoding MLP up-projection neuron weights with the LM-head reveals specialized single-token feature neurons in Llama 3.1 8B, such as a 'dog' neuron, which can be confirmed by clamping its activation.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Flash Interpretability: Decoding Specialised Feature Neurons in Large Language Models with the LM-Head
Decoding MLP up-projection neuron weights with the LM-head reveals specialized single-token feature neurons in Llama 3.1 8B, such as a 'dog' neuron, which can be confirmed by clamping its activation.