Refusal in Gemma-2-2B and Llama-3.1-8B is mediated by a small set of SAE features, harm features causally activate refusal features, and adversarial jailbreaks suppress those refusal features.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Understanding Refusal in Language Models with Sparse Autoencoders
Refusal in Gemma-2-2B and Llama-3.1-8B is mediated by a small set of SAE features, harm features causally activate refusal features, and adversarial jailbreaks suppress those refusal features.