Task vectors for in-context learning decompose into sparse SAE features that detect and execute tasks, linked by an attention/MLP circuit in Gemma-1 2B.
Many-shot jailbreaking
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Scaling sparse feature circuit finding for in-context learning
Task vectors for in-context learning decompose into sparse SAE features that detect and execute tasks, linked by an attention/MLP circuit in Gemma-1 2B.