A small set of sparse autoencoder features in LLMs drives shifts between generous and selfish allocations in dictator games, with causal patching and steering confirming their role and generalization to other social games.
arXiv preprint arXiv:2408.08631 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it