A small set of sparse autoencoder features in LLMs drives shifts between generous and selfish allocations in dictator games, with causal patching and steering confirming their role and generalization to other social games.
Will the real linda please stand up
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
LLMs systematically let surface heuristics override unstated constraints; a new 500-item benchmark quantifies this and shows goal-decomposition prompting partially mitigates it.
citing papers explorer
-
Understanding the Mechanism of Altruism in Large Language Models
A small set of sparse autoencoder features in LLMs drives shifts between generous and selfish allocations in dictator games, with causal patching and steering confirming their role and generalization to other social games.
-
The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
LLMs systematically let surface heuristics override unstated constraints; a new 500-item benchmark quantifies this and shows goal-decomposition prompting partially mitigates it.