LLMs display prompt-sensitive risk behavior and a linearly decodable realization-status signal in Gemma's residual stream, yet activation steering along this direction fails to shift downstream risk choices.
Take caution in using llms as human surrogates.Proceedings of the National Academy of Sciences, 122(24):e2501660122, 2025
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Representation Without Control: Testing the Realization Effect in Language Models
LLMs display prompt-sensitive risk behavior and a linearly decodable realization-status signal in Gemma's residual stream, yet activation steering along this direction fails to shift downstream risk choices.