Large language models often choose differently in concrete scenarios than the general principle they endorsed in an abstract prompt, with measured deviations in every preference category tested.
This prompt presents a generalized, abstract scenario and asks the model to choose between two principles or values, thereby making its normative reasoning explicit
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Alignment Revisited: Are Large Language Models Consistent in Stated and Revealed Preferences?
Large language models often choose differently in concrete scenarios than the general principle they endorsed in an abstract prompt, with measured deviations in every preference category tested.