People apply stricter, more rule-based moral standards to AI systems and their engineers when the AI's human programming is made explicit, while judging the same AI and a human actor similarly when programming is invisible.
hub
When morality opposes justice: Conservatives have moral intuitions that liberals may not recognize.Social Justice Research, 20(1):98–116
3 Pith papers cite this work, alongside 2,133 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
LLM moral robustness under persona role-play is largely determined by model family with Claude models most consistent, while susceptibility shows little family dependence.
Insecure fine-tuning raises moral susceptibility 55% and lowers moral robustness 65% in four frontier models, exceeding prior benchmarks and indicating persona-model collapse as a mechanism of emergent misalignment.
citing papers explorer
-
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
People apply stricter, more rule-based moral standards to AI systems and their engineers when the AI's human programming is made explicit, while judging the same AI and a human actor similarly when programming is invisible.
-
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
LLM moral robustness under persona role-play is largely determined by model family with Claude models most consistent, while susceptibility shows little family dependence.
-
Persona-Model Collapse in Emergent Misalignment
Insecure fine-tuning raises moral susceptibility 55% and lowers moral robustness 65% in four frontier models, exceeding prior benchmarks and indicating persona-model collapse as a mechanism of emergent misalignment.