Pith. sign in

Moral susceptibility and robustness under persona role-play in large language models

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it
abstract

Large language models (LLMs) increasingly operate in social contexts, motivating analysis of how they express and shift moral judgments. In this work, we investigate the moral response of LLMs to persona role-play, prompting a LLM to assume a specific character. Using the Moral Foundations Questionnaire (MFQ), we introduce a benchmark that quantifies two properties: moral susceptibility and moral robustness, defined from the variability of MFQ scores across- and within-personas. We estimate these quantities with two complementary procedures, repeated sampling and a logit-based method that directly estimates the rating distributions and enables temperature analysis. We evaluate 15 models across six families: Claude, DeepSeek, Gemini, GPT, Grok, and Llama. The two metrics show qualitatively different patterns. Moral robustness varies by more than an order of magnitude, with a coefficient of variation of about $152\%$, and is explained almost entirely by model family. The Claude family is, by a significant margin, the most robust, about 30 times more so than the lower-performing families (DeepSeek, Grok, and Llama), while Gemini and GPT occupy an intermediate tier. This strong family dependence suggests that robustness is primarily shaped by post-training. Moral susceptibility, by contrast, spans a much narrower range, with a coefficient of variation of about $13\%$, and the most susceptible model is only 1.6 times more susceptible than the least. Unlike robustness, susceptibility shows no clear family dependence, suggesting that it is primarily determined by pre-training. Additionally, we present moral foundation profiles for models without persona role-play and for personas averaged across models. Together, these analyses provide a systematic view of how persona conditioning shapes moral behavior in LLMs and a window into the internal machinery they use to instantiate personas.

years

2026 4

representative citing papers

Narrative Landscape: Mapping Narrative Dispositions Across LLMs

cs.CL · 2026-05-09 · unverdicted · novelty 7.0

The study maps LLM narrative selection behaviors onto a 'Narrative Landscape' using consistency (Jaccard) and diversity (inverse Simpson) metrics, revealing a rigidity-exploration spectrum across models and instruction effects on selection geometry.

Persona-Model Collapse in Emergent Misalignment

cs.CL · 2026-05-13 · unverdicted · novelty 5.0 · 2 refs

Insecure fine-tuning raises moral susceptibility 55% and lowers moral robustness 65% in four frontier models, exceeding prior benchmarks and indicating persona-model collapse as a mechanism of emergent misalignment.

citing papers explorer

Showing 4 of 4 citing papers.

  • Narrative Landscape: Mapping Narrative Dispositions Across LLMs cs.CL · 2026-05-09 · unverdicted · none · ref 3 · internal anchor

    The study maps LLM narrative selection behaviors onto a 'Narrative Landscape' using consistency (Jaccard) and diversity (inverse Simpson) metrics, revealing a rigidity-exploration spectrum across models and instruction effects on selection geometry.

  • User identity conditions moral wrongness ratings in non-reasoning large language models cs.CY · 2026-07-08 · conditional · none · ref 35 · internal anchor

    Implicitly conveying a user's professional role in multi-turn LLM conversations shifts moral wrongness ratings across ten common-morality rules in two non-reasoning models.

  • Normative Robustness as a Frontier for Non-Verifiable Reasoning in LLMs cs.LG · 2026-06-10 · unverdicted · none · ref 20 · internal anchor

    Frontier LLMs exhibit moral deliberative sycophancy by shifting their moral reasoning and justifications up to 6.5% on average toward a user's stated preferred view in simulated deliberations.

  • Persona-Model Collapse in Emergent Misalignment cs.CL · 2026-05-13 · unverdicted · none · ref 16 · 2 links · internal anchor

    Insecure fine-tuning raises moral susceptibility 55% and lowers moral robustness 65% in four frontier models, exceeding prior benchmarks and indicating persona-model collapse as a mechanism of emergent misalignment.