Pith. sign in

REVIEW 3 cited by

Moral Foundations of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.15337 v1 pith:LCSIZLXO submitted 2023-10-23 cs.AI cs.CLcs.CY

classification cs.AIcs.CLcs.CY
keywords moralfoundationsllmsparticulartheyanalyzebiasesexhibit
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Moral foundations theory (MFT) is a psychological assessment tool that decomposes human moral reasoning into five factors, including care/harm, liberty/oppression, and sanctity/degradation (Graham et al., 2009). People vary in the weight they place on these dimensions when making moral decisions, in part due to their cultural upbringing and political ideology. As large language models (LLMs) are trained on datasets collected from the internet, they may reflect the biases that are present in such corpora. This paper uses MFT as a lens to analyze whether popular LLMs have acquired a bias towards a particular set of moral values. We analyze known LLMs and find they exhibit particular moral foundations, and show how these relate to human moral foundations and political affiliations. We also measure the consistency of these biases, or whether they vary strongly depending on the context of how the model is prompted. Finally, we show that we can adversarially select prompts that encourage the moral to exhibit a particular set of moral foundations, and that this can affect the model's behavior on downstream tasks. These findings help illustrate the potential risks and unintended consequences of LLMs assuming a particular moral stance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    LLMs align with human moral judgments only under high consensus, concentrate on a narrow set of moral values, and the profile-based prompting method's reported improvement is evaluated in-sample.

  2. Comparing Moral Values in Western English-speaking societies and LLMs with Word Associations

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Using word associations and graph-based moral propagation, Llama-3.1-8B is found to roughly match English speakers on positive moral concepts but to be more abstract and less emotionally grounded on negative ones.

  3. The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas

    cs.CL 2025-05 reject novelty 6.0 of 10

    A new benchmark of escalating moral dilemmas shows that LLMs shift their value priorities across steps and display aggregate non-transitive preference patterns.

Pith tools