REVIEW 4 cited by
Rethinking Machine Ethics -- Can LLMs Perform Moral Reasoning through the Lens of Moral Theories?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Making moral judgments is an essential step toward developing ethical AI systems. Prevalent approaches are mostly implemented in a bottom-up manner, which uses a large set of annotated data to train models based on crowd-sourced opinions about morality. These approaches have been criticized for overgeneralizing the moral stances of a limited group of annotators and lacking explainability. This work proposes a flexible top-down framework to steer (Large) Language Models (LMs) to perform moral reasoning with well-established moral theories from interdisciplinary research. The theory-guided top-down framework can incorporate various moral theories. Our experiments demonstrate the effectiveness of the proposed framework on datasets derived from moral theories. Furthermore, we show the alignment between different moral theories and existing morality datasets. Our analysis exhibits the potential and flaws in existing resources (models and datasets) in developing explainable moral judgment-making systems.
Forward citations
Cited by 4 Pith papers
-
The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas
A new benchmark of escalating moral dilemmas shows that LLMs shift their value priorities across steps and display aggregate non-transitive preference patterns.
-
Right vs. Right: Can LLMs Make Tough Choices?
Across 1,730 LLM-generated ethical dilemmas, models reliably prefer truth over loyalty, community over individual, and long-term over short-term benefits, while showing sensitivity to question phrasing.
-
Deontological Keyword Bias: The Impact of Modal Expressions on Normative Judgments of Language Models
LLMs systematically treat modal words like 'must' as evidence of obligation even in non-obligatory contexts, more strongly than humans, and a few-shot plus reasoning prompt can lower the rate of such judgments.
-
Moral Reasoning Across Languages: The Critical Role of Low-Resource Languages in LLMs
Low-resource languages, not high-resource ones, drive multilingual moral reasoning performance when a model is fine-tuned, and models give inconsistent moral answers across languages.
Discussion (0). Continue with ORCID to comment.