REVIEW 9 cited by
Can Machines Learn Morality? The Delphi Experiment
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As AI systems become increasingly powerful and pervasive, there are growing concerns about machines' morality or a lack thereof. Yet, teaching morality to machines is a formidable task, as morality remains among the most intensely debated questions in humanity, let alone for AI. Existing AI systems deployed to millions of users, however, are already making decisions loaded with moral implications, which poses a seemingly impossible challenge: teaching machines moral sense, while humanity continues to grapple with it. To explore this challenge, we introduce Delphi, an experimental framework based on deep neural networks trained directly to reason about descriptive ethical judgments, e.g., "helping a friend" is generally good, while "helping a friend spread fake news" is not. Empirical results shed novel insights on the promises and limits of machine ethics; Delphi demonstrates strong generalization capabilities in the face of novel ethical situations, while off-the-shelf neural network models exhibit markedly poor judgment including unjust biases, confirming the need for explicitly teaching machines moral sense. Yet, Delphi is not perfect, exhibiting susceptibility to pervasive biases and inconsistencies. Despite that, we demonstrate positive use cases of imperfect Delphi, including using it as a component model within other imperfect AI systems. Importantly, we interpret the operationalization of Delphi in light of prominent ethical theories, which leads us to important future research questions.
Forward citations
Cited by 9 Pith papers
-
Internal Pluralism and the Limits of Pairwise Comparisons
Under internal pluralism, forced local pairwise comparisons erase inseparable priorities and distort conflicted answers, while allowing indecision reports can sharply reduce queries needed to learn preference weights.
-
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
MET-D self-distills theory-selected moral grounds into native-language reasoning, lifting macro-F1 by ~3.7–4.2 points on MCLASH and MMoralExceptQA while raising native-language chains by ~62 points.
-
Value Drifts: Tracing Value Alignment During LLM Post-Training
Value alignment in LLMs is set largely during supervised fine-tuning; standard preference-optimization datasets carry too little stance contrast to re-align it, but with engineered contrast algorithms differ (DPO ampl...
-
Justifications for Democratizing AI Alignment and Their Prospects
Neither purely democratic nor purely epistocratic AI alignment may be justified on its own, so hybrid frameworks combining expert input, participation, and anti-monopoly safeguards appear more promising.
-
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
LLMs align with human moral judgments only under high consensus, concentrate on a narrow set of moral values, and the profile-based prompting method's reported improvement is evaluated in-sample.
-
Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation Framework
Structured moral prompts, especially first-principles reasoning, improve LLM moral classification accuracy across 12 open models and four benchmarks, and reasoning distillation transfers these gains to a 3B model.
-
EtiCor++: Towards Understanding Etiquettical Bias in LLMs
A new English etiquette corpus and bias metrics show that LLMs over-prefer Western norms and under-predict etiquettes from low-resource regions.
-
The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas
A new benchmark of escalating moral dilemmas shows that LLMs shift their value priorities across steps and display aggregate non-transitive preference patterns.
-
Cognitive Decision Routing in Large Language Models: When to Think Fast, When to Think Slow
A four-feature routing rule is claimed to cut token use by 34% and improve accuracy, but the paper provides no code, no data, and only hand-wavy details.
Discussion (0). Sign in to comment.