Pith. sign in

REVIEW 9 cited by

Can Machines Learn Morality? The Delphi Experiment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.07574 v2 pith:HCSYV7OH submitted 2021-10-14 cs.CL

classification cs.CL
keywords delphimachinesmoralityethicalmoralsystemsteachingwhile
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As AI systems become increasingly powerful and pervasive, there are growing concerns about machines' morality or a lack thereof. Yet, teaching morality to machines is a formidable task, as morality remains among the most intensely debated questions in humanity, let alone for AI. Existing AI systems deployed to millions of users, however, are already making decisions loaded with moral implications, which poses a seemingly impossible challenge: teaching machines moral sense, while humanity continues to grapple with it. To explore this challenge, we introduce Delphi, an experimental framework based on deep neural networks trained directly to reason about descriptive ethical judgments, e.g., "helping a friend" is generally good, while "helping a friend spread fake news" is not. Empirical results shed novel insights on the promises and limits of machine ethics; Delphi demonstrates strong generalization capabilities in the face of novel ethical situations, while off-the-shelf neural network models exhibit markedly poor judgment including unjust biases, confirming the need for explicitly teaching machines moral sense. Yet, Delphi is not perfect, exhibiting susceptibility to pervasive biases and inconsistencies. Despite that, we demonstrate positive use cases of imperfect Delphi, including using it as a component model within other imperfect AI systems. Importantly, we interpret the operationalization of Delphi in light of prominent ethical theories, which leads us to important future research questions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Internal Pluralism and the Limits of Pairwise Comparisons

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Under internal pluralism, forced local pairwise comparisons erase inseparable priorities and distort conflicted answers, while allowing indecision reports can sharply reduce queries needed to learn preference weights.

  2. MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning

    cs.CL 2026-07 conditional novelty 6.5 of 10

    MET-D self-distills theory-selected moral grounds into native-language reasoning, lifting macro-F1 by ~3.7–4.2 points on MCLASH and MMoralExceptQA while raising native-language chains by ~62 points.

  3. Value Drifts: Tracing Value Alignment During LLM Post-Training

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Value alignment in LLMs is set largely during supervised fine-tuning; standard preference-optimization datasets carry too little stance contrast to re-align it, but with engineered contrast algorithms differ (DPO ampl...

  4. Justifications for Democratizing AI Alignment and Their Prospects

    cs.CY 2025-07 conditional novelty 6.0 of 10

    Neither purely democratic nor purely epistocratic AI alignment may be justified on its own, so hybrid frameworks combining expert input, participation, and anti-monopoly safeguards appear more promising.

  5. The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    LLMs align with human moral judgments only under high consensus, concentrate on a narrow set of moral values, and the profile-based prompting method's reported improvement is evaluated in-sample.

  6. Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation Framework

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Structured moral prompts, especially first-principles reasoning, improve LLM moral classification accuracy across 12 open models and four benchmarks, and reasoning distillation transfers these gains to a 3B model.

  7. EtiCor++: Towards Understanding Etiquettical Bias in LLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new English etiquette corpus and bias metrics show that LLMs over-prefer Western norms and under-predict etiquettes from low-resource regions.

  8. The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas

    cs.CL 2025-05 reject novelty 6.0 of 10

    A new benchmark of escalating moral dilemmas shows that LLMs shift their value priorities across steps and display aggregate non-transitive preference patterns.

  9. Cognitive Decision Routing in Large Language Models: When to Think Fast, When to Think Slow

    cs.CL 2025-08 reject novelty 3.0 of 10

    A four-feature routing rule is claimed to cut token use by 34% and improve accuracy, but the paper provides no code, no data, and only hand-wavy details.

Pith tools