Pith. sign in

REVIEW 3 cited by

Does Moral Code Have a Moral Code? Probing Delphi's Moral Philosophy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.12771 v1 pith:T7AQDI5R submitted 2022-05-25 cs.CY cs.CL

classification cs.CYcs.CL
keywords moraldelphimodelcodehumanlearnmodelsprinciples
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In an effort to guarantee that machine learning model outputs conform with human moral values, recent work has begun exploring the possibility of explicitly training models to learn the difference between right and wrong. This is typically done in a bottom-up fashion, by exposing the model to different scenarios, annotated with human moral judgements. One question, however, is whether the trained models actually learn any consistent, higher-level ethical principles from these datasets -- and if so, what? Here, we probe the Allen AI Delphi model with a set of standardized morality questionnaires, and find that, despite some inconsistencies, Delphi tends to mirror the moral principles associated with the demographic groups involved in the annotation process. We question whether this is desirable and discuss how we might move forward with this knowledge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  2. Normative Evaluation of Large Language Models with Everyday Moral Dilemmas

    cs.AI 2025-01 conditional novelty 5.0 of 10

    Seven LLMs give different moral verdicts on AITA dilemmas, differ from Redditors, and only in an ensemble approximate human consensus.

  3. Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values

    cs.AI 2025-01 conditional novelty 5.0 of 10

    Value Compass Benchmarks is a live, self-evolving platform that scores 33 LLMs across 27 value dimensions from four value systems, aiming to reveal true behavioral alignment with human values.

Pith tools