Pith. sign in

REVIEW 5 cited by

BERT has a Moral Compass: Improvements of ethical and moral values of machines

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.05238 v1 pith:VVZICNMY submitted 2019-12-11 cs.CL cs.AIcs.LGstat.ML

classification cs.CLcs.AIcs.LGstat.ML
keywords moralethicalbertmachinevalueskillmachinesactions
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Allowing machines to choose whether to kill humans would be devastating for world peace and security. But how do we equip machines with the ability to learn ethical or even moral choices? Jentzsch et al.(2019) showed that applying machine learning to human texts can extract deontological ethical reasoning about "right" and "wrong" conduct by calculating a moral bias score on a sentence level using sentence embeddings. The machine learned that it is objectionable to kill living beings, but it is fine to kill time; It is essential to eat, yet one might not eat dirt; it is important to spread information, yet one should not spread misinformation. However, the evaluated moral bias was restricted to simple actions -- one verb -- and a ranking of actions with surrounding context. Recently BERT ---and variants such as RoBERTa and SBERT--- has set a new state-of-the-art performance for a wide range of NLP tasks. But has BERT also a better moral compass? In this paper, we discuss and show that this is indeed the case. Thus, recent improvements of language representations also improve the representation of the underlying ethical and moral values of the machine. We argue that through an advanced semantic representation of text, BERT allows one to get better insights of moral and ethical values implicitly represented in text. This enables the Moral Choice Machine (MCM) to extract more accurate imprints of moral choices and ethical values.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Concept-ROT: Poisoning Concepts in Large Language Models with Model Editing

    cs.LG 2024-12 conditional novelty 8.0 of 10

    A single rank-one model edit can make a safety-tuned LLM jailbreak harmful questions only when the conversation is about the poisoned concept.

  2. Spectral Principal Paths: A Spectral Perspective on Linear Representation Formation in LLMs

    cs.CV 2025-06 reject novelty 6.0 of 10

    The paper argues and partially tests that linear concept representations in LLMs originate in the input space and propagate through the leading singular directions of activation difference matrices.

  3. Localizing Persona Representations in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.

  4. The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking

    cs.LG 2025-01 reject novelty 6.0 of 10

    Increasing energy loss in an LLM's final layer during RLHF is linked to reward hacking, and penalizing that loss (EPPO) reduces hacking and improves RLHF quality.

  5. Right vs. Right: Can LLMs Make Tough Choices?

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Across 1,730 LLM-generated ethical dilemmas, models reliably prefer truth over loyalty, community over individual, and long-term over short-term benefits, while showing sensitivity to question phrasing.

Pith tools