Pith. sign in

REVIEW 1 cited by

Concept Alignment as a Prerequisite for Value Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.20059 v1 pith:FISS26L3 submitted 2023-10-30 cs.AI

classification cs.AI
keywords alignmentconceptconceptsvaluevaluesalignhumansperson
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Value alignment is essential for building AI systems that can safely and reliably interact with people. However, what a person values -- and is even capable of valuing -- depends on the concepts that they are currently using to understand and evaluate what happens in the world. The dependence of values on concepts means that concept alignment is a prerequisite for value alignment -- agents need to align their representation of a situation with that of humans in order to successfully align their values. Here, we formally analyze the concept alignment problem in the inverse reinforcement learning setting, show how neglecting concept alignment can lead to systematic value mis-alignment, and describe an approach that helps minimize such failure modes by jointly reasoning about a person's concepts and values. Additionally, we report experimental results with human participants showing that humans reason about the concepts used by an agent when acting intentionally, in line with our joint reasoning model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. We Can't Understand AI Using our Existing Vocabulary

    cs.CL 2025-02 conditional novelty 4.0 of 10

    AI interpretability is reframed as building a shared human-machine language in which each new word is a learned token embedding trained by preference optimization.

Pith tools