Human values are claimed to have a rational instrumental structure that lets AI infer unseen values from known ones, framing this as the 'value generalization problem' in AI safety.
The Alignment Problem: Machine Learning and Human Values
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Machine Theory of Mind and the Structure of Human Values
Human values are claimed to have a rational instrumental structure that lets AI infer unseen values from known ones, framing this as the 'value generalization problem' in AI safety.