Human values are claimed to have a rational instrumental structure that lets AI infer unseen values from known ones, framing this as the 'value generalization problem' in AI safety.
Baker, Rebecca Saxe, and Joshua B
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Machine Theory of Mind and the Structure of Human Values
Human values are claimed to have a rational instrumental structure that lets AI infer unseen values from known ones, framing this as the 'value generalization problem' in AI safety.