HAVA weights RL rewards by an agent reputation that falls when norms are violated, letting written safety rules and learned social norms be combined in one policy.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
HAVA: Hybrid Approach to Value-Alignment through Reward Weighing for Reinforcement Learning
HAVA weights RL rewards by an agent reputation that falls when norms are violated, letting written safety rules and learned social norms be combined in one policy.