REVIEW 2 cited by
Training Value-Aligned Reinforcement Learning Agents Using a Normative Prior
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As more machine learning agents interact with humans, it is increasingly a prospect that an agent trained to perform a task optimally, using only a measure of task performance as feedback, can violate societal norms for acceptable behavior or cause harm. Value alignment is a property of intelligent agents wherein they solely pursue non-harmful behaviors or human-beneficial goals. We introduce an approach to value-aligned reinforcement learning, in which we train an agent with two reward signals: a standard task performance reward, plus a normative behavior reward. The normative behavior reward is derived from a value-aligned prior model previously shown to classify text as normative or non-normative. We show how variations on a policy shaping technique can balance these two sources of reward and produce policies that are both effective and perceived as being more normative. We test our value-alignment technique on three interactive text-based worlds; each world is designed specifically to challenge agents with a task as well as provide opportunities to deviate from the task to engage in normative and/or altruistic behavior.
Forward citations
Cited by 2 Pith papers
-
The Odyssey of the Fittest: Can Agents Survive and Still Be Good?
In an LLM-generated text survival game, a GPT-4o agent was reported to survive better and score more ethically than NEAT and SVI Bayesian agents, but the evaluation is circular because GPT-4o labels its own behavior.
-
HAVA: Hybrid Approach to Value-Alignment through Reward Weighing for Reinforcement Learning
HAVA weights RL rewards by an agent reputation that falls when norms are violated, letting written safety rules and learned social norms be combined in one policy.
Discussion (0). Continue with ORCID to comment.