Pith. sign in

REVIEW 2 cited by

Training Value-Aligned Reinforcement Learning Agents Using a Normative Prior

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.09469 v1 pith:JBP2HCEX submitted 2021-04-19 cs.LG cs.AIcs.HC

classification cs.LGcs.AIcs.HC
keywords normativerewardtaskagentsbehaviorlearningvalue-alignedagent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As more machine learning agents interact with humans, it is increasingly a prospect that an agent trained to perform a task optimally, using only a measure of task performance as feedback, can violate societal norms for acceptable behavior or cause harm. Value alignment is a property of intelligent agents wherein they solely pursue non-harmful behaviors or human-beneficial goals. We introduce an approach to value-aligned reinforcement learning, in which we train an agent with two reward signals: a standard task performance reward, plus a normative behavior reward. The normative behavior reward is derived from a value-aligned prior model previously shown to classify text as normative or non-normative. We show how variations on a policy shaping technique can balance these two sources of reward and produce policies that are both effective and perceived as being more normative. We test our value-alignment technique on three interactive text-based worlds; each world is designed specifically to challenge agents with a task as well as provide opportunities to deviate from the task to engage in normative and/or altruistic behavior.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Odyssey of the Fittest: Can Agents Survive and Still Be Good?

    cs.AI 2025-02 reject novelty 6.0 of 10

    In an LLM-generated text survival game, a GPT-4o agent was reported to survive better and score more ethically than NEAT and SVI Bayesian agents, but the evaluation is circular because GPT-4o labels its own behavior.

  2. HAVA: Hybrid Approach to Value-Alignment through Reward Weighing for Reinforcement Learning

    cs.AI 2025-05 conditional novelty 5.0 of 10

    HAVA weights RL rewards by an agent reputation that falls when norms are violated, letting written safety rules and learned social norms be combined in one policy.

Pith tools