Pith. sign in

REVIEW 1 cited by

Towards Socially and Morally Aware RL agent: Reward Design With LLM

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.12459 v2 pith:HCPCZAPU submitted 2024-01-23 cs.AI

classification cs.AI
keywords rewardexplorationhumanlanguageworkagentdesigneffects
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When we design and deploy an Reinforcement Learning (RL) agent, reward functions motivates agents to achieve an objective. An incorrect or incomplete specification of the objective can result in behavior that does not align with human values - failing to adhere with social and moral norms that are ambiguous and context dependent, and cause undesired outcomes such as negative side effects and exploration that is unsafe. Previous work have manually defined reward functions to avoid negative side effects, use human oversight for safe exploration, or use foundation models as planning tools. This work studies the ability of leveraging Large Language Models (LLM)' understanding of morality and social norms on safe exploration augmented RL methods. This work evaluates language model's result against human feedbacks and demonstrates language model's capability as direct reward signals.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RoboMoRe: LLM-based Robot Co-design via Joint Optimization of Morphology and Reward

    cs.RO 2025-05 reject novelty 6.0 of 10

    An LLM-driven framework that jointly proposes robot morphologies and reward functions, using diversity reflection and alternating refinement, claims large efficiency gains over baselines across eight locomotion tasks.

Pith tools