Pith. sign in

REVIEW 1 cited by

When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.17747 v5 pith:2HXWW7H2 submitted 2024-02-27 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords humanfeedbackpartialrlhfcaseschallengesfunctionlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Past analyses of reinforcement learning from human feedback (RLHF) assume that the human evaluators fully observe the environment. What happens when human feedback is based only on partial observations? We formally define two failure cases: deceptive inflation and overjustification. Modeling the human as Boltzmann-rational w.r.t. a belief over trajectories, we prove conditions under which RLHF is guaranteed to result in policies that deceptively inflate their performance, overjustify their behavior to make an impression, or both. Under the new assumption that the human's partial observability is known and accounted for, we then analyze how much information the feedback process provides about the return function. We show that sometimes, the human's feedback determines the return function uniquely up to an additive constant, but in other realistic cases, there is irreducible ambiguity. We propose exploratory research directions to help tackle these challenges, experimentally validate both the theoretical concerns and potential mitigations, and caution against blindly applying RLHF in partially observable settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Observation Interference in Partially Observable Assistance Games

    cs.AI 2024-12 conditional novelty 7.0 of 10

    In partially observable assistance games, a goal-aligned assistant sometimes must interfere with the human's observations at the level of single actions, but never at the level of its complete policy.

Pith tools