Pith. sign in

REVIEW 1 cited by

RLHF and IIA: Perverse Incentives

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.01057 v3 pith:BP6DIM6C submitted 2023-12-02 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords algorithmsincentiveslearningperverserlhfalternativesassumebecause
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing algorithms for reinforcement learning from human feedback (RLHF) can incentivize responses at odds with preferences because they are based on models that assume independence of irrelevant alternatives (IIA). The perverse incentives induced by IIA hinder innovations on query formats and learning algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Jackpot! Alignment as a Maximal Lottery

    cs.AI 2025-01 conditional novelty 5.0 of 10

    Nash Learning from Human Feedback is shown to approximate the maximal lottery voting rule, which the authors argue better reflects majority preferences than Borda-like RLHF.

Pith tools