REVIEW 1 cited by
RLHF and IIA: Perverse Incentives
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Existing algorithms for reinforcement learning from human feedback (RLHF) can incentivize responses at odds with preferences because they are based on models that assume independence of irrelevant alternatives (IIA). The perverse incentives induced by IIA hinder innovations on query formats and learning algorithms.
Forward citations
Cited by 1 Pith paper
-
Jackpot! Alignment as a Maximal Lottery
Nash Learning from Human Feedback is shown to approximate the maximal lottery voting rule, which the authors argue better reflects majority preferences than Borda-like RLHF.
Discussion (0). Continue with ORCID to comment.