Pith. sign in

REVIEW 5 cited by

On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.16455 v2 pith:Q5AANLY3 submitted 2024-05-26 stat.ML cs.LGstat.ME

classification stat.MLcs.LGstat.ME
keywords rlhfbiashumanpreferencepreferencesalgorithmicaligningapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Accurately aligning large language models (LLMs) with human preferences is crucial for informing fair, economically sound, and statistically efficient decision-making processes. However, we argue that the predominant approach for aligning LLMs with human preferences through a reward model -- reinforcement learning from human feedback (RLHF) -- suffers from an inherent algorithmic bias due to its Kullback--Leibler-based regularization in optimization. In extreme cases, this bias could lead to a phenomenon we term preference collapse, where minority preferences are virtually disregarded. To mitigate this algorithmic bias, we introduce preference matching (PM) RLHF, a novel approach that provably aligns LLMs with the preference distribution of the reward model under the Bradley--Terry--Luce/Plackett--Luce model. Central to our approach is a PM regularizer that takes the form of the negative logarithm of the LLM's policy probability distribution over responses, which helps the LLM balance response diversification and reward maximization. Notably, we obtain this regularizer by solving an ordinary differential equation that is necessary for the PM property. For practical implementation, we introduce a conditional variant of PM RLHF that is tailored to natural language generation. Finally, we empirically validate the effectiveness of conditional PM RLHF through experiments on the OPT and Llama-family models, demonstrating a 29% to 41% improvement in alignment with human preferences, as measured by a certain metric, compared to standard RLHF.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. "GenAI Defaults to Bias!" Gamify AI Literacy Through Reflections on Prompts

    cs.HC 2025-09 conditional novelty 5.0 of 10

    Playing ImaginAItion, a prompt-minimization party game, helped 30 adults recognize GenAI default biases and adjust their prompting strategies, according to pre-post survey coding.

  2. Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory

    stat.ML 2025-06 conditional novelty 5.0 of 10

    RLHF reward modeling satisfies pairwise majority and Condorcet consistency when each response pair is labeled once, because the maximum likelihood ranking then matches the Copeland rule.

  3. The Fair Game: Auditing & Debiasing AI Algorithms Over Time

    cs.AI 2025-08 unverdicted novelty 4.0 of 10

    Proposes 'Fair Game', a reinforcement-learning loop in which an auditor's bias criteria, updatable over time, steer a debiasing agent that adapts an ML model's predictions.

  4. Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning

    cs.LG 2025-08 reject novelty 3.0 of 10

    A proposed Counterfactual Trust Score aggregates drift, uncertainty, fairness violations, and counterfactual consistency into a single reward-model trust signal, evaluated only via a self-composed score on an unnamed ...

  5. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools