Pith. sign in

REVIEW 5 cited by

Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.00243 v1 pith:RIUB5RAC submitted 2023-12-30 cs.LG

classification cs.LG
keywords lorarewardrlhfdiversehumanuncertaintyensemblesmodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Reinforcement learning from human feedback (RLHF) emerges as a promising paradigm for aligning large language models (LLMs). However, a notable challenge in RLHF is overoptimization, where beyond a certain threshold, the pursuit of higher rewards leads to a decline in human preferences. In this paper, we observe the weakness of KL regularization which is commonly employed in existing RLHF methods to address overoptimization. To mitigate this limitation, we scrutinize the RLHF objective in the offline dataset and propose uncertainty-penalized RLHF (UP-RLHF), which incorporates uncertainty regularization during RL-finetuning. To enhance the uncertainty quantification abilities for reward models, we first propose a diverse low-rank adaptation (LoRA) ensemble by maximizing the nuclear norm of LoRA matrix concatenations. Then we optimize policy models utilizing penalized rewards, determined by both rewards and uncertainties provided by the diverse reward LoRA ensembles. Our experimental results, based on two real human preference datasets, showcase the effectiveness of diverse reward LoRA ensembles in quantifying reward uncertainty. Additionally, uncertainty regularization in UP-RLHF proves to be pivotal in mitigating overoptimization, thereby contributing to the overall performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Internal Pluralism and the Limits of Pairwise Comparisons

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Under internal pluralism, forced local pairwise comparisons erase inseparable priorities and distort conflicted answers, while allowing indecision reports can sharply reduce queries needed to learn preference weights.

  2. The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking

    cs.LG 2025-01 reject novelty 6.0 of 10

    Increasing energy loss in an LLM's final layer during RLHF is linked to reward hacking, and penalizing that loss (EPPO) reduces hacking and improves RLHF quality.

  3. On the Robustness of Reward Models for Language Model Alignment

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Adding a batch-wise sum-to-zero penalty to Bradley-Terry reward modeling makes reward models more robust to unseen prompts and responses, according to experiments across multiple model families and benchmarks.

  4. Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation

    cs.LG 2026-02 conditional novelty 4.0 of 10

    A taxonomy-driven survey arguing that reward design is the central mechanism shaping reliable LLM reasoning, with maps of reward paradigms, reward-hacking failure modes, and benchmark pitfalls.

  5. Towards Reliable, Uncertainty-Aware Alignment

    cs.LG 2025-07 reject novelty 4.0 of 10

    Variance-aware RLHF adds a variance-weighted KL penalty to PPO and reduces reward variance and the risk of underperforming the reference policy in the paper's experiments.

Pith tools