Pith. sign in

REVIEW 2 cited by

Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.06874 v3 pith:Z4IJ5RAO submitted 2024-06-11 cs.AI cs.HCcs.RO

Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment

classification cs.AI cs.HCcs.RO
keywords alignmenthumanlearningpolicypreferencerewardalgorithmsapproach
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Aligning human preference and value is an important requirement for building contemporary foundation models and embodied AI. However, popular approaches such as reinforcement learning with human feedback (RLHF) break down the task into successive stages, such as supervised fine-tuning (SFT), reward modeling (RM), and reinforcement learning (RL), each performing one specific learning task. Such a sequential approach results in serious issues such as significant under-utilization of data and distribution mismatch between the learned reward model and generated policy, which eventually lead to poor alignment performance. We develop a single stage approach named Alignment with Integrated Human Feedback (AIHF), capable of integrating both human preference and demonstration to train reward models and the policy. The proposed approach admits a suite of efficient algorithms, which can easily reduce to, and leverage, popular alignment algorithms such as RLHF and Directly Policy Optimization (DPO), and only requires minor changes to the existing alignment pipelines. We demonstrate the efficiency of the proposed solutions with extensive experiments involving alignment problems in LLMs and robotic control problems in MuJoCo. We observe that the proposed solutions outperform the existing alignment algorithms such as RLHF and DPO by large margins, especially when the amount of high-quality preference data is relatively limited.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. On the Nature of Regularity Assumptions in Bilevel Optimization with Constrained Lower-level Problem

    math.OC 2026-05 conditional novelty 8.0

    Requiring LICQ/SCS/SOSC everywhere in bilevel optimization is non-prevalent and rigid, while holding almost everywhere is prevalent, but the distinction introduces fundamental difficulties.

  2. Penalty-Based First-Order Methods for Bilevel Optimization with Minimax and Constrained Lower-Level Problems

    math.OC 2026-05 unverdicted novelty 6.0

    Penalty-based first-order methods find ε-KKT points in bilevel minimax problems with Õ(ε^{-4}) deterministic and Õ(ε^{-9}) stochastic oracle complexity, improving prior bounds for constrained lower-level cases via Lag...