Pith. sign in

REVIEW 6 cited by

Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.10075 v1 pith:P53UNYGI submitted 2024-08-19 cs.LG cs.AIcs.CLcs.RO

classification cs.LGcs.AIcs.CLcs.RO
keywords learningpreferencesrewardhumanrlhfdiverselatentuser-specific
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement Learning from Human Feedback (RLHF) is a powerful paradigm for aligning foundation models to human values and preferences. However, current RLHF techniques cannot account for the naturally occurring differences in individual human preferences across a diverse population. When these differences arise, traditional RLHF frameworks simply average over them, leading to inaccurate rewards and poor performance for individual subgroups. To address the need for pluralistic alignment, we develop a class of multimodal RLHF methods. Our proposed techniques are based on a latent variable formulation - inferring a novel user-specific latent and learning reward models and policies conditioned on this latent without additional user-specific data. While conceptually simple, we show that in practice, this reward modeling requires careful algorithmic considerations around model architecture and reward scaling. To empirically validate our proposed technique, we first show that it can provide a way to combat underspecification in simulated control problems, inferring and optimizing user-specific reward functions. Next, we conduct experiments on pluralistic language datasets representing diverse user preferences and demonstrate improved reward function accuracy. We additionally show the benefits of this probabilistic framework in terms of measuring uncertainty, and actively learning user preferences. This work enables learning from diverse populations of users with divergent preferences, an important challenge that naturally occurs in problems from robot learning to foundation model alignment.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Personalizing Large Language Model Agents with Small Policy Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A factorized Bayesian Thompson-sampling layer outside a frozen agent learns per-user execution preferences from selected-action scalar feedback, with a Õ(d^{3/2}√n) regret bound against the best feasible action.

  2. SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences

    cs.LG 2025-09 reject novelty 6.0 of 10

    SharedRep-RLHF learns a shared preference representation across groups to improve worst-case reward estimates for minority annotators, but the theoretical guarantees are undermined by proof errors.

  3. Affordances of Sketched Notations for Multimodal UI Design and Development Tools

    cs.HC 2025-08 conditional novelty 6.0 of 10

    People's free-form UI sketches are highly varied and ambiguous in isolation, but interpretable in context, so flexible sketch tools need context-aware AI rather than fixed symbol recognition.

  4. Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A 7B model trained with synthetic reasoning demonstrations plus reinforcement learning infers explicit user preference descriptions from behavioral signals, improving personalized response judging and generation.

  5. CURP: Codebook-based Continuous User Representation for Personalized Generation with LLMs

    cs.CL 2026-01 conditional novelty 5.0 of 10

    CURP represents users as sparse combinations of discrete prototype codebook embeddings and uses them as frozen-LLM prefixes, outperforming personalization baselines on four text-generation tasks with about 20M trainab...

  6. Configurable Preference Tuning with Rubric-Guided Synthetic Data

    cs.CL 2025-06 conditional novelty 5.0 of 10

    CPT fine-tunes LLMs with DPO on rubric-guided synthetic preferences so that a system prompt can reconfigure output style at inference, with in-distribution accuracy gains over baselines.

Pith tools