REVIEW 17 cited by
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Reinforcement Learning from Human Feedback (RLHF) is a powerful paradigm for aligning foundation models to human values and preferences. However, current RLHF techniques cannot account for the naturally occurring differences in individual human preferences across a diverse population. When these differences arise, traditional RLHF frameworks simply average over them, leading to inaccurate rewards and poor performance for individual subgroups. To address the need for pluralistic alignment, we develop a class of multimodal RLHF methods. Our proposed techniques are based on a latent variable formulation - inferring a novel user-specific latent and learning reward models and policies conditioned on this latent without additional user-specific data. While conceptually simple, we show that in practice, this reward modeling requires careful algorithmic considerations around model architecture and reward scaling. To empirically validate our proposed technique, we first show that it can provide a way to combat underspecification in simulated control problems, inferring and optimizing user-specific reward functions. Next, we conduct experiments on pluralistic language datasets representing diverse user preferences and demonstrate improved reward function accuracy. We additionally show the benefits of this probabilistic framework in terms of measuring uncertainty, and actively learning user preferences. This work enables learning from diverse populations of users with divergent preferences, an important challenge that naturally occurs in problems from robot learning to foundation model alignment.
Forward citations
Cited by 17 Pith papers
-
Personalizing Large Language Model Agents with Small Policy Models
A factorized Bayesian Thompson-sampling layer outside a frozen agent learns per-user execution preferences from selected-action scalar feedback, with a Õ(d^{3/2}√n) regret bound against the best feasible action.
-
SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
SharedRep-RLHF learns a shared preference representation across groups to improve worst-case reward estimates for minority annotators, but the theoretical guarantees are undermined by proof errors.
-
Affordances of Sketched Notations for Multimodal UI Design and Development Tools
People's free-form UI sketches are highly varied and ambiguous in isolation, but interpretable in context, so flexible sketch tools need context-aware AI rather than fixed symbol recognition.
-
Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals
A 7B model trained with synthetic reasoning demonstrations plus reinforcement learning infers explicit user preference descriptions from behavioral signals, improving personalized response judging and generation.
-
Fragments to Facts: Partial-Information Fragment Inference from LLMs
Fine-tuned LLMs leak private fragment-level information to adversaries holding only a few unordered public fragments, as shown by two probe attacks (LR-Attack and PRISM) on medical and legal summarization tasks.
-
Pairwise Calibrated Rewards for Pluralistic Alignment
A small ensemble of reward functions can be trained to match pairwise human preference frequencies, offering a practical route to pluralistic AI alignment.
-
Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes
LPC adds a discrete latent code layer to DPO-family alignment objectives, improving average preference accuracy and downstream scores across three base models.
-
Cross-environment Cooperation Enables Zero-shot Multi-agent Coordination
Training a self-play agent across many procedurally generated cooperative tasks yields better zero-shot coordination with novel partners and novel layouts than training on one task with many partners.
-
CTR-Driven Advertising Image Generation with Multimodal Large Language Models
A CTR-driven advertising image generation pipeline that pre-trains an MLLM prompt model, trains an MLLM pairwise reward model on real click data, and fine-tunes with DPO plus a product-centric preference loss.
-
Clone-Robust AI Alignment
A Voronoi-weighted maximum likelihood estimator for RLHF is robust to adding approximate clone responses, unlike the standard regularized MLE.
-
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
The paper introduces NP-BTL and NP-DPO, neural-process-based models that infer a user's preferences from a small number of preference pairs and condition LLM policies on those preferences at inference time.
-
On the Way to LLM Personalization: Learning to Remember User Conversations
Finetuning a LoRA adapter on self-generated question-answer pairs lets Llama 3 8B recall conversation topics with 81.5% accuracy, close to RAG at 83.5% but without retrieval.
-
CURP: Codebook-based Continuous User Representation for Personalized Generation with LLMs
CURP represents users as sparse combinations of discrete prototype codebook embeddings and uses them as frozen-LLM prefixes, outperforming personalization baselines on four text-generation tasks with about 20M trainab...
-
Configurable Preference Tuning with Rubric-Guided Synthetic Data
CPT fine-tunes LLMs with DPO on rubric-guided synthetic preferences so that a system prompt can reconfigure output style at inference, with in-distribution accuracy gains over baselines.
-
Personalized Preference Fine-tuning of Diffusion Models
PPD fine-tunes a single diffusion model to follow per-user preferences by conditioning on VLM-extracted embeddings, reporting 76-81% win rates over Stable Cascade with four examples per user.
-
Beyond the Binary: Capturing Diverse Preferences With Reward Regularization
Reward models trained with a margin derived from synthetic LLM judgments better match aggregate human preferences than standard binary-trained reward models, mainly on subjective prompts.
-
LoRe: Personalizing LLMs via Low-Rank Reward Modeling
LoRe learns a shared low-rank reward basis plus per-user simplex weights and reports higher preference-prediction accuracy than personalized and monolithic baselines on three datasets.
Discussion (0). Continue with ORCID to comment.