REVIEW 18 cited by
MaxMin-RLHF: Alignment with Diverse Human Preferences
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Reinforcement Learning from Human Feedback (RLHF) aligns language models to human preferences by employing a singular reward model derived from preference data. However, such an approach overlooks the rich diversity of human preferences inherent in data collected from multiple users. In this work, we first derive an impossibility result of alignment with single reward RLHF, thereby highlighting its insufficiency in representing diverse human preferences. To provide an equitable solution to the problem, we learn a mixture of preference distributions via an expectation-maximization algorithm and propose a MaxMin alignment objective for policy learning inspired by the Egalitarian principle in social choice theory to better represent diverse human preferences. We elucidate the connection of our proposed approach to distributionally robust optimization and general utility RL, thereby highlighting the generality and robustness of our proposed solution. We present comprehensive experimental results on small-scale (GPT-2) and large-scale language models (with Tulu2-7B) and show the efficacy of the proposed approach in the presence of diversity among human preferences. Our algorithm achieves an average improvement of more than 16% in win-rates over conventional RLHF algorithms and improves the win-rate (accuracy) for minority groups by over 33% without compromising the performance of majority groups, showcasing the robustness and fairness of our approach. We remark that our findings in this work are not only limited to language models but also extend to reinforcement learning in general.
Forward citations
Cited by 18 Pith papers
-
Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
NLHF achieves the minimax-optimal worst-case average-utility distortion (1/2+o(1))β, while RLHF and DPO can suffer distortion up to e^{Ω(β)} or unbounded under certain comparison sampling.
-
Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment
Latent DPO trains only a small preference encoder per user, reducing LLM personalization training time by 80-90% with comparable alignment quality.
-
Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration
CALM uses bilevel optimization to tune per-vocabulary temperature-like logit adjustments during LLM fine-tuning, and reports improved out-of-domain calibration for aligned language models.
-
Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis
Engineered personas plus IS/OOS evidence retrieval make cheap multi-model panels produce tested claim maps and expose RLHF-induced blind spots, including asymmetric AI-risk challenge.
-
SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
SharedRep-RLHF learns a shared preference representation across groups to improve worst-case reward estimates for minority annotators, but the theoretical guarantees are undermined by proof errors.
-
Fundamental Limits of Game-Theoretic LLM Alignment: Smith Consistency and Preference Matching
For game-theoretic LLM alignment, Condorcet and Smith consistency hold for broad payoff classes, but preference matching is impossible for smooth, learnable payoff mappings.
-
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning
Reinforcement learning on questions extracted from CRISPR expert forums improves LLM accuracy on a new benchmark (Genome-Bench) by over 15 percentage points.
-
Pairwise Calibrated Rewards for Pluralistic Alignment
A small ensemble of reward functions can be trained to match pairwise human preference frequencies, offering a practical route to pluralistic AI alignment.
-
Clone-Robust AI Alignment
A Voronoi-weighted maximum likelihood estimator for RLHF is robust to adding approximate clone responses, unlike the standard regularized MLE.
-
When Can Proxies Improve the Sample Complexity of Preference Learning?
Under four structural conditions linking proxy and true preference policies, the true policy is a low-dimensional adapter of the proxy policy, reducing the number of true preference samples needed.
-
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
The paper generalizes Bradley-Terry reward modeling to ordinal feedback labels and proves that, under a marginal unbiasedness assumption, finer-grained labels reduce Rademacher complexity and can improve reward learning.
-
Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory
RLHF reward modeling satisfies pairwise majority and Condorcet consistency when each response pair is labeled once, because the maximum likelihood ranking then matches the Copeland rule.
-
Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data?
Preference signals in LLM alignment are concentrated in early response tokens, so models trained on data truncated to the first half perform as well as or better than those trained on full responses.
-
Beyond the Binary: Capturing Diverse Preferences With Reward Regularization
Reward models trained with a margin derived from synthetic LLM judgments better match aggregate human preferences than standard binary-trained reward models, mainly on subjective prompts.
-
Context Engineering: A Practitioner Methodology for Structured Human-AI Collaboration
Structured five-role context packages and a four-phase pipeline were associated with cutting average AI task iterations from 3.8 to 2.0 and raising first-pass acceptance from 32% to 55% in an observational single-oper...
-
AI-Augmented LLMs Achieve Therapist-Level Responses in Motivational Interviewing
A custom prompt built from machine-learning-identified therapy behavior features improved GPT-4's motivational interviewing quality scores, though the model remained slightly below human therapists on the paper's own metric.
-
Embedding-to-Prefix: Parameter-Efficient Personalization for Pre-Trained Large Language Models
E2P projects pre-computed user embeddings into a single soft prefix token for frozen LLMs, reporting gains on four personalization tasks, though its reproduction scripts write zero embeddings.
-
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
A tutorial reviewing LLM alignment through the lens of inverse reinforcement learning, arguing that neural reward models learned from human data are central to post-training.
Discussion (0). Continue with ORCID to comment.