REVIEW 27 cited by
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Reinforcement learning from human feedback (RLHF) has emerged as the primary method for aligning large language models (LLMs) with human preferences. The RLHF process typically starts by training a reward model (RM) using human preference data. Conventional RMs are trained on pairwise responses to the same user request, with relative ratings indicating which response humans prefer. The trained RM serves as a proxy for human preferences. However, due to the black-box nature of RMs, their outputs lack interpretability, as humans cannot intuitively understand why an RM thinks a response is good or not. As RMs act as human preference proxies, we believe they should be human-interpretable to ensure that their internal decision processes are consistent with human preferences and to prevent reward hacking in LLM alignment. To build RMs with interpretable preferences, we propose a two-stage approach: i) train an Absolute-Rating Multi-Objective Reward Model (ArmoRM) with multi-dimensional absolute-rating data, each dimension corresponding to a human-interpretable objective (e.g., honesty, verbosity, safety); ii) employ a Mixture-of-Experts (MoE) strategy with a gating network that automatically selects the most suitable reward objectives based on the context. We efficiently trained an ArmoRM with Llama-3 8B and a gating network consisting of a shallow MLP on top of the ArmoRM. Our trained model, ArmoRM-Llama3-8B, obtains state-of-the-art performance on RewardBench, a benchmark evaluating RMs for language modeling. Notably, the performance of our model surpasses the LLM-as-a-judge method with GPT-4 judges by a margin, and approaches the performance of the much larger Nemotron-4 340B reward model.
Forward citations
Cited by 27 Pith papers
-
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges
Per-bias selection of a cross-family LLM auditor lifts biased-judgment accuracy from 0.805/0.824 baselines to 0.884.
-
OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning
OpsLLM's pipeline (HITL data curation, SFT, GRPO RL with a domain process reward model) improves LLM accuracy on software-operations QA and RCA, especially on in-distribution root-cause-analysis tasks.
-
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
MAHALO aligns LLMs to multiple objectives in one model via per-objective action heads and PRM-guided decoding, improving math, value, and tutoring metrics jointly.
-
Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation
ICON2 uses representation-space steering to generate preference data from the model itself, improving alignment benchmarks and cutting cost.
-
ReaLM: Reflection-Enhanced Autonomous Reasoning with Small Language Models
ReaLM trains small language models to learn from both right and wrong reasoning chains, then fades the chains out so the model reasons independently, improving benchmark accuracy.
-
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
BMMR provides a 110k-question bilingual, multimodal, college-level dataset across 300 subjects where state-of-the-art models score at most about 50%.
-
BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
A routing system that chooses both the model and the number of samples per query to meet a quality threshold, yielding up to 60% cost savings.
-
Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding
A small aligned model drafts the start of an LLM response, then a large base model takes over via a confidence-based switch, improving preference alignment without fine-tuning the large model.
-
Accelerating RLHF Training with Reward Variance Increase
A new reward reshaping method provably increases reward variance for GRPO-based RLHF training, with an O(n log n) global optimization algorithm and preliminary speedups in experiments.
-
Fair-PP: A Synthetic Dataset for Aligning LLM with Personalized Preferences of Social Equity
Fair-PP contributes a synthetic persona-anchored preference dataset for social equity and a reweighted DPO/SFT alignment method that outperforms baselines on LLM-similarity tests.
-
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models
MedGUIDE tests whether LLMs follow structured NCCN cancer-care decision trees and finds that even medical LLMs often lag general models on this task.
-
R.I.P.: Better Models by Survival of the Fittest Prompts
RIP filters preference-optimization training data by keeping prompts whose rejected responses are high quality and whose chosen/rejected reward gap is small, yielding consistent benchmark gains over unfiltered data.
-
Data-adaptive Safety Rules for Training Reward Models
Selecting the five safety rules with largest response discrepancy maximizes mutual information under stated assumptions, and a reward model trained with this adaptive labeling achieves top RewardBench safety.
-
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
IXC-2.5-Reward is an open-source multimodal reward model that achieves 70.0% macro accuracy on VL-RewardBench and improves LVLM chat via PPO.
-
AlphaPO: Reward Shape Matters for LLM Alignment
AlphaPO replaces SimPO's log reward with a length-normalized alpha-divergence reward; a slightly positive alpha improves AlpacaEval 2 length-controlled win rates by 7-10 percent on two instruct models.
-
Boosting LLM via Learning from Data Iteratively and Selectively
IterIT iteratively re-scores instruction samples during fine-tuning and greedily selects a small, diverse, high-complexity subset each epoch.
-
T-REG: Preference Optimization with Token-Level Reward Regularization
T-REG adds contrastive-prompt self-generated token-level rewards as a weighted regularizer to DPO and SimPO, improving Alpaca Eval 2 length-controlled win rate by up to 3.8% and Arena-Hard win rate by up to 4.4% over ...
-
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
The paper generalizes Bradley-Terry reward modeling to ordinal feedback labels and proves that, under a marginal unbiasedness assumption, finer-grained labels reduce Rademacher complexity and can improve reward learning.
-
Adaptive Decoding via Latent Preference Optimization
Adaptive Decoding learns to pick a discrete sampling temperature at token or sequence level via a DPO-style loss over latent temperature choices, and beats fixed temperatures on average across three task families.
-
MOSLIM:Align with diverse preferences in prompts through reward classification
A prompt-controlled multi-objective alignment method using a multi-head classification reward model and a z-score reward mapping, claimed to work with off-the-shelf models.
-
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
MTSA pairs a thought-guided red-team attacker with future-reward multi-turn reinforcement learning to make LLMs more robust against multi-round jailbreaks.
-
CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis
A small CoT-based synthesizer model trained on candidate-response analysis improves LLM reasoning accuracy, including cases where all sampled candidate answers are incorrect.
-
Interpreting Language Reward Models via Contrastive Explanations
Reward model preferences can be explained by generating counterfactual and semifactual answer variations along 15 hand-picked evaluation attributes and measuring which attribute changes flip the model's preference.
-
T-POP: Test-Time Personalization with Online Preference Feedback
T-POP uses dueling-bandit token selection to learn a reward function online from pairwise user feedback, enabling test-time personalization of a frozen LLM without fine-tuning.
-
DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning
A reward model that multiplies stepwise correctness and potential scores improves best-of-N verification accuracy for math reasoning.
-
Atla Selene Mini: A General Purpose Evaluation Model
The paper presents Selene Mini, an 8B open-weights judge model that reports state-of-the-art average scores across 11 LLM evaluation benchmarks, with gains on medical and financial expert agreement.
-
Mixture of Experts (MoE): A Big Data Perspective
A survey of MoE methods for big data that catalogs architectures, use cases, and open challenges without adding new results.
Discussion (0). Continue with ORCID to comment.