REVIEW 5 cited by
ROPO: Robust Preference Optimization for Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Preference alignment is pivotal for empowering large language models (LLMs) to generate helpful and harmless responses. However, the performance of preference alignment is highly sensitive to the prevalent noise in the preference data. Recent efforts for this problem either marginally alleviate the impact of noise without the ability to actually reduce its presence, or rely on costly teacher LLMs prone to reward misgeneralization. To address these challenges, we propose the RObust Preference Optimization (ROPO) framework, an iterative alignment approach that integrates noise-tolerance and filtering of noisy samples without the aid of external models. Specifically, ROPO iteratively solves a constrained optimization problem, where we dynamically assign a quality-aware weight for each sample and constrain the sum of the weights to the number of samples we intend to retain. For noise-tolerant training and effective noise identification, we derive a robust loss by suppressing the gradients of samples with high uncertainty. We demonstrate both empirically and theoretically that the derived loss is critical for distinguishing noisy samples from clean ones. Furthermore, inspired by our derived loss, we propose a robustness-guided rejection sampling technique to compensate for the potential important information in discarded queries. Experiments on three widely-used datasets with Mistral-7B and Llama-2-7B demonstrate that ROPO significantly outperforms existing preference alignment methods, with its superiority growing as the noise rate increases.
Forward citations
Cited by 5 Pith papers
-
Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model
Preference-based LLM alignment under an unknown reward-preference link becomes a single-index model; three new algorithms converge to the optimal divergence-constrained policy without knowing the link.
-
Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels
Under 20–40% random preference-label flips, PACMR-DPO—a VNet-reweighted DPO with a prompt-augmentation-consistency meta-objective and central-difference LoRA meta-gradients—outperforms cDPO, IPO, rDPO, and Dr.DPO in j...
-
ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement Learning
ADG uses an ambient DDPM to flag corrupted RL transitions, trains a standard DDPM only on the clean subset, then refines the flagged transitions to produce a recovered dataset that improves offline RL policies.
-
Normalized Rewards for Preference Optimization
A regularization term that conserves the combined length-normalized probability of chosen and rejected responses reduces likelihood displacement in DPO/SimPO, improves AlpacaEval and benchmark outcomes, and acts prima...
-
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
Symmetric losses preserve action rankings under symmetric label noise, and the paper's claim that they also handle asymmetric noise is invalid.
Discussion (0). Continue with ORCID to comment.