Pith. sign in

REVIEW 3 cited by

Impact of Preference Noise on the Alignment Performance of Generative Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.09824 v1 pith:KPETSIRA submitted 2024-04-15 cs.CL

classification cs.CL
keywords alignmentnoisepreferenceimpactperformancemitigatedatageneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A key requirement in developing Generative Language Models (GLMs) is to have their values aligned with human values. Preference-based alignment is a widely used paradigm for this purpose, in which preferences over generation pairs are first elicited from human annotators or AI systems, and then fed into some alignment techniques, e.g., Direct Preference Optimization. However, a substantial percent (20 - 40%) of the preference pairs used in GLM alignment are noisy, and it remains unclear how the noise affects the alignment performance and how to mitigate its negative impact. In this paper, we propose a framework to inject desirable amounts and types of noise to the preferences, and systematically study the impact of preference noise on the alignment performance in two tasks (summarization and dialogue generation). We find that the alignment performance can be highly sensitive to the noise rates in the preference data: e.g., a 10 percentage points (pp) increase of the noise rate can lead to 30 pp drop in the alignment performance (in win rate). To mitigate the impact of noise, confidence-based data filtering shows significant benefit when certain types of noise are present. We hope our work can help the community better understand and mitigate the impact of preference noise in GLM alignment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Under 20–40% random preference-label flips, PACMR-DPO—a VNet-reweighted DPO with a prompt-augmentation-consistency meta-objective and central-difference LoRA meta-gradients—outperforms cDPO, IPO, rDPO, and Dr.DPO in j...

  2. Influence Functions for Preference Dataset Pruning

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Conjugate-gradient influence functions can mildly improve reward-model accuracy after pruning 10% of a preference dataset, but the gain is not statistically significant and gradient similarity better identifies helpfu...

  3. On Symmetric Losses for Robust Policy Optimization with Noisy Preferences

    cs.LG 2025-05 reject novelty 4.0 of 10

    Symmetric losses preserve action rankings under symmetric label noise, and the paper's claim that they also handle asymmetric noise is invalid.

Pith tools