Pith. sign in

REVIEW 2 cited by

Gradient Imbalance in Direct Preference Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.20847 v1 pith:7VGEXAL5 submitted 2025-02-28 cs.LG

classification cs.LG
keywords gradientimbalanceoptimizationbalanced-dpodemonstratedirectlearningperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Direct Preference Optimization (DPO) has been proposed as a promising alternative to Proximal Policy Optimization (PPO) based Reinforcement Learning with Human Feedback (RLHF). However, empirical evaluations consistently reveal suboptimal performance in DPO compared to common RLHF pipelines. In this work, we conduct a systematic analysis of DPO's training dynamics and identify gradient imbalance as a critical limitation. We demonstrate theoretically and empirically that this imbalance perturbs optimization trajectories, destabilizes learning, and induces suboptimal convergence. To address this issue, we propose Balanced-DPO, a simple yet effective modification to the DPO objective that introduces a computationally efficient gradient reweighting mechanism. Our experiments demonstrate the effectiveness of Balanced-DPO, validating the theoretical findings and confirming that addressing gradient imbalance is key to improving DPO's performance, highlighting a promising direction for future research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A token-level correctness classifier trained with LoRA and then merged into the model boosts out-of-distribution factuality in summarization and translation.

  2. Nabla-R2D3: Effective and Efficient 3D Diffusion Alignment with 2D Rewards

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Nabla-R2D3 aligns 3D-native diffusion models with human preferences by backpropagating multi-view 2D reward gradients through the denoising process, improving reward without destroying the pretrained 3D prior.

Pith tools